Skip to content

just_dna_format.gene_metrics

just_dna_format.gene_metrics

The source-independent gene-constraint table (0.5).

gene_metrics.csv is the gene-level sibling of frequencies.csv: population-constraint facts (pLI, LOEUF, missense Z) for the genes a module already mentions. Filled by just-dna-enricher's gnomAD pass — from a small offline snapshot first, the live API second — consumed and hashed by the compiler, never fetched by it.

Gene-level and variant-level facts get separate tables rather than gene metrics repeated on every variant row (Principle 5, and the "one CSV = one concern" rule): a module with forty variants across six genes carries six rows here, not forty duplicated ones, and the two axes stay independently updatable.

Unlike the frequency table's integer counts, these are floats by nature — a constraint score has no integer form to round-trip through — so the canonical-formatting discipline applies to the whole row rather than to one column, and the reverse writer's round-trip is covered by a test.

GeneMetricsRow

Bases: BaseModel

One gene's population-constraint metrics, keyed by the gene symbol a module authored.

Standalone (not an AuthoredModel) for the same reason ResolutionRow/FrequencyRow are — a derived reference fact, not an authored annotation — with extra="forbid" so a typo'd column is caught rather than dropped.

normalize_constraint_flags

normalize_constraint_flags(value: object) -> str | None

gnomAD's caveat list, from either producer, as one pipe-joined string or None (RM110).

The two routes hand this the same fact in two encodings, and until 0.7 only one of them was read. The live GraphQL field returns flags as a JSON array, which arrives here as a Python list; the bulk v4.1 TSV writes the array literal into the cell, so the snapshot route stored the two-character string "[]" and the four-token string '["no_exp_lof","no_exp_mis","no_exp_syn","no_variants"]' verbatim. Neither is a pipe-joined flag list, and constraint_flags is inside GENE_METRICS_FACT_FIELDS, so the same gene fetched two ways produced two gene_metrics.signature values.

Measured over the published v4.1 snapshot rather than estimated: of 18,111 rows, not one is null or empty — 17,403 carry "[]" and 708 carry a real array literal — so a consumer writing the obvious if row.constraint_flags: read 100% of snapshot rows as flagged where the true figure is 3.9%, and one splitting on | got a single bogus token instead of two flags.

So the empty case is only half of it: the non-empty cells need parsing too, which is why this takes the cell apart rather than testing it against a null set. "[]" → None alone would have left 708 rows still lying about their own shape.

Total and idempotent, because it runs on both legs and on a snapshot rebuilt after this shipped: a list joins, a JSON array literal parses then joins, an already-pipe-joined string splits then re-joins, and everything empty answers None rather than "" — the house rule that an absence is not a value. Sorted, because the live producer has always sorted and a signature must not depend on which route filled the cell.

A string that starts like an array and does not parse is kept verbatim rather than dropped or guessed at: this normalizes an encoding it can recognise and does not invent a reading for one it cannot, and a cell that survives here unchanged is visible to whoever reads the table.

It lives in this tier, not in the enricher, because the decision was that the normalization goes in the CELL. A helper the fetching tier called would fix rows written after 0.7 and leave every gene_metrics.csv already carrying "[]" — including the one in this repo's own reference corpus — still contradicting the column's description and still hashing differently from the same gene fetched the other way. Bound as a mode="before" validator below, it reaches every producer there will ever be, including a hand-written table and a re-read of a file some earlier release wrote.

Source code in schema/src/just_dna_format/gene_metrics.py
def normalize_constraint_flags(value: object) -> str | None:
    """gnomAD's caveat list, from either producer, as one pipe-joined string or `None` (RM110).

    **The two routes hand this the same fact in two encodings**, and until 0.7 only one of them was
    read. The live GraphQL field returns `flags` as a JSON array, which arrives here as a Python
    `list`; the bulk v4.1 TSV writes the array *literal* into the cell, so the snapshot route stored
    the two-character string ``"[]"`` and the four-token string
    ``'["no_exp_lof","no_exp_mis","no_exp_syn","no_variants"]'`` verbatim. Neither is a pipe-joined
    flag list, and `constraint_flags` is inside `GENE_METRICS_FACT_FIELDS`, so the same gene fetched
    two ways produced two `gene_metrics.signature` values.

    **Measured over the published v4.1 snapshot rather than estimated**: of 18,111 rows, *not one* is
    null or empty — 17,403 carry ``"[]"`` and 708 carry a real array literal — so a consumer writing
    the obvious ``if row.constraint_flags:`` read **100%** of snapshot rows as flagged where the true
    figure is 3.9%, and one splitting on ``|`` got a single bogus token instead of two flags.

    So the empty case is only half of it: the non-empty cells need parsing too, which is why this
    takes the cell apart rather than testing it against a null set. `"[]"` → `None` alone would have
    left 708 rows still lying about their own shape.

    Total and idempotent, because it runs on both legs and on a snapshot rebuilt after this shipped:
    a list joins, a JSON array literal parses then joins, an already-pipe-joined string splits then
    re-joins, and everything empty answers `None` rather than `""` — the house rule that an absence is
    not a value. Sorted, because the live producer has always sorted and a signature must not depend
    on which route filled the cell.

    A string that starts like an array and does not parse is kept verbatim rather than dropped or
    guessed at: this normalizes an encoding it can recognise and does not invent a reading for one it
    cannot, and a cell that survives here unchanged is visible to whoever reads the table.

    **It lives in this tier, not in the enricher, because the decision was that the normalization
    goes in the CELL.** A helper the fetching tier called would fix rows written after 0.7 and
    leave every `gene_metrics.csv` already carrying `"[]"` — including the one in this repo's own
    reference corpus — still contradicting the column's description and still hashing differently
    from the same gene fetched the other way. Bound as a `mode="before"` validator below, it
    reaches every producer there will ever be, including a hand-written table and a re-read of a
    file some earlier release wrote.
    """
    if value is None:
        return None
    if isinstance(value, (list, tuple, set)):
        tokens: list[str] = [str(v).strip() for v in value]
    else:
        text = str(value).strip()
        if text in _FLAG_NULLS:
            return None
        if text.startswith("[") and text.endswith("]"):
            try:
                parsed = json.loads(text)
            except ValueError:
                return text
            if not isinstance(parsed, list):
                return text
            tokens = [str(v).strip() for v in parsed]
        else:
            tokens = [part.strip() for part in text.split("|")]
    kept = sorted(t for t in tokens if t)
    return "|".join(kept) if kept else None