just_dna_format.gene_metrics¶
just_dna_format.gene_metrics ¶
The source-independent gene-constraint table (0.5).
gene_metrics.csv is the gene-level sibling of frequencies.csv: population-constraint facts
(pLI, LOEUF, missense Z) for the genes a module already mentions. Filled by just-dna-enricher's
gnomAD pass — from a small offline snapshot first, the live API second — consumed and hashed by the
compiler, never fetched by it.
Gene-level and variant-level facts get separate tables rather than gene metrics repeated on every variant row (Principle 5, and the "one CSV = one concern" rule): a module with forty variants across six genes carries six rows here, not forty duplicated ones, and the two axes stay independently updatable.
Unlike the frequency table's integer counts, these are floats by nature — a constraint score has no integer form to round-trip through — so the canonical-formatting discipline applies to the whole row rather than to one column, and the reverse writer's round-trip is covered by a test.
GeneMetricsRow ¶
Bases: BaseModel
One gene's population-constraint metrics, keyed by the gene symbol a module authored.
Standalone (not an AuthoredModel) for the same reason ResolutionRow/FrequencyRow are — a
derived reference fact, not an authored annotation — with extra="forbid" so a typo'd column is
caught rather than dropped.
normalize_constraint_flags ¶
gnomAD's caveat list, from either producer, as one pipe-joined string or None (RM110).
The two routes hand this the same fact in two encodings, and until 0.7 only one of them was
read. The live GraphQL field returns flags as a JSON array, which arrives here as a Python
list; the bulk v4.1 TSV writes the array literal into the cell, so the snapshot route stored
the two-character string "[]" and the four-token string
'["no_exp_lof","no_exp_mis","no_exp_syn","no_variants"]' verbatim. Neither is a pipe-joined
flag list, and constraint_flags is inside GENE_METRICS_FACT_FIELDS, so the same gene fetched
two ways produced two gene_metrics.signature values.
Measured over the published v4.1 snapshot rather than estimated: of 18,111 rows, not one is
null or empty — 17,403 carry "[]" and 708 carry a real array literal — so a consumer writing
the obvious if row.constraint_flags: read 100% of snapshot rows as flagged where the true
figure is 3.9%, and one splitting on | got a single bogus token instead of two flags.
So the empty case is only half of it: the non-empty cells need parsing too, which is why this
takes the cell apart rather than testing it against a null set. "[]" → None alone would have
left 708 rows still lying about their own shape.
Total and idempotent, because it runs on both legs and on a snapshot rebuilt after this shipped:
a list joins, a JSON array literal parses then joins, an already-pipe-joined string splits then
re-joins, and everything empty answers None rather than "" — the house rule that an absence is
not a value. Sorted, because the live producer has always sorted and a signature must not depend
on which route filled the cell.
A string that starts like an array and does not parse is kept verbatim rather than dropped or guessed at: this normalizes an encoding it can recognise and does not invent a reading for one it cannot, and a cell that survives here unchanged is visible to whoever reads the table.
It lives in this tier, not in the enricher, because the decision was that the normalization
goes in the CELL. A helper the fetching tier called would fix rows written after 0.7 and
leave every gene_metrics.csv already carrying "[]" — including the one in this repo's own
reference corpus — still contradicting the column's description and still hashing differently
from the same gene fetched the other way. Bound as a mode="before" validator below, it
reaches every producer there will ever be, including a hand-written table and a re-read of a
file some earlier release wrote.