just_dna_format.gene_validity¶
just_dna_format.gene_validity ¶
The source-independent gene–disease validity table (0.6, RM24).
gene_validity.csv is the fifth derived-fact sidecar, and the second keyed on a gene rather than a
variant. gene_metrics.csv answers how constrained is this gene; this one answers does variation
in this gene cause this disease, and how sure is anyone. Filled by just-dna-enricher's gene-validity
pass (ClinGen gene–disease validity, GenCC's aggregate of nineteen submitters), consumed and hashed by
the compiler, never fetched by it.
Why a table and not columns on gene_metrics.csv. The grain is gene × disease term × mode of
inheritance, not gene. Dosage sensitivity went the other way in 0.5 for exactly that reason — a
haploinsufficiency rating is one value per gene, so it is two columns on the gene row — while
RYR1 carries a definitive assertion for malignant hyperthermia and a separate one for a congenital
myopathy, and neither is a property of the gene alone. Columns cannot hold two.
Three facts about the real files that decide the shape, all from reading them rather than their
documentation (ClinGen's gene-validity/download, 3,659 rows, and GenCC's submissions-export-csv,
30,410 rows, both on 2026-08-13):
- Mode of inheritance is part of the key. 59 (gene, disease) pairs in ClinGen carry two rows that
differ only by MOI — ACO2 and mitochondrial disease, ACTA1 and nemaline myopathy — and keying
without it silently keeps one. Adding
(gene, disease, moi)leaves zero collisions. - GenCC is an aggregate, so
submitteris in the key too. Nineteen submitters, and the same gene–disease pair is routinely asserted by several at different strengths (AARS1 / Charcot-Marie-Tooth 2N is Definitive to ClinGen and Strong to Labcorp). One row per submitter is the data; picking one would be the bare-triple mistakePharmVariantRowalready paid for once. - The two vocabularies disagree in spelling and agree in meaning, so they are normalized here
(
vocab.VALID_GENE_VALIDITY,vocab.VALID_INHERITANCE_MODE) rather than stored verbatim. Verbatim is right for an identity — a star allele, an accession — and wrong for a value a consumer will filter or sort on, which is the same line ClinGen's dosage codes fall on.
No classification for a source that publishes none. A submitter can assert an association
without grading it; the cell is then empty, which is this codebase's answer to an unknown everywhere
else. It is not the same as no_known_disease_relationship, which is a graded verdict against.
GeneValidityRow ¶
Bases: BaseModel
One curated gene–disease assertion, keyed by (gene, disease, mode of inheritance, submitter).
Standalone (not an AuthoredModel) for the same reason ResolutionRow/FrequencyRow/
GeneMetricsRow/LiteratureRow are — a machine-produced reference fact rather than an authored
annotation — with extra="forbid" so a typo'd column is caught rather than silently dropped.
currency_group ¶
The group a row's currency is decided within. See CURRENCY_GROUP_FIELDS.
classify_currency ¶
Per row: CURRENT, SUPERSEDED, or None where nothing orders its group.
Newest classification_date wins, and nothing is deleted. That is S45's answer carried to a
weaker signal, and taking it means accepting one thing this format had not accepted before — that
a date is authoritative for currency. The concession is narrower than it looks: the date decides
ordering and nothing else. It never says a classification is right, both rows stay in the file so
the drift stays visible, and a consumer wanting the history still has it.
Publishing both facts and leaving the consumer to choose was the honest alternative and lost on one point: every consumer then implements the same date comparison, and they will not all implement it the same way.
Two edges, and both withhold rather than inventing an order:
- a tie on
classification_date— two curations of one claim stamped the same instant, and nothing in the row says which came second; - any row in the group carrying no date — an undated row cannot be placed, and calling it superseded because a dated one exists would assert a fact about a curation on the strength of a cell the source left empty.
In both cases every row in the group answers None. Breaking a tie on assertion_id was
rejected: an identifier carries no chronology, and sorting on one would manufacture a winner out
of a spelling.
A group of one is CURRENT, dated or not — there is nothing to order it against, and a lone
assertion is the live one by construction. That is what keeps this quiet on the ordinary module:
a check that cannot fail must not report (@tautology-zero), and almost every real group is a
singleton.
Source code in schema/src/just_dna_format/gene_validity.py
superseded_groups ¶
The groups where a later curation replaced an earlier one, in first-seen order.
Deterministic order because a warning built from it is a published string: insertion order is the
order the groups were first met, never a set iteration (@dont-discard-computed next door — the
ordering rules the compiler keeps).
Source code in schema/src/just_dna_format/gene_validity.py
undecidable_groups ¶
The multi-row groups nothing orders — a tie, or a member with no classification_date.
Reported separately from the superseded ones, because they ask the reader for different
things: a superseded row is the archive having moved on, and an unorderable group is the archive
not having said enough to tell. Collapsing them would publish one number meaning two facts, which
is the shape @unreachable-not-absent exists about.