Skip to content

just_dna_format.gene_validity

just_dna_format.gene_validity

The source-independent gene–disease validity table (0.6, RM24).

gene_validity.csv is the fifth derived-fact sidecar, and the second keyed on a gene rather than a variant. gene_metrics.csv answers how constrained is this gene; this one answers does variation in this gene cause this disease, and how sure is anyone. Filled by just-dna-enricher's gene-validity pass (ClinGen gene–disease validity, GenCC's aggregate of nineteen submitters), consumed and hashed by the compiler, never fetched by it.

Why a table and not columns on gene_metrics.csv. The grain is gene × disease term × mode of inheritance, not gene. Dosage sensitivity went the other way in 0.5 for exactly that reason — a haploinsufficiency rating is one value per gene, so it is two columns on the gene row — while RYR1 carries a definitive assertion for malignant hyperthermia and a separate one for a congenital myopathy, and neither is a property of the gene alone. Columns cannot hold two.

Three facts about the real files that decide the shape, all from reading them rather than their documentation (ClinGen's gene-validity/download, 3,659 rows, and GenCC's submissions-export-csv, 30,410 rows, both on 2026-08-13):

  • Mode of inheritance is part of the key. 59 (gene, disease) pairs in ClinGen carry two rows that differ only by MOI — ACO2 and mitochondrial disease, ACTA1 and nemaline myopathy — and keying without it silently keeps one. Adding (gene, disease, moi) leaves zero collisions.
  • GenCC is an aggregate, so submitter is in the key too. Nineteen submitters, and the same gene–disease pair is routinely asserted by several at different strengths (AARS1 / Charcot-Marie-Tooth 2N is Definitive to ClinGen and Strong to Labcorp). One row per submitter is the data; picking one would be the bare-triple mistake PharmVariantRow already paid for once.
  • The two vocabularies disagree in spelling and agree in meaning, so they are normalized here (vocab.VALID_GENE_VALIDITY, vocab.VALID_INHERITANCE_MODE) rather than stored verbatim. Verbatim is right for an identity — a star allele, an accession — and wrong for a value a consumer will filter or sort on, which is the same line ClinGen's dosage codes fall on.

No classification for a source that publishes none. A submitter can assert an association without grading it; the cell is then empty, which is this codebase's answer to an unknown everywhere else. It is not the same as no_known_disease_relationship, which is a graded verdict against.

GeneValidityRow

Bases: BaseModel

One curated gene–disease assertion, keyed by (gene, disease, mode of inheritance, submitter).

Standalone (not an AuthoredModel) for the same reason ResolutionRow/FrequencyRow/ GeneMetricsRow/LiteratureRow are — a machine-produced reference fact rather than an authored annotation — with extra="forbid" so a typo'd column is caught rather than silently dropped.

currency_group

currency_group(row: GeneValidityRow) -> tuple

The group a row's currency is decided within. See CURRENCY_GROUP_FIELDS.

Source code in schema/src/just_dna_format/gene_validity.py
def currency_group(row: "GeneValidityRow") -> tuple:
    """The group a row's currency is decided within. See `CURRENCY_GROUP_FIELDS`."""
    return tuple(getattr(row, name) for name in CURRENCY_GROUP_FIELDS)

classify_currency

classify_currency(
    rows: Sequence[GeneValidityRow],
) -> list[str | None]

Per row: CURRENT, SUPERSEDED, or None where nothing orders its group.

Newest classification_date wins, and nothing is deleted. That is S45's answer carried to a weaker signal, and taking it means accepting one thing this format had not accepted before — that a date is authoritative for currency. The concession is narrower than it looks: the date decides ordering and nothing else. It never says a classification is right, both rows stay in the file so the drift stays visible, and a consumer wanting the history still has it.

Publishing both facts and leaving the consumer to choose was the honest alternative and lost on one point: every consumer then implements the same date comparison, and they will not all implement it the same way.

Two edges, and both withhold rather than inventing an order:

  • a tie on classification_date — two curations of one claim stamped the same instant, and nothing in the row says which came second;
  • any row in the group carrying no date — an undated row cannot be placed, and calling it superseded because a dated one exists would assert a fact about a curation on the strength of a cell the source left empty.

In both cases every row in the group answers None. Breaking a tie on assertion_id was rejected: an identifier carries no chronology, and sorting on one would manufacture a winner out of a spelling.

A group of one is CURRENT, dated or not — there is nothing to order it against, and a lone assertion is the live one by construction. That is what keeps this quiet on the ordinary module: a check that cannot fail must not report (@tautology-zero), and almost every real group is a singleton.

Source code in schema/src/just_dna_format/gene_validity.py
def classify_currency(rows: "Sequence[GeneValidityRow]") -> list[str | None]:
    """Per row: `CURRENT`, `SUPERSEDED`, or `None` where nothing orders its group.

    **Newest `classification_date` wins, and nothing is deleted.** That is S45's answer carried to a
    weaker signal, and taking it means accepting one thing this format had not accepted before — that
    a date is authoritative for currency. The concession is narrower than it looks: the date decides
    *ordering* and nothing else. It never says a classification is right, both rows stay in the file so
    the drift stays visible, and a consumer wanting the history still has it.

    Publishing both facts and leaving the consumer to choose was the honest alternative and lost on one
    point: every consumer then implements the same date comparison, and they will not all implement it
    the same way.

    **Two edges, and both withhold** rather than inventing an order:

    * a **tie** on `classification_date` — two curations of one claim stamped the same instant, and
      nothing in the row says which came second;
    * **any row in the group carrying no date** — an undated row cannot be placed, and calling it
      superseded because a dated one exists would assert a fact about a curation on the strength of a
      cell the source left empty.

    In both cases every row in the group answers `None`. Breaking a tie on `assertion_id` was
    rejected: an identifier carries no chronology, and sorting on one would manufacture a winner out
    of a spelling.

    **A group of one is `CURRENT`**, dated or not — there is nothing to order it against, and a lone
    assertion is the live one by construction. That is what keeps this quiet on the ordinary module:
    a check that cannot fail must not report (`@tautology-zero`), and almost every real group is a
    singleton.
    """
    groups: dict[tuple, list[int]] = {}
    for index, row in enumerate(rows):
        groups.setdefault(currency_group(row), []).append(index)

    verdicts: list[str | None] = [None] * len(rows)
    for members in groups.values():
        if len(members) == 1:
            verdicts[members[0]] = CURRENT
            continue
        dates = [rows[i].classification_date for i in members]
        if any(d is None for d in dates):
            continue  # undated row in the group — nothing can be placed
        newest = max(dates)
        if dates.count(newest) > 1:
            continue  # a tie orders nothing
        for index in members:
            verdicts[index] = CURRENT if rows[index].classification_date == newest else SUPERSEDED
    return verdicts

superseded_groups

superseded_groups(
    rows: Sequence[GeneValidityRow],
) -> list[tuple]

The groups where a later curation replaced an earlier one, in first-seen order.

Deterministic order because a warning built from it is a published string: insertion order is the order the groups were first met, never a set iteration (@dont-discard-computed next door — the ordering rules the compiler keeps).

Source code in schema/src/just_dna_format/gene_validity.py
def superseded_groups(rows: "Sequence[GeneValidityRow]") -> list[tuple]:
    """The groups where a later curation replaced an earlier one, in first-seen order.

    Deterministic order because a warning built from it is a published string: insertion order is the
    order the groups were first met, never a set iteration (`@dont-discard-computed` next door — the
    ordering rules the compiler keeps).
    """
    verdicts = classify_currency(rows)
    seen: list[tuple] = []
    marked: set[tuple] = set()
    for row, verdict in zip(rows, verdicts, strict=True):
        key = currency_group(row)
        if verdict == SUPERSEDED and key not in marked:
            marked.add(key)
            seen.append(key)
    return seen

undecidable_groups

undecidable_groups(
    rows: Sequence[GeneValidityRow],
) -> list[tuple]

The multi-row groups nothing orders — a tie, or a member with no classification_date.

Reported separately from the superseded ones, because they ask the reader for different things: a superseded row is the archive having moved on, and an unorderable group is the archive not having said enough to tell. Collapsing them would publish one number meaning two facts, which is the shape @unreachable-not-absent exists about.

Source code in schema/src/just_dna_format/gene_validity.py
def undecidable_groups(rows: "Sequence[GeneValidityRow]") -> list[tuple]:
    """The multi-row groups nothing orders — a tie, or a member with no `classification_date`.

    Reported **separately** from the superseded ones, because they ask the reader for different
    things: a superseded row is the archive having moved on, and an unorderable group is the archive
    not having said enough to tell. Collapsing them would publish one number meaning two facts, which
    is the shape `@unreachable-not-absent` exists about.
    """
    verdicts = classify_currency(rows)
    counts: dict[tuple, int] = {}
    order: list[tuple] = []
    undecided: set[tuple] = set()
    for row, verdict in zip(rows, verdicts, strict=True):
        key = currency_group(row)
        if key not in counts:
            counts[key] = 0
            order.append(key)
        counts[key] += 1
        if verdict is None:
            undecided.add(key)
    return [key for key in order if key in undecided and counts[key] > 1]