Skip to content

just_dna_enricher.identifiers

just_dna_enricher.identifiers

Identifier currency — are the labels a module keys on still the ones its registries serve?

Four registries, one shape: ask whether an authored identifier is current, and report what the registry says instead. Nothing here repairs anything, and for rsIDs that is not a stylistic preference but a hard requirement — see check_rsids.

The fourth is the PGS Catalog (RM163), and it arrived last for a reason worth keeping: pgs_id is the column PgsRow is keyed on, so it was the one authored identifier in the format that no pass ever put to its registry. Its client and its source-shape knowledge live in pgs.py; the statuses, the comparison and the two attestations are here, beside the other three.

This is the generalization of COMPILER.md's "is the source stale?" blind spot from datasets to identifiers. The compiler cannot see the world move; a dbSNP merge, an EFO term retirement and an HGNC symbol rename all leave a module perfectly well-formed and quietly out of date.

dbSNP: three states, and one of them means two opposite things. Probed rather than assumed:

State esummary db=snp
live snp_id == requested, merged_sort='0'
merged snp_id != requested, merged_sort='1'
absent {'uid': …, 'error': 'cannot get document summary'}

The absent response for rs11273140 — an rsID withdrawn after a clustering error — is byte-identical to the one for rs2000000000, which was never assigned. For an author these mean opposite things (fix the typo, versus the variant itself was retracted and the annotation resting on it may be worthless), and no live endpoint separates them: esearch has no withdrawn filter, and latest_release/misc/rs_unsupported_b157.txt looks like a withdrawn registry but is a one-off build-157 ClinVar-parsing incident list that does not contain rs11273140. So the message names both readings and asserts neither. Guessing "typo" would send an author to fix the wrong thing.

NCBI is the oracle, not Ensembl. Ensembl REST resolves some merges (rs77121243 → rs334) and returns HTTP 400 on others (rs3216883, which dbSNP correctly reports as merged into rs3051860), so Ensembl alone would misclassify a merged rsID as unresolvable.

IdentifierCheckError

Bases: RuntimeError

A currency lookup failed outright (transport or an unusable response).

IdentifierUnavailable

Bases: IdentifierCheckError

A source this pass asks was not reachable, so the question could not be put (RM101).

A subclass rather than a second exception, for AcmgListUnavailable's reason: every existing except IdentifierCheckError still catches it, and the two histories this module can fail with stop being one type. Only this one means a source was asked; an invalid local variants.csv is the plain parent, and a caller separating them had to read __cause__ to do it.

It carries no skip member, unlike AcmgListUnavailable, and the reason is not that nothing attests — check-identifiers does, and getting this wrong once is how the first draft of this class shipped a false docstring. The difference is where the member is decided. acmg raises from several places that mean different skips, so the raiser is the only thing that knows which; here every raise means the same one, and unreachable_records() already owns it for all three checks at once. Putting it on the exception would be a second place to keep the same fact.

RsidStatus dataclass

RsidStatus(
    rsid: str, state: str, current: str | None = None
)

What dbSNP currently says about one authored rsID.

is_fatal property

is_fatal: bool

Whether this refuses in both modes rather than only under strict.

Only withdrawn does. A merged rsID still resolves to the right locus (the module is dated, not wrong) and an absent one may be a typo, a very new id, or API lag — all leave the annotation itself intact. A withdrawn one is dbSNP repudiating the variant, so the annotation may be describing nothing.

Nothing in this module produces withdrawn today: a retraction is byte-identical to a never-assigned id through every live endpoint, so classify_rsid reports absent and names both readings. The state exists for a curator who has established the retraction by hand, and for a future source that can tell them apart.

TraitStatus dataclass

TraitStatus(
    curie: str,
    state: str,
    label: str | None = None,
    replaced_by: str | None = None,
)

What the ontology currently says about one authored trait id.

GeneStatus dataclass

GeneStatus(
    symbol: str,
    state: str,
    current: str | None = None,
    hgnc_id: str | None = None,
    location: str | None = None,
)

What HGNC currently says about one authored gene symbol.

chromosome property

chromosome: str | None

The contig the band names, or None when HGNC gave none or it cannot be read.

Band syntax is <chrom><arm><band>, so the contig is everything before the first p/q — except mitochondria, which HGNC spells out and this codebase calls MT. Alternate-reference loci (19q13.42 alternate reference locus) keep their leading band, so the split still works. None for anything unparsed, because a guess here becomes a false accusation about a row.

GeneLocusConflict dataclass

GeneLocusConflict(
    gene: str,
    gene_chrom: str,
    variant_key: str,
    variant_chrom: str,
)

An authored gene whose HGNC chromosome is not the chromosome the row's variant sits on.

The relationship is false while both halves are individually true — the rsID resolves, the symbol is approved — which is why nothing caught it before (S24). It is the signature of a generated citation: a real gene name beside an invented rs number, which resolves anyway because dbSNP is dense enough that almost any seven-digit number hits something. Four of one reporter's seven rows were this, each pairing a real symbol with a variant on another chromosome.

PgsStatus dataclass

PgsStatus(
    pgs_id: str,
    state: str,
    name: str | None = None,
    date_release: str | None = None,
    trait_efo_ids: tuple[str, ...] = (),
    variants_number: int | None = None,
    license: str | None = None,
)

What the PGS Catalog currently says about one authored pgs_id (RM163).

Three states, and the first of them is settled before any request goes out. The REST surface answers HTTP 200 with an empty body for a never-assigned accession and for a malformed one alike, so the service itself cannot separate them — but the format's own grammar can, and asking the Catalog about a string that is not an accession would spend a request to learn nothing.

PgsDrift dataclass

PgsDrift(
    pgs_id: str,
    field_name: str,
    authored: str,
    published: str,
)

An authored cell beside a pgs_id that the score's own Catalog record contradicts (RM163).

Reported, never repaired (@enrichment-is-validation): which of the two is right is not something this tier can know, and both training_ancestry and training_cohort are authored judgements a curator may have made deliberately against the Catalog's own summary.

PgsComparison dataclass

PgsComparison(
    drift: list[PgsDrift] = list(),
    compared: list[tuple[str, str, str]] = list(),
    withheld: dict[tuple[str, str, str], str] = dict(),
    authored: list[tuple[str, str, str]] = list(),
)

What the two-field drift check compared, what it withheld, and why — the denominator (RM163).

withheld is the point of the type, and it is the @dont-discard-computed rule: a cell this check looked at and could not settle is neither a finding nor a clean comparison, and folding it into either would publish a coverage figure larger or smaller than the question that was put.

IdentifierReport dataclass

IdentifierReport(
    rsids: list[RsidStatus] = list(),
    traits: list[TraitStatus] = list(),
    genes: list[GeneStatus] = list(),
    gene_loci: list[GeneLocusConflict] = list(),
    gene_loci_not_checked: str | None = None,
    gene_loci_compared: int = 0,
    trait_tables_read: list[str] = list(),
    trait_tables_not_read: dict[str, str] = dict(),
    gene_tables_read: list[str] = list(),
    gene_tables_not_read: dict[str, str] = dict(),
    pgs: list[PgsStatus] = list(),
    pgs_metadata: PgsComparison = PgsComparison(),
    pgs_tables_read: list[str] = list(),
    pgs_tables_not_read: dict[str, str] = dict(),
    pgs_release: str | None = None,
    pgs_not_checked: tuple[str, str] | None = None,
    pgs_sources: list[SourceRow] = list(),
)

unreadable_tables property

unreadable_tables: dict[str, str]

Tables that exist and would not parse, across both columns — never merely absent ones.

The actionable half: an absent optional table is the normal shape of every module, while one that is present and unreadable means ids the module carries went unchecked.

clean property

clean: Verdict

Whether every identifier this pass put to a registry is one the registry still serves.

pgs_metadata.drift is deliberately NOT in here, and stale_pgs deliberately is. The CLI gates --strict's exit code on this property, and the two findings are different kinds: an accession the Catalog does not hold is a broken identifier, exactly like a retired HGNC symbol; a training_ancestry the Catalog's own summary disagrees with is two authorities differing about a curated judgement, and PgsDrift.__str__ tells the author in as many words to leave it if their curation is deliberate. Failing a build on that would make the format arbitrate between its own sources with no way for the author to clear it — the reason @clinsig-never-escalates and @a-source-recuring-is-not-a-strict-matter keep the same shape out of the strict gate. It is reported in both modes; see metadata_disagrees.

A Verdict rather than a bool since 2026-09-13 (RM235), for the reason verdict.py states: --strict gates an exit code on this, so a run that could not certify has to be a no, and the reason for the no rides beside the verdict instead of inside it.

tables_unreadable is the new arm and the one behaviour change. A table that carries identifiers and will not parse leaves those ids unchecked, and the five lists above are empty for exactly the same reason they are empty when everything agreed — so clean answered True over ids nobody looked at. The command has printed the unreadable table since S86 and then exited 0 under --strict with all identifiers current above it: the diagnosis was there and the verdict disagreed with it. A library caller reading this property sees only the verdict, which is how the sibling defect on AcmgReport arrived (S100 — these are wrapped as MCP tools, so the dataclass is what is read and never the terminal).

What is deliberately not an arm. A check the caller switched off, and a module carrying no id-bearing table at all, are honest non-answers rather than errors: they keep the verdict true and are reported through the read/not-read rosters that already exist, so --strict --no-traits does not start failing builds for doing as it was told. unreachable is not an arm either, and could not be — a registry outage raises IdentifierUnavailable before this report exists, and the CLI already exits 1 there with an unreachable attestation for all five checks. The item that proposed widening this property for outages had measured hand-built models rather than the real path (@a-disagreement-with-a-document-may-be-in-the-instrument).

metadata_disagrees property

metadata_disagrees: bool

Whether a source disagreed with an authored cell — reported, never gated.

Separate from clean so a caller can say checked, and something differs without the exit code claiming the module is broken. The CLI uses it to withhold the green line.

OntologyClient dataclass

OntologyClient(
    ols4_base: str = DEFAULT_OLS4_BASE,
    hgnc_base: str = DEFAULT_HGNC_BASE,
    min_request_interval: float = 0.2,
    timeout: float = 30.0,
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

OLS4 + HGNC lookups. One class because both are unbatched GET-per-id REST services.

Neither publishes a documented rate limit, so the gate is a courtesy rather than a budget — but it is here rather than absent, because a module with two hundred genes would otherwise open two hundred connections as fast as it could.

trait

trait(curie: str) -> TraitStatus

Existence, obsolescence and the replacement term for one ontology CURIE.

Source code in enricher/src/just_dna_enricher/identifiers.py
def trait(self, curie: str) -> TraitStatus:
    """Existence, obsolescence and the replacement term for one ontology CURIE."""
    match = _CURIE.match(curie.strip())
    if match is None:
        return TraitStatus(curie=curie, state="unchecked")
    prefix, local = match.group(1).upper(), match.group(2)
    route = _ONTOLOGY_IRI.get(prefix)
    if route is None:
        return TraitStatus(curie=curie, state="unchecked")
    ontology, iri_prefix = route
    response = self._get(
        f"{self.ols4_base.rstrip('/')}/ontologies/{ontology}/terms",
        {"iri": f"{iri_prefix}{local}"},
    )
    if response.status_code == 404:
        return TraitStatus(curie=curie, state="absent")
    terms = (_json(response).get("_embedded") or {}).get("terms") or []
    if not terms:
        return TraitStatus(curie=curie, state="absent")
    term = terms[0]
    if term.get("is_obsolete"):
        return TraitStatus(
            curie=curie,
            state="obsolete",
            label=term.get("label"),
            replaced_by=_curie_from_iri(term.get("term_replaced_by")),
        )
    return TraitStatus(curie=curie, state="current", label=term.get("label"))

gene

gene(symbol: str) -> GeneStatus

Approved / retired / unknown for one gene symbol.

Uses the exact fetch/ endpoints, never search/: search/BRCA1 is fuzzy and returns 19 hits including ABRAXAS1, which would make every symbol look ambiguous. fetch/symbol answers currency; fetch/prev_symbol answers retirement, and is only consulted when the first misses.

Source code in enricher/src/just_dna_enricher/identifiers.py
def gene(self, symbol: str) -> GeneStatus:
    """Approved / retired / unknown for one gene symbol.

    Uses the **exact** `fetch/` endpoints, never `search/`: `search/BRCA1` is fuzzy and returns 19
    hits including `ABRAXAS1`, which would make every symbol look ambiguous. `fetch/symbol` answers
    currency; `fetch/prev_symbol` answers retirement, and is only consulted when the first misses.
    """
    approved = self._hgnc(f"symbol/{symbol}")
    if approved:
        return GeneStatus(
            symbol=symbol,
            state="approved",
            current=approved[0].get("symbol"),
            hgnc_id=approved[0].get("hgnc_id"),
            location=approved[0].get("location"),
        )
    previous = self._hgnc(f"prev_symbol/{symbol}")
    if previous:
        return GeneStatus(
            symbol=symbol,
            state="retired",
            current=previous[0].get("symbol"),
            hgnc_id=previous[0].get("hgnc_id"),
            location=previous[0].get("location"),
        )
    return GeneStatus(symbol=symbol, state="unknown")

IdentifierRoster dataclass

IdentifierRoster(
    ids: list[str] = list(),
    read: list[str] = list(),
    not_read: dict[str, str] = dict(),
    read_errors: dict[str, str] = dict(),
)

The ids an authored spec names, with the tables that were and were not read (S86).

not_read is the point of the type. A bare roster cannot distinguish this module declares no trait from this module's traits are in a table nobody read, and both render as 0 checked, 0 clean, 0 flagged — which reads as a clean run and let a module ship a retired CURIE with every gate green. Same three-valued rule as gene_loci_not_checked two fields down in the report, and as unconsulted_rsids one tier down: a question never put is not an answer.

unreadable property

unreadable: dict[str, str]

Only the tables that exist and failed to parse — the half worth warning about.

classify_rsid

classify_rsid(
    rsid: str, record: dict[str, Any]
) -> RsidStatus

One esummary record → live / merged / absent.

Split out from the batch so the state machine is testable against the recorded payload directly. snp_id comes back as an int and the uid as a string, so both sides are normalized before comparison — comparing them raw would classify every live rsID as merged.

Source code in enricher/src/just_dna_enricher/identifiers.py
def classify_rsid(rsid: str, record: dict[str, Any]) -> RsidStatus:
    """One esummary record → `live` / `merged` / `absent`.

    Split out from the batch so the state machine is testable against the recorded payload directly.
    `snp_id` comes back as an int and the uid as a string, so both sides are normalized before
    comparison — comparing them raw would classify every live rsID as merged.
    """
    if is_missing(record):
        return RsidStatus(rsid=rsid, state="absent")
    requested = rsid.removeprefix("rs")
    served = record.get("snp_id")
    merged = str(record.get("merged_sort") or "0") == "1"
    if served is not None and str(served) != requested:
        return RsidStatus(rsid=rsid, state="merged", current=f"rs{served}")
    if merged:
        # dbSNP flagged a merge but served the requested id back — record it rather than silently
        # calling it live, since the two signals disagree.
        logger.info("%s: merged_sort=1 but snp_id is the requested id; treated as live", rsid)
    return RsidStatus(rsid=rsid, state="live")

check_rsids

check_rsids(
    rsids: list[str], *, client: EutilsClient | None = None
) -> list[RsidStatus]

Current/merged/absent for each authored rsID, batched through esummary db=snp.

Report, never repair — and here that is load-bearing rather than tidy. weights.parquet carries both variant_key and rsid, and for an rsid-authored row they are the same label. Writing the merged-into id back would not be a one-time digest move but an identity migration performed by a network lookup: reverse would emit the new rsID into variants.csv, the next compile would key on it, and variant_key itself would change with no authored edit anywhere. The module's identity would drift and the round-trip would stop being a fixed point (Principle 7).

Source code in enricher/src/just_dna_enricher/identifiers.py
def check_rsids(rsids: list[str], *, client: EutilsClient | None = None) -> list[RsidStatus]:
    """Current/merged/absent for each authored rsID, batched through `esummary db=snp`.

    **Report, never repair — and here that is load-bearing rather than tidy.** `weights.parquet`
    carries both `variant_key` and `rsid`, and for an rsid-authored row they are the same label.
    Writing the merged-into id back would not be a one-time digest move but an *identity migration
    performed by a network lookup*: reverse would emit the new rsID into `variants.csv`, the next
    compile would key on it, and `variant_key` itself would change with no authored edit anywhere.
    The module's identity would drift and the round-trip would stop being a fixed point (Principle 7).
    """
    wanted = dedupe(r for r in rsids if r)
    if not wanted:
        return []
    owned = client is None
    eutils = client or EutilsClient()
    try:
        records = eutils.esummary("snp", [r.removeprefix("rs") for r in wanted])
    except EutilsError as exc:
        # dbSNP is a *different* module's client, so `EutilsError` is not this module's type
        # and `except IdentifierCheckError` around `check_rsids` never fired for a 5xx.
        raise IdentifierUnavailable(f"dbSNP could not be reached: {exc}") from exc
    finally:
        if owned:
            eutils.close()
    return [classify_rsid(rsid, records.get(rsid.removeprefix("rs"), {})) for rsid in wanted]

module_trait_ids

module_trait_ids(variants: list[VariantRow]) -> list[str]

Every ontology CURIE the module's variants.csv rows name, in first-occurrence order.

trait_efo_id is a multi-valued cell (comma/semicolon/pipe-separated), so one row can name several traits.

This is the roster of one table, and eleven authored tables carry the column. Kept as-is because a caller holding rows is the only thing it can serve, and widened where the tables are actually reachable: authored_identifiers below takes a spec directory and reads all of them. A caller passing variants= gets this narrow roster and is told so (S86).

Source code in enricher/src/just_dna_enricher/identifiers.py
def module_trait_ids(variants: list[VariantRow]) -> list[str]:
    """Every ontology CURIE the module's **`variants.csv` rows** name, in first-occurrence order.

    `trait_efo_id` is a multi-valued cell (comma/semicolon/pipe-separated), so one row can name
    several traits.

    **This is the roster of one table, and eleven authored tables carry the column.** Kept as-is
    because a caller holding rows is the only thing it can serve, and widened where the tables are
    actually reachable: `authored_identifiers` below takes a spec directory and reads all of them.
    A caller passing `variants=` gets this narrow roster and is told so (S86).
    """
    out: list[str] = []
    for variant in variants:
        if not variant.trait_efo_id:
            continue
        out.extend(part.strip() for part in MULTI_SEP.split(variant.trait_efo_id) if part.strip())
    return dedupe(out)

authored_identifiers

authored_identifiers(
    spec_dir: Path, column: str
) -> IdentifierRoster

Every value of column across every authored table that declares it, and what was read.

The roster check_identifiers runs on when it is given a spec directory. Rows are loaded with the same load_csv_rows the compiler uses, so a table this cannot parse is reported as unread rather than skipped silently — the distinction the whole item is about.

Multi-valued cells are split on MULTI_SEP for both columns. A gene cell is single-valued in practice, and splitting it costs nothing and cannot mis-read one: a symbol contains no separator.

Source code in enricher/src/just_dna_enricher/identifiers.py
def authored_identifiers(spec_dir: Path, column: str) -> IdentifierRoster:
    """Every value of `column` across every authored table that declares it, and what was read.

    The roster `check_identifiers` runs on when it is given a spec directory. Rows are loaded with the
    same `load_csv_rows` the compiler uses, so a table this cannot parse is reported as unread rather
    than skipped silently — the distinction the whole item is about.

    Multi-valued cells are split on `MULTI_SEP` for **both** columns. A `gene` cell is single-valued in
    practice, and splitting it costs nothing and cannot mis-read one: a symbol contains no separator.
    """
    return authored_rows(spec_dir, column)[1]

authored_rows

authored_rows(
    spec_dir: Path, column: str
) -> tuple[list[BaseModel], IdentifierRoster]

The rows behind that roster, in table-then-file order, beside the roster itself.

One loader for both shapes. RM163's drift check needs the rows rather than the ids — an accession's training_ancestry and training_cohort sit on the same row as its pgs_id — and walking the same tables a second time would be a second place for the read / not-read bookkeeping to go wrong, which is precisely the bookkeeping S86 exists to keep honest.

Source code in enricher/src/just_dna_enricher/identifiers.py
def authored_rows(spec_dir: Path, column: str) -> tuple[list[BaseModel], IdentifierRoster]:
    """The rows behind that roster, in table-then-file order, beside the roster itself.

    One loader for both shapes. RM163's drift check needs the **rows** rather than the ids — an
    accession's `training_ancestry` and `training_cohort` sit on the same row as its `pgs_id` — and
    walking the same tables a second time would be a second place for the read / not-read bookkeeping
    to go wrong, which is precisely the bookkeeping S86 exists to keep honest.
    """
    roster = IdentifierRoster()
    kept: list[BaseModel] = []
    for name, model in sorted(_id_bearing_tables(column).items()):
        try:
            table = resolve_sidecar(spec_dir, name) or spec_dir / name
        except SidecarCollision as exc:
            roster.not_read[name] = str(exc)
            continue
        if not table.exists():
            roster.not_read[name] = NOT_PRESENT
            continue
        rows, errors, _ = load_csv_rows(table, model, table.name)
        if errors:
            # Read-only pass: an unparseable table is a *reason a question was not put*, never a
            # failure of this check. Dying here would say nothing about the ids the module does carry.
            roster.not_read[name] = f"could not be read ({errors[0]})"
            roster.read_errors[name] = errors[0]
            continue
        roster.read.append(name)
        kept.extend(rows)
        for row in rows:
            value = getattr(row, column, None)
            if not value:
                continue
            roster.ids.extend(part.strip() for part in MULTI_SEP.split(value) if part.strip())
    roster.ids = dedupe(roster.ids)
    return kept, roster

classify_pgs_accession

classify_pgs_accession(
    pgs_id: str, payload: dict
) -> PgsStatus

One accession plus whatever the Catalog served for it → known / unrecognised / malformed.

Pure and split out from the request, so the state machine is testable against a recorded payload the way classify_rsid is. The grammar arm comes first and is settled without a request: the Catalog answers 200 with {} for PGSXXXX exactly as it does for PGS999999, so its answer about a malformed id carries no information, and spending a request to learn nothing would also let the message read the absence as a possible withdrawal. Callers pass {} for an id they did not ask about.

Source code in enricher/src/just_dna_enricher/identifiers.py
def classify_pgs_accession(pgs_id: str, payload: dict) -> PgsStatus:
    """One accession plus whatever the Catalog served for it → `known` / `unrecognised` / `malformed`.

    Pure and split out from the request, so the state machine is testable against a recorded payload
    the way `classify_rsid` is. **The grammar arm comes first and is settled without a request:** the
    Catalog answers 200 with `{}` for `PGSXXXX` exactly as it does for `PGS999999`, so its answer
    about a malformed id carries no information, and spending a request to learn nothing would also
    let the message read the absence as a possible withdrawal. Callers pass `{}` for an id they did
    not ask about.
    """
    accession = (pgs_id or "").strip()
    if not PGS_ID_PATTERN.match(accession):
        return PgsStatus(pgs_id=pgs_id, state="malformed")
    if not payload:
        # The Catalog's own negative, read off the body because the status is 200 either way.
        return PgsStatus(pgs_id=accession, state="unrecognised")
    traits = tuple(
        str(t.get("id")) for t in payload.get("trait_efo") or [] if isinstance(t, dict) and t.get("id")
    )
    number = payload.get("variants_number")
    return PgsStatus(
        pgs_id=accession,
        state="known",
        name=payload.get("name") or None,
        date_release=payload.get("date_release") or None,
        trait_efo_ids=traits,
        variants_number=number if isinstance(number, int) else None,
        license=payload.get("license") or None,
    )

compare_pgs_metadata

compare_pgs_metadata(
    rows: list[BaseModel], records: dict[str, dict]
) -> PgsComparison

The two-field drift check over the authored rows, against the records the Catalog served.

Two fields, and the other two authored columns beside them are deliberately not checked. match_rate_floor is described in its own Field as an author-set floor and research_tier is a two-member curator judgement — the Catalog publishes neither, so there is nothing to drift them against, and a check with no source-side value is a check that cannot fail (@tautology-zero). Written down here so nobody adds them later on the symmetry.

A row whose accession the Catalog does not hold is out of scope entirely: pgs_accession_currency has already said so, and a second finding about the same absence would count one problem twice under two denominators.

Source code in enricher/src/just_dna_enricher/identifiers.py
def compare_pgs_metadata(rows: list[BaseModel], records: dict[str, dict]) -> PgsComparison:
    """The two-field drift check over the authored rows, against the records the Catalog served.

    **Two fields, and the other two authored columns beside them are deliberately not checked.**
    `match_rate_floor` is described in its own `Field` as an author-set floor and `research_tier` is a
    two-member curator judgement — the Catalog publishes neither, so there is nothing to drift them
    against, and a check with no source-side value is a check that cannot fail (`@tautology-zero`).
    Written down here so nobody adds them later on the symmetry.

    A row whose accession the Catalog does not hold is out of scope entirely: `pgs_accession_currency`
    has already said so, and a second finding about the same absence would count one problem twice
    under two denominators.
    """
    comparison = PgsComparison()
    seen: set[tuple[str, str, str]] = set()
    for row in rows:
        pgs_id = (getattr(row, "pgs_id", "") or "").strip()
        payload = records.get(pgs_id)
        if not payload:
            continue
        ancestry = getattr(row, "training_ancestry", None) or []
        cohort = (getattr(row, "training_cohort", "") or "").strip()
        # `(field, the authored value as written, the comparison)`. An absent cell is skipped rather
        # than withheld: no authored value is no claim to disagree with, and putting a cell nobody
        # wrote into the could-not-settle denominator would overstate what this check could not do.
        legs = []
        if ancestry:
            legs.append(
                ("training_ancestry", ",".join(ancestry), _compare_ancestry(pgs_id, ancestry, payload))
            )
        if cohort:
            legs.append(("training_cohort", cohort, _compare_cohort(pgs_id, cohort, payload)))
        for field_name, authored, (drift, reason) in legs:
            cell = (pgs_id, field_name, authored)
            if cell in seen:
                # The same accession stating the same value on two rows — a pleiotropic score reported
                # against two traits is two `PgsRow`s under one `pgs_id` — is one claim to this check,
                # and counting it twice would inflate both halves of the record. Two rows stating
                # *different* values are two claims and both are put: keying on the accession alone
                # dropped the second silently, and which one survived depended on row order.
                continue
            seen.add(cell)
            comparison.authored.append(cell)
            if reason is not None:
                comparison.withheld[cell] = reason
                continue
            comparison.compared.append(cell)
            if drift is not None:
                comparison.drift.append(drift)
    return comparison

check_identifiers

check_identifiers(
    variants: list[VariantRow] | None = None,
    *,
    spec_dir: Path | None = None,
    check_traits: bool = True,
    check_genes: bool = True,
    check_pgs: bool = True,
    client: OntologyClient | None = None,
    pgs_client: PgsCatalogClient | None = None,
    write: bool = True,
    declared_use: str = "unstated",
) -> IdentifierReport

Ontology-term, gene-symbol and PGS-accession currency for one module's authored identifiers.

rsIDs are deliberately not done here: they are checked inside enrich(), because their verdict lands on resolution.csv columns rather than being a standalone report. See check_rsids.

The PGS leg (RM163) needs a spec directory and is a no-op without one, for a structural reason rather than a policy: pgs_id lives on PgsRow and VariantRow has no such column, so a caller holding variants rows has no accession to put. It is also the one leg here that writes — sources.csv, because the Catalog's terms are per score and a module carrying an academic-use-only score must not compile claiming the generic ones. write=False turns that off for a caller that only wants the report.

Pass either variants or spec_dir (RM41) — the row-taking form is right for an in-process caller that already holds the rows, and spec_dir= is the shape every other pass in this tier has. Exactly one, never both: two answers in mind, and silently preferring one is a guess.

Source code in enricher/src/just_dna_enricher/identifiers.py
def check_identifiers(
    variants: list[VariantRow] | None = None,
    *,
    spec_dir: Path | None = None,
    check_traits: bool = True,
    check_genes: bool = True,
    check_pgs: bool = True,
    client: OntologyClient | None = None,
    pgs_client: PgsCatalogClient | None = None,
    write: bool = True,
    declared_use: str = "unstated",
) -> IdentifierReport:
    """Ontology-term, gene-symbol and PGS-accession currency for one module's authored identifiers.

    rsIDs are deliberately **not** done here: they are checked inside `enrich()`, because their verdict
    lands on `resolution.csv` columns rather than being a standalone report. See `check_rsids`.

    The PGS leg (RM163) needs a spec directory and is a no-op without one, for a structural reason
    rather than a policy: `pgs_id` lives on `PgsRow` and `VariantRow` has no such column, so a caller
    holding `variants` rows has no accession to put. It is also the one leg here that **writes** —
    `sources.csv`, because the Catalog's terms are per score and a module carrying an
    academic-use-only score must not compile claiming the generic ones. `write=False` turns that off
    for a caller that only wants the report.

    **Pass either `variants` or `spec_dir` (RM41)** — the row-taking form is right for an in-process
    caller that already holds the rows, and `spec_dir=` is the shape every other pass in this tier has.
    Exactly one, never both: two answers in mind, and silently preferring one is a guess.
    """
    if (variants is None) == (spec_dir is None):
        raise ValueError(
            "pass exactly one of variants= (rows you already hold) or spec_dir= (a module spec "
            "directory, loaded with the module's declared genome_build)"
        )
    if variants is None:
        # **An absent `variants.csv` is a module shape, not an error** — the format has never made the
        # table mandatory, and the four PGx kinds a module can be built entirely out of are among the
        # nine this pass now reads. Loading unconditionally raised `variants.csv is invalid: ... not
        # found` on every one of them, so the roster S86 widened could not be reached from a spec
        # directory at all: the CLI's own guard was returning before this line and hiding it. These
        # rows are wanted for exactly one thing — placing a gene symbol against the chromosome its
        # variant resolves to — and a module with none simply has nothing to place, which
        # `_gene_locus_conflicts` reports as its own reason rather than as a silent zero.
        #
        # Present-and-unparseable still raises. That is a module whose rows exist and cannot be read,
        # which is the author's to fix and not a shape.
        if (Path(spec_dir) / "variants.csv").exists():
            variants, errors, _ = load_spec_variants(Path(spec_dir))
            if errors:
                raise ValueError(f"variants.csv is invalid: {errors[0]}")
        else:
            variants = []
    report = IdentifierReport()
    # **The roster is every authored table carrying the column, not `variants.csv` alone (S86).** It
    # read one table while eleven declare `trait_efo_id`/`gene`, so a module whose traits live in
    # `studies.csv` — where `StudyRow` has carried the column since 0.3 — reported `0 checked, 0 clean,
    # 0 flagged` and shipped a retired CURIE with every gate green. `0` said two things at once: *this
    # module declares no trait* and *its traits are in a table nobody read*.
    #
    # A caller passing `variants=` still gets the narrow roster, because rows in hand are all it has;
    # `tables_not_read` then records that, so the narrow case is stated rather than silently equal to
    # the wide one.
    if spec_dir is not None:
        trait_roster = authored_identifiers(Path(spec_dir), "trait_efo_id")
        gene_roster = authored_identifiers(Path(spec_dir), "gene")
    else:
        trait_roster = IdentifierRoster(
            ids=module_trait_ids(variants),
            read=["variants.csv"],
            not_read={
                name: ROWS_PASSED_IN
                for name in sorted(_id_bearing_tables("trait_efo_id"))
                if name != "variants.csv"
            },
        )
        gene_roster = IdentifierRoster(
            ids=dedupe(v.gene for v in variants if v.gene),
            read=["variants.csv"],
            not_read={
                name: ROWS_PASSED_IN for name in sorted(_id_bearing_tables("gene")) if name != "variants.csv"
            },
        )
    traits = trait_roster.ids if check_traits else []
    genes = gene_roster.ids if check_genes else []
    # Recorded whether or not anything was found, and **before** the early return below: a roster that
    # came back empty is exactly the case where a reader needs to know which tables were behind it.
    if check_traits:
        report.trait_tables_read = trait_roster.read
        report.trait_tables_not_read = trait_roster.not_read
    if check_genes:
        report.gene_tables_read = gene_roster.read
        report.gene_tables_not_read = gene_roster.not_read
    for column, roster in (("trait_efo_id", trait_roster), ("gene", gene_roster)):
        for name, why in sorted(roster.unreadable.items()):
            # An absent table is the normal case and says nothing; one that exists and will not parse
            # means the check silently skipped ids the module really does carry.
            logger.warning(
                "%s carries %s and could not be read (%s), so its identifiers were not checked. "
                "The counts below are out of the tables that were.",
                name,
                column,
                why,
            )
    # **The ontology legs run first, and the early return became a guard rather than a return.** The
    # PGS leg reads a different roster from a different registry, so returning before it would leave a
    # module whose only identifiers are accessions — `pgs.csv` and nothing else — having asked nothing,
    # which is S86's unreadable zero arriving in a fourth column. And this order is the one that keeps
    # a registry outage local: `OntologyClient` raises, so an OLS4 failure must abort before the PGS
    # leg spends requests and writes a licence table the run is about to discard.
    if traits or genes:
        owned = client is None
        ontology = client or OntologyClient()
        try:
            report.traits = [ontology.trait(curie) for curie in traits]
            report.genes = [ontology.gene(symbol) for symbol in genes]
        finally:
            if owned:
                ontology.close()
        if check_genes:
            (
                report.gene_loci,
                report.gene_loci_compared,
                report.gene_loci_not_checked,
            ) = _gene_locus_conflicts(variants, report.genes, Path(spec_dir) if spec_dir else None)

    if check_pgs and spec_dir is not None:
        _check_pgs(report, Path(spec_dir), client=pgs_client, write=write, declared_use=declared_use)
    elif check_pgs:
        # **`unsupported`, never `nothing_to_check`.** With rows passed in there is no `pgs_id`-bearing
        # table to read at all — `VariantRow` has no such column — so saying "no row carries a pgs_id"
        # would assert a fact about the module that this call never established. Same distinction
        # `IdentifierRoster.not_read` draws one level down.
        report.pgs_not_checked = (
            "unsupported",
            "rows were passed in rather than a spec directory, and no row model a caller can hold "
            "carries a pgs_id — pass spec_dir= to check accessions",
        )
    return report

verification_records

verification_records(
    report: IdentifierReport,
    *,
    check_traits: bool,
    check_genes: bool,
    check_pgs: bool = True,
) -> list[VerificationRecord]

The five checks this pass puts, as records verification.json can carry (RM72, RM163).

Five records rather than one, because they are five questions over five different subject sets: the trait CURIEs the module names, the gene symbols it names, the rows where a symbol and a chromosome could both be established, the PGS accessions it keys on, and the authored cells beside those accessions. One record averaging them would be a number nothing could act on — the same argument literature._verification_records makes for its own three.

The two PGS records are two and not one, deliberately. pgs_accession_currency asks whether an id still names a score and pgs_metadata_agreement asks whether two cells beside it still match: different questions, different subjects, and therefore different denominators. Folding them would publish a single findings count over two populations, which is a number that means nothing.

Every count is read off report, and nothing is recounted here. The denominator belongs to the check that produced it; a count recomputed beside a check is one that can disagree with it, and then the document's two halves contradict each other.

A switched-off check is not_requested, an empty subject set is nothing_to_check, and a comparison whose input was missing is no_reference — never ran(0, 0). A zero out of zero reads as a clean bill, which is the whole confusion RM45 exists to end. Only the first of those three is cleared by re-running with a different flag, which is why the skip vocabulary keeps them apart rather than folding them into one absence.

This pass fetches nothing on the caller's behalf beyond OLS4 and HGNC, so there is no offline branch to record: the command has no such flag, and inventing one here would name a state the run cannot be in.

Source code in enricher/src/just_dna_enricher/identifiers.py
def verification_records(
    report: IdentifierReport,
    *,
    check_traits: bool,
    check_genes: bool,
    check_pgs: bool = True,
) -> list[VerificationRecord]:
    """The five checks this pass puts, as records `verification.json` can carry (RM72, RM163).

    Five records rather than one, because they are five questions over five different subject
    sets: the trait CURIEs the module names, the gene symbols it names, the rows where a symbol
    and a chromosome could both be established, the PGS accessions it keys on, and the authored cells
    beside those accessions. One record averaging them would be a number nothing could act on — the
    same argument `literature._verification_records` makes for its own three.

    **The two PGS records are two and not one, deliberately.** `pgs_accession_currency` asks whether
    an id still names a score and `pgs_metadata_agreement` asks whether two cells beside it still
    match: different questions, different subjects, and therefore different denominators. Folding
    them would publish a single findings count over two populations, which is a number that means
    nothing.

    **Every count is read off `report`, and nothing is recounted here.** The denominator belongs to
    the check that produced it; a count recomputed beside a check is one that can disagree with it,
    and then the document's two halves contradict each other.

    **A switched-off check is `not_requested`, an empty subject set is `nothing_to_check`, and a
    comparison whose input was missing is `no_reference` — never `ran(0, 0)`.** A zero out of zero
    reads as a clean bill, which is the whole confusion RM45 exists to end. Only the first of those
    three is cleared by re-running with a different flag, which is why the skip vocabulary keeps them
    apart rather than folding them into one absence.

    This pass fetches nothing on the caller's behalf beyond OLS4 and HGNC, so there is no `offline`
    branch to record: the command has no such flag, and inventing one here would name a state the run
    cannot be in.
    """
    records = [_trait_record(report, check_traits), _gene_symbol_record(report, check_genes)]
    records.append(_gene_locus_record(report, check_genes))
    records.append(_pgs_accession_record(report, check_pgs))
    records.append(_pgs_metadata_record(report, check_pgs))
    return records

unreachable_records

unreachable_records(
    *,
    check_traits: bool,
    check_genes: bool,
    check_pgs: bool = True,
    detail: str,
) -> list[VerificationRecord]

The same five checks, for a run whose registry never answered.

A request that failed is unreachable, never an absence — S20's distinction, and the reason the twin command records it too. Without this the command would advertise records that the question was put and then write nothing on precisely the run where a reader most needs to know the question went unanswered.

Derived from verification_records over an empty report rather than restated: the flag a caller set is still the truer reason for a check they switched off (a source being down does not make --no-traits a network problem), and the source names come from the one place that assigns them.

Source code in enricher/src/just_dna_enricher/identifiers.py
def unreachable_records(
    *, check_traits: bool, check_genes: bool, check_pgs: bool = True, detail: str
) -> list[VerificationRecord]:
    """The same five checks, for a run whose registry never answered.

    A request that failed is `unreachable`, never an absence — S20's distinction, and the reason the
    twin command records it too. Without this the command would advertise *records that the question
    was put* and then write nothing on precisely the run where a reader most needs to know the
    question went unanswered.

    Derived from `verification_records` over an empty report rather than restated: the flag a caller
    set is still the truer reason for a check they switched off (a source being down does not make
    `--no-traits` a network problem), and the source names come from the one place that assigns them.
    """
    return [
        record
        if record.skipped == "not_requested"
        else skipped(record.check, "unreachable", detail=detail, source=record.source)
        for record in verification_records(
            IdentifierReport(),
            check_traits=check_traits,
            check_genes=check_genes,
            check_pgs=check_pgs,
        )
    ]

pgs_withheld_sentences

pgs_withheld_sentences(
    comparison: PgsComparison,
) -> list[str]

One sentence per reason for the cells that could not be settled — never one per cell.

@dont-discard-computed: these were counted by the comparison and would otherwise vanish, and a reader who cannot see them reads 2 of 2 agree as a clean bill over six cells.

Source code in enricher/src/just_dna_enricher/identifiers.py
def pgs_withheld_sentences(comparison: PgsComparison) -> list[str]:
    """One sentence per *reason* for the cells that could not be settled — never one per cell.

    `@dont-discard-computed`: these were counted by the comparison and would otherwise vanish, and a
    reader who cannot see them reads `2 of 2 agree` as a clean bill over six cells.
    """
    by_reason: dict[str, list[str]] = {}
    for (pgs_id, field_name, _authored), reason in comparison.withheld.items():
        by_reason.setdefault(reason, []).append(f"{pgs_id}/{field_name}")
    return [
        f"{len(by_reason[reason])} cell(s) withheld because {reason}: " + examples(sorted(by_reason[reason]))
        for reason in sorted(by_reason)
    ]