just_dna_enricher.identifiers¶
just_dna_enricher.identifiers ¶
Identifier currency — are the labels a module keys on still the ones its registries serve?
Four registries, one shape: ask whether an authored identifier is current, and report what the
registry says instead. Nothing here repairs anything, and for rsIDs that is not a stylistic
preference but a hard requirement — see check_rsids.
The fourth is the PGS Catalog (RM163), and it arrived last for a reason worth keeping: pgs_id is the
column PgsRow is keyed on, so it was the one authored identifier in the format that no pass ever
put to its registry. Its client and its source-shape knowledge live in pgs.py; the statuses, the
comparison and the two attestations are here, beside the other three.
This is the generalization of COMPILER.md's "is the source stale?" blind spot from datasets to identifiers. The compiler cannot see the world move; a dbSNP merge, an EFO term retirement and an HGNC symbol rename all leave a module perfectly well-formed and quietly out of date.
dbSNP: three states, and one of them means two opposite things. Probed rather than assumed:
| State | esummary db=snp |
|---|---|
| live | snp_id == requested, merged_sort='0' |
| merged | snp_id != requested, merged_sort='1' |
| absent | {'uid': …, 'error': 'cannot get document summary'} |
The absent response for rs11273140 — an rsID withdrawn after a clustering error — is
byte-identical to the one for rs2000000000, which was never assigned. For an author these mean
opposite things (fix the typo, versus the variant itself was retracted and the annotation resting on
it may be worthless), and no live endpoint separates them: esearch has no withdrawn filter, and
latest_release/misc/rs_unsupported_b157.txt looks like a withdrawn registry but is a one-off
build-157 ClinVar-parsing incident list that does not contain rs11273140. So the message names both
readings and asserts neither. Guessing "typo" would send an author to fix the wrong thing.
NCBI is the oracle, not Ensembl. Ensembl REST resolves some merges (rs77121243 → rs334) and
returns HTTP 400 on others (rs3216883, which dbSNP correctly reports as merged into rs3051860), so
Ensembl alone would misclassify a merged rsID as unresolvable.
IdentifierCheckError ¶
Bases: RuntimeError
A currency lookup failed outright (transport or an unusable response).
IdentifierUnavailable ¶
Bases: IdentifierCheckError
A source this pass asks was not reachable, so the question could not be put (RM101).
A subclass rather than a second exception, for AcmgListUnavailable's reason: every existing
except IdentifierCheckError still catches it, and the two histories this module can fail with
stop being one type. Only this one means a source was asked; an invalid local variants.csv is
the plain parent, and a caller separating them had to read __cause__ to do it.
It carries no skip member, unlike AcmgListUnavailable, and the reason is not that nothing
attests — check-identifiers does, and getting this wrong once is how the first draft of this
class shipped a false docstring. The difference is where the member is decided. acmg raises
from several places that mean different skips, so the raiser is the only thing that knows which;
here every raise means the same one, and unreachable_records() already owns it for all three
checks at once. Putting it on the exception would be a second place to keep the same fact.
RsidStatus
dataclass
¶
What dbSNP currently says about one authored rsID.
is_fatal
property
¶
Whether this refuses in both modes rather than only under strict.
Only withdrawn does. A merged rsID still resolves to the right locus (the module is dated,
not wrong) and an absent one may be a typo, a very new id, or API lag — all leave the
annotation itself intact. A withdrawn one is dbSNP repudiating the variant, so the annotation
may be describing nothing.
Nothing in this module produces withdrawn today: a retraction is byte-identical to a
never-assigned id through every live endpoint, so classify_rsid reports absent and names
both readings. The state exists for a curator who has established the retraction by hand, and
for a future source that can tell them apart.
TraitStatus
dataclass
¶
What the ontology currently says about one authored trait id.
GeneStatus
dataclass
¶
GeneStatus(
symbol: str,
state: str,
current: str | None = None,
hgnc_id: str | None = None,
location: str | None = None,
)
What HGNC currently says about one authored gene symbol.
chromosome
property
¶
The contig the band names, or None when HGNC gave none or it cannot be read.
Band syntax is <chrom><arm><band>, so the contig is everything before the first p/q —
except mitochondria, which HGNC spells out and this codebase calls MT. Alternate-reference
loci (19q13.42 alternate reference locus) keep their leading band, so the split still works.
None for anything unparsed, because a guess here becomes a false accusation about a row.
GeneLocusConflict
dataclass
¶
An authored gene whose HGNC chromosome is not the chromosome the row's variant sits on.
The relationship is false while both halves are individually true — the rsID resolves, the symbol is approved — which is why nothing caught it before (S24). It is the signature of a generated citation: a real gene name beside an invented rs number, which resolves anyway because dbSNP is dense enough that almost any seven-digit number hits something. Four of one reporter's seven rows were this, each pairing a real symbol with a variant on another chromosome.
PgsStatus
dataclass
¶
PgsStatus(
pgs_id: str,
state: str,
name: str | None = None,
date_release: str | None = None,
trait_efo_ids: tuple[str, ...] = (),
variants_number: int | None = None,
license: str | None = None,
)
What the PGS Catalog currently says about one authored pgs_id (RM163).
Three states, and the first of them is settled before any request goes out. The REST surface answers HTTP 200 with an empty body for a never-assigned accession and for a malformed one alike, so the service itself cannot separate them — but the format's own grammar can, and asking the Catalog about a string that is not an accession would spend a request to learn nothing.
PgsDrift
dataclass
¶
An authored cell beside a pgs_id that the score's own Catalog record contradicts (RM163).
Reported, never repaired (@enrichment-is-validation): which of the two is right is not something
this tier can know, and both training_ancestry and training_cohort are authored judgements a
curator may have made deliberately against the Catalog's own summary.
PgsComparison
dataclass
¶
PgsComparison(
drift: list[PgsDrift] = list(),
compared: list[tuple[str, str, str]] = list(),
withheld: dict[tuple[str, str, str], str] = dict(),
authored: list[tuple[str, str, str]] = list(),
)
What the two-field drift check compared, what it withheld, and why — the denominator (RM163).
withheld is the point of the type, and it is the @dont-discard-computed rule: a cell this
check looked at and could not settle is neither a finding nor a clean comparison, and folding it
into either would publish a coverage figure larger or smaller than the question that was put.
IdentifierReport
dataclass
¶
IdentifierReport(
rsids: list[RsidStatus] = list(),
traits: list[TraitStatus] = list(),
genes: list[GeneStatus] = list(),
gene_loci: list[GeneLocusConflict] = list(),
gene_loci_not_checked: str | None = None,
gene_loci_compared: int = 0,
trait_tables_read: list[str] = list(),
trait_tables_not_read: dict[str, str] = dict(),
gene_tables_read: list[str] = list(),
gene_tables_not_read: dict[str, str] = dict(),
pgs: list[PgsStatus] = list(),
pgs_metadata: PgsComparison = PgsComparison(),
pgs_tables_read: list[str] = list(),
pgs_tables_not_read: dict[str, str] = dict(),
pgs_release: str | None = None,
pgs_not_checked: tuple[str, str] | None = None,
pgs_sources: list[SourceRow] = list(),
)
unreadable_tables
property
¶
Tables that exist and would not parse, across both columns — never merely absent ones.
The actionable half: an absent optional table is the normal shape of every module, while one that is present and unreadable means ids the module carries went unchecked.
clean
property
¶
Whether every identifier this pass put to a registry is one the registry still serves.
pgs_metadata.drift is deliberately NOT in here, and stale_pgs deliberately is. The CLI
gates --strict's exit code on this property, and the two findings are different kinds: an
accession the Catalog does not hold is a broken identifier, exactly like a retired HGNC symbol;
a training_ancestry the Catalog's own summary disagrees with is two authorities differing
about a curated judgement, and PgsDrift.__str__ tells the author in as many words to leave it
if their curation is deliberate. Failing a build on that would make the format arbitrate
between its own sources with no way for the author to clear it — the reason
@clinsig-never-escalates and @a-source-recuring-is-not-a-strict-matter keep the same shape
out of the strict gate. It is reported in both modes; see metadata_disagrees.
A Verdict rather than a bool since 2026-09-13 (RM235), for the reason verdict.py
states: --strict gates an exit code on this, so a run that could not certify has to be a
no, and the reason for the no rides beside the verdict instead of inside it.
tables_unreadable is the new arm and the one behaviour change. A table that carries
identifiers and will not parse leaves those ids unchecked, and the five lists above are empty
for exactly the same reason they are empty when everything agreed — so clean answered True
over ids nobody looked at. The command has printed the unreadable table since S86 and then
exited 0 under --strict with all identifiers current above it: the diagnosis was there and
the verdict disagreed with it. A library caller reading this property sees only the verdict,
which is how the sibling defect on AcmgReport arrived (S100 — these are wrapped as MCP tools,
so the dataclass is what is read and never the terminal).
What is deliberately not an arm. A check the caller switched off, and a module carrying no
id-bearing table at all, are honest non-answers rather than errors: they keep the verdict true
and are reported through the read/not-read rosters that already exist, so --strict --no-traits
does not start failing builds for doing as it was told. unreachable is not an arm either, and
could not be — a registry outage raises IdentifierUnavailable before this report exists, and
the CLI already exits 1 there with an unreachable attestation for all five checks. The item
that proposed widening this property for outages had measured hand-built models rather than the
real path (@a-disagreement-with-a-document-may-be-in-the-instrument).
metadata_disagrees
property
¶
Whether a source disagreed with an authored cell — reported, never gated.
Separate from clean so a caller can say checked, and something differs without the exit
code claiming the module is broken. The CLI uses it to withhold the green line.
OntologyClient
dataclass
¶
OntologyClient(
ols4_base: str = DEFAULT_OLS4_BASE,
hgnc_base: str = DEFAULT_HGNC_BASE,
min_request_interval: float = 0.2,
timeout: float = 30.0,
gate: PacingGate | None = None,
_client: Client | None = None,
)
OLS4 + HGNC lookups. One class because both are unbatched GET-per-id REST services.
Neither publishes a documented rate limit, so the gate is a courtesy rather than a budget — but it is here rather than absent, because a module with two hundred genes would otherwise open two hundred connections as fast as it could.
trait ¶
Existence, obsolescence and the replacement term for one ontology CURIE.
Source code in enricher/src/just_dna_enricher/identifiers.py
gene ¶
Approved / retired / unknown for one gene symbol.
Uses the exact fetch/ endpoints, never search/: search/BRCA1 is fuzzy and returns 19
hits including ABRAXAS1, which would make every symbol look ambiguous. fetch/symbol answers
currency; fetch/prev_symbol answers retirement, and is only consulted when the first misses.
Source code in enricher/src/just_dna_enricher/identifiers.py
IdentifierRoster
dataclass
¶
IdentifierRoster(
ids: list[str] = list(),
read: list[str] = list(),
not_read: dict[str, str] = dict(),
read_errors: dict[str, str] = dict(),
)
The ids an authored spec names, with the tables that were and were not read (S86).
not_read is the point of the type. A bare roster cannot distinguish this module declares no
trait from this module's traits are in a table nobody read, and both render as 0 checked,
0 clean, 0 flagged — which reads as a clean run and let a module ship a retired CURIE with every
gate green. Same three-valued rule as gene_loci_not_checked two fields down in the report, and
as unconsulted_rsids one tier down: a question never put is not an answer.
unreadable
property
¶
Only the tables that exist and failed to parse — the half worth warning about.
classify_rsid ¶
One esummary record → live / merged / absent.
Split out from the batch so the state machine is testable against the recorded payload directly.
snp_id comes back as an int and the uid as a string, so both sides are normalized before
comparison — comparing them raw would classify every live rsID as merged.
Source code in enricher/src/just_dna_enricher/identifiers.py
check_rsids ¶
Current/merged/absent for each authored rsID, batched through esummary db=snp.
Report, never repair — and here that is load-bearing rather than tidy. weights.parquet
carries both variant_key and rsid, and for an rsid-authored row they are the same label.
Writing the merged-into id back would not be a one-time digest move but an identity migration
performed by a network lookup: reverse would emit the new rsID into variants.csv, the next
compile would key on it, and variant_key itself would change with no authored edit anywhere.
The module's identity would drift and the round-trip would stop being a fixed point (Principle 7).
Source code in enricher/src/just_dna_enricher/identifiers.py
module_trait_ids ¶
Every ontology CURIE the module's variants.csv rows name, in first-occurrence order.
trait_efo_id is a multi-valued cell (comma/semicolon/pipe-separated), so one row can name
several traits.
This is the roster of one table, and eleven authored tables carry the column. Kept as-is
because a caller holding rows is the only thing it can serve, and widened where the tables are
actually reachable: authored_identifiers below takes a spec directory and reads all of them.
A caller passing variants= gets this narrow roster and is told so (S86).
Source code in enricher/src/just_dna_enricher/identifiers.py
authored_identifiers ¶
Every value of column across every authored table that declares it, and what was read.
The roster check_identifiers runs on when it is given a spec directory. Rows are loaded with the
same load_csv_rows the compiler uses, so a table this cannot parse is reported as unread rather
than skipped silently — the distinction the whole item is about.
Multi-valued cells are split on MULTI_SEP for both columns. A gene cell is single-valued in
practice, and splitting it costs nothing and cannot mis-read one: a symbol contains no separator.
Source code in enricher/src/just_dna_enricher/identifiers.py
authored_rows ¶
The rows behind that roster, in table-then-file order, beside the roster itself.
One loader for both shapes. RM163's drift check needs the rows rather than the ids — an
accession's training_ancestry and training_cohort sit on the same row as its pgs_id — and
walking the same tables a second time would be a second place for the read / not-read bookkeeping
to go wrong, which is precisely the bookkeeping S86 exists to keep honest.
Source code in enricher/src/just_dna_enricher/identifiers.py
classify_pgs_accession ¶
One accession plus whatever the Catalog served for it → known / unrecognised / malformed.
Pure and split out from the request, so the state machine is testable against a recorded payload
the way classify_rsid is. The grammar arm comes first and is settled without a request: the
Catalog answers 200 with {} for PGSXXXX exactly as it does for PGS999999, so its answer
about a malformed id carries no information, and spending a request to learn nothing would also
let the message read the absence as a possible withdrawal. Callers pass {} for an id they did
not ask about.
Source code in enricher/src/just_dna_enricher/identifiers.py
compare_pgs_metadata ¶
The two-field drift check over the authored rows, against the records the Catalog served.
Two fields, and the other two authored columns beside them are deliberately not checked.
match_rate_floor is described in its own Field as an author-set floor and research_tier is a
two-member curator judgement — the Catalog publishes neither, so there is nothing to drift them
against, and a check with no source-side value is a check that cannot fail (@tautology-zero).
Written down here so nobody adds them later on the symmetry.
A row whose accession the Catalog does not hold is out of scope entirely: pgs_accession_currency
has already said so, and a second finding about the same absence would count one problem twice
under two denominators.
Source code in enricher/src/just_dna_enricher/identifiers.py
check_identifiers ¶
check_identifiers(
variants: list[VariantRow] | None = None,
*,
spec_dir: Path | None = None,
check_traits: bool = True,
check_genes: bool = True,
check_pgs: bool = True,
client: OntologyClient | None = None,
pgs_client: PgsCatalogClient | None = None,
write: bool = True,
declared_use: str = "unstated",
) -> IdentifierReport
Ontology-term, gene-symbol and PGS-accession currency for one module's authored identifiers.
rsIDs are deliberately not done here: they are checked inside enrich(), because their verdict
lands on resolution.csv columns rather than being a standalone report. See check_rsids.
The PGS leg (RM163) needs a spec directory and is a no-op without one, for a structural reason
rather than a policy: pgs_id lives on PgsRow and VariantRow has no such column, so a caller
holding variants rows has no accession to put. It is also the one leg here that writes —
sources.csv, because the Catalog's terms are per score and a module carrying an
academic-use-only score must not compile claiming the generic ones. write=False turns that off
for a caller that only wants the report.
Pass either variants or spec_dir (RM41) — the row-taking form is right for an in-process
caller that already holds the rows, and spec_dir= is the shape every other pass in this tier has.
Exactly one, never both: two answers in mind, and silently preferring one is a guess.
Source code in enricher/src/just_dna_enricher/identifiers.py
1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 | |
verification_records ¶
verification_records(
report: IdentifierReport,
*,
check_traits: bool,
check_genes: bool,
check_pgs: bool = True,
) -> list[VerificationRecord]
The five checks this pass puts, as records verification.json can carry (RM72, RM163).
Five records rather than one, because they are five questions over five different subject
sets: the trait CURIEs the module names, the gene symbols it names, the rows where a symbol
and a chromosome could both be established, the PGS accessions it keys on, and the authored cells
beside those accessions. One record averaging them would be a number nothing could act on — the
same argument literature._verification_records makes for its own three.
The two PGS records are two and not one, deliberately. pgs_accession_currency asks whether
an id still names a score and pgs_metadata_agreement asks whether two cells beside it still
match: different questions, different subjects, and therefore different denominators. Folding
them would publish a single findings count over two populations, which is a number that means
nothing.
Every count is read off report, and nothing is recounted here. The denominator belongs to
the check that produced it; a count recomputed beside a check is one that can disagree with it,
and then the document's two halves contradict each other.
A switched-off check is not_requested, an empty subject set is nothing_to_check, and a
comparison whose input was missing is no_reference — never ran(0, 0). A zero out of zero
reads as a clean bill, which is the whole confusion RM45 exists to end. Only the first of those
three is cleared by re-running with a different flag, which is why the skip vocabulary keeps them
apart rather than folding them into one absence.
This pass fetches nothing on the caller's behalf beyond OLS4 and HGNC, so there is no offline
branch to record: the command has no such flag, and inventing one here would name a state the run
cannot be in.
Source code in enricher/src/just_dna_enricher/identifiers.py
unreachable_records ¶
unreachable_records(
*,
check_traits: bool,
check_genes: bool,
check_pgs: bool = True,
detail: str,
) -> list[VerificationRecord]
The same five checks, for a run whose registry never answered.
A request that failed is unreachable, never an absence — S20's distinction, and the reason the
twin command records it too. Without this the command would advertise records that the question
was put and then write nothing on precisely the run where a reader most needs to know the
question went unanswered.
Derived from verification_records over an empty report rather than restated: the flag a caller
set is still the truer reason for a check they switched off (a source being down does not make
--no-traits a network problem), and the source names come from the one place that assigns them.
Source code in enricher/src/just_dna_enricher/identifiers.py
pgs_withheld_sentences ¶
One sentence per reason for the cells that could not be settled — never one per cell.
@dont-discard-computed: these were counted by the comparison and would otherwise vanish, and a
reader who cannot see them reads 2 of 2 agree as a clean bill over six cells.