just-dna-enricher — the network tier¶
The package reference for just-dna-enricher: the only tier in the workspace allowed to fetch. It
produces the source-independent resolution table (resolution.csv) that the compiler consumes, and
carries the publisher surface for pushing compiled modules to HuggingFace. New in 0.5.
Enrichment is partly validation, and that is a goal rather than a side effect. Filling in what a
module left out is only half of what this tier does; the other half is checking what the module
asserted against what the sources actually say. It is the only tier that can: the format and compiler
tiers are inject-only by charter (Principle 2) and hold no reference to check anything against. So
surfacing a discrepancy — an authored ref that contradicts the genome, a source id that disagrees with
a locally-minted one, an rsID that maps somewhere else — is part of the job. Two rules govern every such
check: it reports and never repairs (silently rewriting an authored value destroys the evidence that
something upstream is wrong, and turns a loud data problem into a quiet one), and its severity follows
the mode (best_effort warns and carries on; strict refuses). What exists today:
| Check | Compares | Where | Attests as |
|---|---|---|---|
| Reference allele | authored ref vs the actual reference sequence |
sequences.verify_reference_alleles |
reference_allele |
| Wrong build | a ref-mismatched row vs the same coordinate on GRCh37 (0.6) | grch37.diagnose_wrong_build (warns in both modes) |
genome_build_agreement |
| VRS cross-check | a source's own vrs_id vs the locally-minted one |
vrs.mint_resolution_rows |
vrs_allele_id |
| rsid↔coordinate | an authored pair vs what the reference says | compiler/resolution.py::_verify over the injected table (warning), and enrich() against the injected Ensembl snapshot (resolver.check_rsid_coordinates, warning in both modes) — one question, two tiers, so one attestation name, and the enricher's half is the one that attests |
rsid_coordinate_agreement |
| Ambiguous back-fill | ≥2 rsIDs for one exact allele → recorded, never guessed | resolver._lookup_rsid_candidates |
— (recorded onto the row, not attested) |
| Clinical significance | authored clin_sig vs every annotation authority consulted — ClinVar and, since 0.7, PubMind — allele-exactly |
clinical.verify_clin_sig (warns in both modes), persisted as the N-authority concordance record by clinical.clin_sig_concordance (0.7, RM130 + RM134 § B) |
clinical_significance |
| Answered-call currency | an authority's call now vs what it said when the author's overrides.csv answer was written (0.7, RM151) |
clinical.answered_call_shift → concordance.shifted_authority_calls (warns in both modes; the baseline is the previous run's clin_sig_authority_calls.csv, so it is read before the commit rewrites it and a move is observable exactly once) |
— (logged; the record it reads is the attestation) |
| PGx evidence level | authored evidence_level vs ClinPGx's own for that annotation |
clinpgx.enrich_clinpgx (refuses in strict — the only enricher cross-check that does — on a row's own cited annotation only; RM297) |
pgx_evidence_level |
| Citation existence | a cited pmid vs PubMed |
literature.enrich_literature |
citation_existence |
| Identifier agreement | an authored doi vs the registry's for that PMID |
literature.enrich_literature |
citation_identifier |
| PMC id agreement | an authored PMC… in the pmid cell vs PubMed's for that record (0.6) |
literature._pmcid_conflicts (attested under citation_identifier, the same question one registry over) |
citation_identifier |
| Article licence | the cited article's own terms, recorded per article (0.6) | literature.enrich_literature → licensing.article_terms |
— (a recording pass: it writes a source's answer and compares nothing) |
| Provenance quote | provenance_quote/provenance_regex vs open-access fulltext |
literature.enrich_literature (warning; partial coverage) |
provenance_quote |
| rsID currency | an authored rsID vs dbSNP (live / merged / absent) | identifiers.check_rsids |
rsid_currency |
| Trait currency | trait_efo_id vs OLS4 (obsolete + replacement) |
identifiers.OntologyClient.trait (attested by check-identifiers since RM72) |
trait_currency |
| Gene symbol currency | gene vs HGNC approved / previous symbols |
identifiers.OntologyClient.gene (attested by check-identifiers since RM72) |
gene_symbol_currency |
| Gene ↔ locus agreement | the row's gene vs the chromosome its variant sits on (0.5.4) |
identifiers.check_identifiers → GeneLocusConflict (attested since RM72) |
gene_locus_agreement |
| PGS accession currency | an authored pgs_id vs the PGS Catalog's own record for it (0.7, RM163) — the Catalog answers 200 with {} for a never-assigned id and for a malformed one, so the check reads the body |
identifiers._check_pgs (attested by check-identifiers; --no-pgs records not_requested) |
pgs_accession_currency |
| PGS metadata agreement | authored training_ancestry / training_cohort vs the score record's ancestry_distribution / samples_training (0.7, RM163) — its own member, because currency asks whether the id still names a score and this asks whether two cells beside it still match |
identifiers._check_pgs (attested by check-identifiers) |
pgs_metadata_agreement |
| ACMG secondary findings | authored acmg_sf vs the published SF gene list (v3.3 via --sf-list; the scraped v3.2 page reports unverifiable) |
acmg.check_acmg_sf (attested by check-acmg since RM72) |
acmg_secondary_findings |
| Repeat bands | an authored repeat_alleles.csv band table vs STRchive's benign_*/intermediate_*/pathogenic_* (0.7, RM165) |
strchive.check_repeat_bands (warns in both modes; the catalogue's pathogenic_max is reported and never written) |
repeat_band_agreement |
| Regulator drug labels | a (gene[, allele], drug) claim vs the Testing Level five drug regulators' labels carry, at two join tiers (0.7, RM166) |
drug_labels.check_drug_labels (warns in both modes; a blank level is unknown and never No Clinical PGx) |
regulator_label_agreement |
| Published refutation | an authored direction vs the refutations CIViC publishes for the same variant (0.7, RM170) |
civic_refutation.compare_refutations → enrich() (warns in both modes; the record names its status_basis, because on the accepted basis the class is empty by construction) |
published_refutation |
| Literature coverage | which papers a variant-literature index holds for a module's alleles, and at which tier (0.7, RM167) | litvar.check_literature_coverage (reports only; allele-resolved, position-only and absent are three outcomes) |
literature_coverage |
| Evidence-status currency | a recorded StudyRow.confidence — a curation status a citing source published when the row was drafted — vs what that source says about the same item now (0.7, RM160) |
civic_citations.check_evidence_status_currency → enrich() (warns in both modes; a status accepted or rejected since, or a citation added since. Not dataset_currency: that one asks which release, this one asks whether one judgement moved) |
evidence_status_currency |
| Allele function | authored function_status vs PharmVar and CPIC |
pgx.enrich_pgx (warns in both modes) |
allele_function |
| Declared use | the caller's --use vs a source's terms |
licensing.check_declared_use (refuses in both modes) |
— (an acquisition gate, not a comparison) |
| Drafted vs authored rows | a source's current row vs the one already in the CSV | just_dna_compiler.draft.append_rows (reports differs; never rewrites) |
— (a drafting report; it runs outside enrich()) |
| A drafted table's licence row | does sources.csv record what licensed the rows this append is landing? |
drafting.licence_commit, handed to every append_* as before_commit (RM232) |
refuses the run — the table is not committed, so a module never carries drafted rows with no licence record |
| Source coverage | is the locus inside the source's callset at all? not_covered ≠ not_found |
gnomad.covers_locus → frequencies.enrich_frequencies (not a strict failure) |
— (a recording pass: coverage is a fact about the callset) |
| Dataset currency | a recorded SourceRow.dataset vs the release that source publishes now (0.7) |
currency.check_dataset_currency → enrich() (reports; strict refuses over a superseded release, never over an unreachable source) |
dataset_currency |
| Variant impact | an authored variant against the score AlphaGenome published for it — the three questions the local AVI artifact provably cannot answer (0.7, RM193) | alphagenome_check.check_variant_impact (warns in both modes; non-commercial, so an undeclared run asks nothing. Its own member rather than a second writer of reference_allele, which the Atlas also answers — two registries answering an overlapping question get two names, or one source's outage writes a skip against the other's) |
variant_impact_agreement |
The fourth column is the join, and it is asserted rather than promised. vocab.py's own comment
says VALID_VERIFICATION_CHECKS was audited against this table — so a member with an emitter and no
row here is a check a reader of verification.json cannot look up, which is what happened to
variant_impact_agreement for a release. Every member that has an emitter names a row above
(test_enricher_doc_registries.py walks it), and the two the vocabulary marks RESERVED —
gene_disease_validity and dosage_sensitivity — deliberately have none: their passes record a
source's verdict and compare nothing authored, and the member is held for a future pass that does. The
containment runs one way only: rows whose fourth column is — are real work that attests nothing,
either because it is an acquisition gate, a recording pass, or a report that runs outside enrich().
Every check that runs records what it did — verification.json (RM45, 0.6). Until 0.6 the table
above described work whose result died with the process: a check's findings reached a log line and an
EnrichmentResult field, and the compiled module could not say whether the check had been put at all.
just_dna_enricher.verification.record_verification is the load-merge-write that closes that, in the
same shape licensing.record_source_terms has and for the same reason — a count of call sites goes
stale, one function does not. Four things to hold onto when wiring a new pass into it:
- A record carries two counts and a closed skip key.
ran(check, subjects=…, findings=…)when it ran (subjects=0is a legitimate answer meaning nothing was in scope) andskipped(check, reason, detail=…)when it did not. Both vocabularies are closed (VALID_VERIFICATION_CHECKS,VALID_VERIFICATION_SKIPS); the human sentence rides indetail, beside the machine key and never instead of it. - The denominator comes from the check, never from the caller.
verify_reference_allelesreturns aRefCheckandcompare_clin_sigaClinSigComparison(orNone), so the count travels with the finding it belongs to: a count recomputed beside a check can disagree with it, and then the manifest's own two halves disagree._verification_recordsdeliberately takes neithervariantsnorrows— a function that cannot see the tables cannot be tempted to count them. Wiring this up surfaced a real hole in two passes: each had an internal skip (no sequence access; a snapshot present but not queryable) returning an empty list indistinguishable from a clean pass, which is S4's defect surviving inside the machinery S4 built. - The denominator is what was EXAMINED, not what existed. The wrong-build pass is bounded
(
DEFAULT_DIAGNOSIS_LIMIT), so on a panel authored wholesale on hg19 it asks about a sample; recordingtotalthere would claim rows it chose not to ask about.sampledis why the two can differ, and the record'sdetailsays so when they do. - A downstream check inherits the reason its upstream did not run. The wrong-build pass reads the
ref-mismatch list, and
diagnose_wrong_build([])answersno_ref_mismatchesfor an empty list whatever emptied it — a ref check that ran clean, or one that never ran. Recording the first unconditionally publishes "no authored ref disagreed with the reference" beside areference_allelerecord saying nothing was compared: one document contradicting itself, and the false half is the answered-absence-versus-unasked-question collapse S20 exists to prevent. - Every early return records its skip — as long as the check APPLIES.
enrich_clinpgxis the worked example: a licensing refusal and a missing snapshot both go through one_attesthelper, because each says "this check applies to your module and did not run", and a pass that records its findings while staying silent about not having run leaves the manifest unable to tell those apart. Two paths deliberately attest nothing. Astrictrefusal raises, so no artifact was produced and there is nothing to attest a check against. And a module carrying nopharm_variants.csvis not a skip at all — the check does not apply, there is no claim to have an opinion about, and recording one would mine a nonce and create averification.jsonon a module that never asked for one.nothing_to_checkstays for a table that is present with no row in scope, which is a real answer. RM72's two commands follow that line exactly: novariants.csvattests nothing (acmg_sf,geneandtrait_efo_idare all columns of it, so the check does not apply), while a failed lookup attestsunreachable— the run with no report to print is the one where a reader most needs the record.check-acmgdistinguishes the two ways it can raise for that purpose:AcmgListUnavailablecarries the skip member (unreachablefor a request that never answered,no_referencefor a source that was there and carried no readable list), while astrictrefusal stays the plain exception and attests nothing — the list was read and the question was answered, so recording a skip would say it was never put. When an attestation fails on the way out of an already-failing run it is reported, not raised: the reason the source could not be read is the sentence the author needs, and a layout complaint replacing it points at the wrong problem. - Two authorities answering one check still make one record, and the counts have to follow.
allele_functionis compared against PharmVar and CPIC, so three things differ from the single-source passes.subjectsis the alleles an authority named back, never the authored rows: a curator may state a function for an allele a point-in-time slice does not list, and counting that row would publish a comparison nobody put (the shortfall is named indetailinstead).findingscounts alleles in dispute, not conflicts — one authored status contradicted by both panels is one allele reported twice, andVerificationRecordrightly refusesfindings > subjects. Andsourceis filled only when exactly one authority is implicated, because it is a single join key into the licensing table: naming one of two would hide the other, and a comma-joined value would break the join the column exists for. With no leg answering, the reason is picked by precedence (not_permitted→offline→unreachable→not_requested) and the sentence names every absent leg — a licensing refusal is cleared by a declaration, so reporting it asofflinewould send a reader hunting a network problem that does not exist. - A check whose input has no producer records the skip, not the pass's own numbers.
vrs mintends with a number out of a number, and that number is coverage, not a comparison:vrs_allele_idnames the cross-check of a source's ownga4gh:VA.…against the locally minted one, whose input ismint_resolution_rows(source_ids=…)— andresolution.csvrecords ids without recording where an id came from, so no caller in the workspace fills that map. The command therefore recordsnothing_to_checkwith the coverage indetail, grouped by reason class. Recording the alleles it named assubjectswould assert a comparison that was never made, which is exactly the trap the reference-allele pass fell into on an unbuilt assembly. (The coverage numbers are published already, by the compiler, asmanifest.compilation.vrs_alleles/vrs_alleles_identified— this record is not their second home.) - A check with no lookup of its own still needs a reason ladder. The rsid↔coordinate pass costs no
request: it widens the rsID batch the resolver chain was already sending and compares the authored
pairs against what came back (
resolver.check_rsid_coordinates, one direction — the converse needs a position→id lookup atchrom:start:refgranularity, and the reverse mapenrich()holds is allele-exact, so asking it would report a spelling difference as a contradiction). The verdict is three-valued and compared atchrom:start:refis optional on the PGx models —pgx_draftwrites exactly that shape — and arefthat disagrees with the reference is the reference-allele check's finding, so including it would give one defect two names. A differing position is a contradiction only where every side is a substitution; an indel re-anchors legitimately (RM31), so that pair is undecided rather than reported. What the cheapness pays back is ways not to run: no row authoring both halves (nothing_to_check, tested first — a module with no pair has no assembly question either), a non-GRCh38 module (unsupported, since every resolver link is gated on the build), no snapshot opened (offlineorno_reference), and a snapshot that settles none of the pairs —no_referenceagain, neverran(0, 0), because a record of a check that ran over nothing reads as a clean bill. Everything not compared stays outside the denominator and inside the record'sdetail, grouped by reason. - One proof-of-work per call, so a pass collects its records and writes once.
enrich()writes all five of its checks at the end of the run,enrich_literatureits three,check-identifiersits three. The merge is what keeps every command in one document, replacing per check and never erasing a check this run did not put — and untilliteraturewas wired in there was no second writer at all, so the merge was machinery tested only against a document nothing produced. - A skip does not displace an answer the document already holds (RM72, 0.6) — while the authored
bytes stand still. The merge was unconditional newest-wins, so
check-acmg --offlineafter a real run, orcheck-identifiers --no-traitsafter a full one, rewrote a truesubjects=13, findings=0verdict tosubjects=0, skipped=offline: an answer turned into "never asked" by a run that learned nothing. The argument is the merge's own, one step further out — a skip is a run that did not put the question, the same fact as an absent check spelled as a record instead of as a silence, so the protection that already covered the absent check covers this one. Newest-wins still holds between two records of the same disposition. This mattered more than it looked: every check wired after it merges through the same function, so wiring four more members into a rule that can silently downgrade them multiplies the defect by four. The condition is what keeps the fix from being a regression: an answer earns protection by still describing this module, so once the authored bytes have moved (VerificationDoc.module_hashagainst the current binding) a fresh "could not ask" wins — which is exactly whatliteraturerelies on when a module's citations change, and a stale finding no authored edit could clear is a defect this project has removed elsewhere. The counter-argument — that a reader may want to know today's run could not reach the source — is answered rather than dismissed: that is a fact about the run, and this is a per-check document, so it would need a run-level place, which is a different question and is not opened. - The merge carries the author's CLOSURE across, and only while the authored bytes stand still
(RM73, 0.6).
just-dna-compiler closewrites aclosureblock into the same document, so this pass — which rebuilds it rather than editing it — would have dropped a field it did not know about, silently and on every run. Enrichment writes only derived sidecars, which are outside the binding, so the ordinary case is that nothing the author closed has changed and un-closing it would train an author to stop closing. Where the authored bytes have moved, the closure is dropped rather than re-bound: carrying it would have this pass assert a human declared these bytes final about bytes that human never saw. Only the author may make that claim again. This is the never-clobber trap one column over fromSourceRow.datasetanddraft_digest, and the rule generalizes — when another tier starts writing into a document this one rebuilds, ask what the rebuild silently discards. - A pass has ONE subject set, and the gate and the record both read it.
literatureis the worked example and it got this wrong three ways at once, all the same shape.literature.csvis the pin — merge-not-clobber means a re-run refetches nothing — so anything counted inside the fetch loop is a count of this run's requests: thestrictgates saw an empty list and refused nothing while the attestation, reading the rows, reported the finding, so whether--strictrefused depended on how many times--best-efforthad run first, and the module could ship a manifest naming a defect its own strict compile had just blessed. The subject set is now the citations the module makes now, answered by the rows the sidecar holds now, and every number — CLI report, refusal, record — comes off it. Two consequences worth copying. A comparison the sidecar does not pin is re-made over every row rather than over the fetched ones (that is how an existingliterature.csvcame to hide every DOI/PMC id conflict). And the set is the cited rows, not the whole file: the sidecar keeps a row for a citation the author has since deleted, and counting it produced a finding no authored edit could clear, which this project treats as a defect wherever else it appears. - A no-op run records nothing; it does not record a zero.
literature --offlinefetches nothing and re-examines nothing, so on a module that already carries aliterature.csvpinning the citations it makes now, it writes no records at all and the earlier ones stand: a run that has said nothing has nothing to record. With no pin covering those citations there is no answer this run could protect, so the three skips are written and say why — and after an online run that state is reached only by the citations having changed, which is the case the merge rule above deliberately lets a skip win. (This bullet used to argue the silence from the merge replacing per check. That was the merge's defect rather than an argument, and RM72 fixed it.) literatureputs three questions with three different denominators, and one record cannot carry them. Every citation gets an existence verdict; only some carry an authored DOI or PMC id to compare; fewer still carry a quote whose article could be read. Soprovenance_quotecounts the quotes a retrieved text settled — never the quotes authored — with the remainder indetail, and it is a skip (no_reference) rather thansubjects=0when nothing was settled at all, since a zero out of zero on a check that could not run is exactly the clean-looking pass this whole item exists to stop. An abstract-only search settles a hit and nothing else. The remainder itself splits in two and the two must not be one sentence: an article whose text could not be read, and a quote nobody went looking for because the sidecar's pinned counts do not describe the quotes the module carries now. Only the second is cleared by deleting the sidecar, and calling it a retrieval failure is flatly false for an open-access article that was read on the previous run. The pinned count has to match, not merely be non-zero — a row pinned at two quotes says nothing attributable about the one that is left, so the whole row goes unexamined rather than carrying a finding about a quote the module no longer makes. Same rule on the identifier check: a citation drops out of the comparison because PubMed named no identifier or because a curator wrote that row, and reporting the second as the first is a claim about PubMed nobody established. And a DOI has three states rather than two: resolved, a pinned verdict about a DOI the module no longer cites (LiteratureRow.doi_checkedis what makes that visible — without it, correcting a bad DOI made the finding follow the correction), and one no run ever put the question for, which is what a--no-doirun followed by a plain one leaves behind. Both ride indetailrather than shrinking the denominator, since the PubMed half of those citations really was answered.- The attestation is bound to the module's authored bytes. Edit
variants.csvafterwards and the compiler drops the block with a warning — correctly, because the checks were put against rows that no longer exist. Re-running the pass re-attests. Currency of the source is a different question and is read off each record's ownrelease.
Which of these attest, and which are recording passes rather than checks. Every member of
VALID_VERIFICATION_CHECKS either has an emitter or says RESERVED beside itself, and
test_verification_record.py walks the vocabulary and asserts exactly that — no total is stated
here, because the two that used to be were wrong twice each and a number in prose is a registry
nothing iterates.
Nor is a per-command one, and that is RM243's repair rather than a stylistic preference. This
paragraph used to name four commands with a count beside each. They summed to 17 of the 24 emitting
members — in the paragraph that had just refused to state a total, on the grounds that a number in
prose is a registry nothing iterates. Two of the four were stale in the same way: enrich() gained
published_refutation and evidence_status_currency, check-identifiers gained
pgs_accession_currency and pgs_metadata_agreement with the PGS Catalog (RM163), and no addition
moved a figure five lines up. Three commands — clinpgx check-labels, litvar coverage,
alphagenome check — were never in the sentence at all. The figures are not restated here even as
history, because a stale number in quotation marks beside the rule reads, to a skimming reader,
exactly like the rule (RM218 learned that from its own repair).
So the attribution lives in the check table above, where it was already richer than the sentence, and
test_enricher_doc_registries.py walks it: every emitting member has a row, no row names a member the
vocabulary does not hold, no reserved member is shown as attested, and each row's Where cell names
a module that exists. Note what that column is — the site of the comparison, which is often not
the module that writes the record: reference_allele is compared in sequences and attested by
enrich. Where a command rather than a function is the useful pointer, the cell says so in prose. What
this paragraph keeps is the part no walk can say: why a pass is or is not in the vocabulary.
The last four members were wired in RM72 (0.6): the two check commands put a real
authored-versus-source question, reported it to stdout, and let the record die with the process, which
is the sentence RM45 exists to end. What blocked them was their own printed promise to write
nothing; that promise is now the narrower and truer one — writes no authored cell, and records
that the question was put — because there is no other home for was this check run. It cannot be
offloaded downstream (a consumer holding the artifact cannot tell "asked and clean" from "never
asked", which is the whole of RM45) and it cannot go in the authored section (an author does not
attest to a check they did not run). The record is unconditional, with no --attest flag: an
optional record is ambiguous between the check was not run and it ran without the flag, which
reintroduces the two-readings-of-one-absence defect the skip vocabulary was built to end. Filling
acmg_sf or a gene symbol from the registry being asked about it stays refused whatever the
attestation does — that is hints.REDUNDANCY_BEARING, and it is why these are open-ended checks
rather than an apply route. The two remaining members are reserved, and each says so in the code
(vocab.VALID_VERIFICATION_CHECKS): gene_disease_validity and dosage_sensitivity have no emitter
because enrich_gene_validity and enrich_dosage_sensitivity record ClinGen/GenCC verdicts into a
derived table and compare nothing authored. Wiring them would report a check where no question was
put. The line
that decides whether a pass belongs in VALID_VERIFICATION_CHECKS at all is whether it compares
something the module asserts — so gene_validity.csv and clinical_assertions.csv, which record
what ClinGen and ClinVar say and adjudicate nothing, have no check name and must not gain one: a
member for them would let a manifest report a check where no question was put. frequencies.csv,
gene_metrics.csv and the per-article licence columns are the same class. The three rows in the table
above that are recording rather than comparing — Ambiguous back-fill, Article licence and
Source coverage — are kept here because a reader wants the whole surface in one place, and they are
named as the exception rather than left to be inferred.
Two of these break the severity rule in opposite directions, and both are deliberate. The
allele-function check joins the clinical cross-check in warning under strict too: PharmVar and CPIC
are different expert panels — one assigns a molecular function, the other a clinical one — and they
genuinely disagree about some alleles, so failing a compile would make the format arbitrate between
the two authorities it depends on. The declared-use gate goes the other way and refuses in both
modes, because it is not a finding about the data at all: it is a statement that the fetch is not
permitted, and best_effort means "resolve what you can", never "take what you may not".
The clinical cross-check warns in strict too. Every other check compares an authored value against a fact — the
genome's bases, a deterministic digest, a registry's own id — where the source is simply right. A
clin_sig disagreement is two opinions differing, and ClinVar is not truth: a curator who has read the
primary literature and disagrees with a one-star submission is doing their job. Failing the compile
would make the format arbitrate a clinical dispute, which the data-agnostic charter forbids. The
finding therefore carries ClinVar's review-star count so a reader can weigh it themselves.
The division of labour with the compiler is a consequence of Principle 2, not a coincidence: a check
that can be settled by computation over injected data belongs in the compiler (it runs on every
compile, offline, and cannot be bypassed), while a check that needs a reference can only live here.
ref validation is the clean example — the compiler can catch two rows contradicting each other
about a reference base, and only the enricher can catch a row contradicting the genome. See
COMPILER.md § what the compiler can and cannot validate for the full division, including
the blind spots neither tier can close.
The dependency arrow points inward — enricher → compiler → format — so httpx / tenacity /
huggingface_hub never enter the compile path. just-dna-format and just-dna-compiler stay strictly
inject-only (CONSTITUTION Principle 2). HuggingFace is permitted only here (the 0.5 amendment scopes
the Non-goal HF ban to format + compiler); the compiler reaches this package solely through a guarded
lazy import on its deprecated ensembl_cache path — it declares no dependency on the enricher.
Companion docs: COMPILER.md (what consumes
resolution.csv), SCHEMAS.md (ResolutionRowand the three hashes), CONSTITUTION.md (the 0.5 amendment). Import from the submodule where a symbol lives;__init__.pyhas no re-exports.
Read beside this: the 2026-09-11 code-first re-derivation¶
A second reading of this tier, written from the code alone on 2026-09-11, is in
audit/ENRICHER_FROM_CODE.md — every command counted from --help, the
resolver chain with what --offline changes per pass, every cache lane and every client's exception
contract. The 2026-08-18 round found RM97–RM100 (0.6.1); this one found RM208 and RM209. Evidence,
not contract: this document is the maintained one. The method is
BLIND_REDERIVATION.md.
Install¶
pip install just-dna-enricher # runtime: enrich (caches, live Ensembl, live gnomAD, VRS minting)
pip install 'just-dna-enricher[atlas]' # + the AlphaGenome Atlas client (grpcio + protobuf)
pip install 'just-dna-enricher[dev]' # + publisher surface (module/reference upload), snapshot builders (polars), tests
[atlas] is two packages, and it is checkout-only today (RM192, RM196). The AlphaGenome Atlas
serves precomputed variant scores over gRPC; its score fields are plain bytes and its request
filter is an AIP-160 string, so grpcio + protobuf reach every RPC and struct.unpack from the
standard library decodes the payload — 19 MB and +2 packages, measured in a clean venv on
2026-09-10, against 550 MB and 47 packages for uv add alphagenome (measured 2026-09-11). The bindings are generated
from the Apache-2.0 .proto sources, which since RM196 are not vendored: the repository
carries a commit id and a sha256 per file, atlas generate fetches and verifies them, and a file
that does not match its pin is refused rather than used. docs/vendor/alphagenome_protos/ is where
the copy used to be and now holds only a README saying so (RM221 — this sentence still said
"vendored in" five weeks after the sources left). Run just-dna-enricher atlas generate once per
checkout (it needs grpcio-tools,
which is in [dev], not in [atlas] — the runtime imports the bindings without it). An installed
wheel carries neither the sources nor the generated tree, so the client's import is guarded and says
so; RM196 is where that trade gets decided.
ga4gh.vrs is a core dependency, not an extra. Minting a substitution's VRS allele id is stdlib
and lives in the format tier, but justifying an indel needs the reference sequence — and reading
sequence is network access, which is this tier's whole job. The plan for this work budgeted for
ga4gh.vrs[extras] (seqrepo + pysam + hgvs: a compiled extension plus a multi-gigabyte local
sequence store) on the assumption that a local seqrepo was required. Probing showed it is not — core
ga4gh.vrs with the seqrepo REST data proxy normalizes indels over HTTP for 14 pure-Python
packages and no compiled dependencies. At that price there is no reason to make complete allele
identity opt-in, so it is the default and --offline is the only thing that turns it off.
Requires Python >=3.13 (the compiler's floor — deliberately not ensembl-mcp's >=3.14; the query
core was ported, not depended on, dropping fastmcp/eliot). In the workspace:
uv sync --package just-dna-enricher --group dev.
Module map¶
| Module | Role | Notable deps |
|---|---|---|
enrich |
orchestration: enrich() runs the resolver chain, writes resolution.csv |
compiler _load_csv_rows, format ResolutionRow |
transaction |
0.7 (RM128): the run's durability — staged answers beside the target, the advisory flock over the read-modify-write window, and the (done, total) progress unit |
format atomic_writer, stdlib fcntl |
resolver |
the DuckDB rsid↔coord resolver (moved from the compiler in 0.5) | duckdb, format |
clinvar |
the DuckDB ClinVar resolver link (lookup_loci) + the annotation reader (lookup_clin_sig) |
duckdb, format |
clinical |
the clin_sig cross-check over the ClinVar snapshot (offline, reports only) |
format |
concordance |
0.7 (RM130): the other half of that check — the two verdicts, the per-authority calls behind them, and the writer that puts clin_sig_concordance.csv + clin_sig_authority_calls.csv beside the spec |
format concordance, compiler load_csv_rows |
clin_sig |
0.7: the one raw-significance → VALID_CLIN_SIG normalizer, shared by every source that reports one. Dependency-free on purpose, so a runtime pass reads it without the [dev] extra |
format VALID_CLIN_SIG |
net |
shared HTTP politeness: PacingGate, batched, dedupe, attempt_floor, and stream_to_file — the one body every bulk download in this package uses, atomic, translated and retried (RM187) |
httpx, tenacity |
eutils |
NCBI E-utilities client (esummary), shared by the literature and rsID checks | httpx, tenacity |
literature |
pass 4: a module's citations (studies.csv + binning pmids) → literature.csv (PubMed + Europe PMC), fulltext quote match, per-article licence, PMCID→PMID |
httpx, tenacity |
identifiers |
rsID / trait-CURIE / gene-symbol / PGS-accession currency (dbSNP, OLS4, HGNC, PGS Catalog) | httpx, tenacity |
currency |
RM85: has the source a module was drafted from published since? Reads SourceRow.dataset against what the source says now — which release, not whether one judgement moved |
the per-source release records |
provenance |
RM73: telling a value still copied from a source from one a human has edited — the axis every tautology skip reads | compiler draft, format |
verification |
RM45: the load-merge-write that records what each pass checked into verification.json, so the manifest can say it |
format verification |
verdict |
0.7 (RM235): a gate's yes/no with the reason codes that made it a no — the one place the house tri-state is deliberately not used, because --strict must choose an exit code |
— |
producers |
0.7: which pass fills each machine-produced table, from which source, checking what — the derived counterpart to drafting.DRAFT_PROVIDERS, and what lets the docs site's per-table reference name a writer instead of hand-keeping one |
licensing |
pgs |
0.7 (RM163): the PGS Catalog REST client, its release record, and the per-score shape identifiers compares against |
httpx, tenacity |
licensing |
per-source terms + the declared-use gate; emits SourceRow |
format SourceRow |
clingen |
ClinGen dosage sensitivity → gene_metrics.csv rows (CC0, so a module stays sellable) |
httpx, format |
gene_validity |
RM24: curated gene–disease assertions → gene_validity.csv (ClinGen expert panels, GenCC's aggregate; both CC0) |
httpx, format |
assertions |
RM25: resolution.csv + the ClinVar snapshot → clinical_assertions.csv (the call and the review tier) |
duckdb via clinvar, format |
gwas |
RM90: the GWAS Catalog REST API → gwas_effects.csv (published effect sizes with their units). Fills no weight |
httpx, format |
expression |
RM194/RM200: the Atlas ListDenseVariantScores RNA_SEQ interval → expression_effects.csv (per-gene direction, with the distance beside it). Non-commercial, so an undeclared run writes nothing |
atlas_client, gene_spans, format |
gene_spans |
RM194: gene symbol → GRCh38 span, out of the operator-built MANE snapshot. A module with a plain name rather than a private helper, so the second caller can find it | polars, mane lane |
atlas_client |
RM192: the AlphaGenome Atlas gRPC client — ListDenseVariantScores and the interval RPCs, on [atlas] (grpcio + protobuf, 19 MB measured) rather than the 550 MB upstream wheel |
grpcio, protobuf (extra) |
atlas_protos |
RM192/RM196: fetch the Apache-2.0 .proto sources at a pinned commit and generate the bindings. The repository carries the pin — a commit id and a sha256 per file — and neither the sources nor the generated code |
grpc_tools (extra) |
alphagenome_check |
RM193: the Atlas as a resolver, for the three questions the local AVI artifact provably cannot answer. Attests variant_impact_agreement; reports, never repairs |
atlas_client, format |
alphagenome_avi_build |
RM191/RM197/RM198 builder: the 88.5 GB AVI artifact the operator already holds → the alphagenome_avi lane. --input is required and has no default — acquisition is the operator's act |
polars (lazy) |
pgx_draft |
the first drafting provider: CPIC → haplotypes/allele_function/diplotypes rows |
cpic, compiler draft |
clinpgx_draft |
RM26: ClinPGx snapshot → pharm_variants.csv rows (offline, inject-only) |
clinpgx, compiler draft |
drafting |
RM228: the shared drafting scaffold every *_draft.py goes through — the DRAFT_PROVIDERS registry, the derived identity verdict (constructing the row model, because authoring_requirements' any_of grammar cannot express ref/alts require a position), a provider's own SourcePrecondition with its reason as a field, and the licence-row/stale-dataset pair that two providers had implemented as one half each |
compiler draft, licensing |
clinvar_draft |
RM26: ClinVar snapshot → variants.csv partial rows; genotype left to a human |
clinvar, compiler draft |
pubmind |
RM134 § B: the runtime reader over the snapshot pubmind_build writes — duckdb, core install, fetches nothing |
duckdb |
pubmind_draft |
RM134 § C: PubMind snapshot → variants.csv rows, as draft-panel --source pubmind rather than a second command |
pubmind, compiler draft |
clingen_allele |
RM153: the ClinGen Allele Registry — a CAID resolved to an identity this format can carry. The route that places a GRCh37-only row without lifting anything over | httpx, tenacity |
grch37 |
RM48: the old assembly, online half — rs-number recovery and the wrong-build diagnosis over a ref-mismatched row. Warns in both modes | httpx, ensembl |
acmg |
RM72: the acmg_sf check against the published secondary-findings list; the scraped v3.2 page reports unverifiable |
httpx, format |
acmg_build |
[dev] builder: ACMG's published v3.3 workbook → the acmg lane, because NCBI's adaptation is a year behind |
polars (lazy), httpx |
civic_build |
[dev] builder (RM152): a dated CIViC bulk release → snapshot parquet + release.json. CC0, so unlike PubMind this one may be published |
polars (lazy), httpx |
civic_vcf |
RM169: the dated accepted_and_submitted VCF beside the TSV pair — the submitted half, pinnable and byte-reproducible, with no API read |
polars (lazy) |
civic_api |
RM160: CIViC's GraphQL API — the one surface carrying evidence attached to a variant no dated file can place, because a VCF record needs a POS | httpx, tenacity |
civic_identities |
RM159: identities CIViC states in a variant's name and never in its identifier columns (N150fs (c.448delA), IVS2+1G>A) |
clingen_allele, ensembl |
civic_citations |
RM160: the citations those variants carry, and the canary that says when the dated file has caught up | civic_api, format |
civic_draft |
RM152: CIViC snapshot → direction-axis rows. The axis is direction, not clin_sig — the measurement refused the latter |
polars, compiler draft |
civic_refutation |
RM170: an authored direction over a variant CIViC has published a refutation of — the pair contested_variants cannot see. Warns in both modes |
polars, format |
mitomap |
RM171: MITOMAP's published pg_dump reader, and the two-token grammar its status column is written in |
duckdb |
mitomap_build |
RM171 builder: the 63 MB gzipped SQL dump → snapshot parquet + release.json. CC BY 3.0 with commercial use stated free, so a deployment may publish it |
polars (lazy), httpx |
mitomap_miss_build |
RM171: the derived lane — MITOMAP minus the ClinVar cache, recomputed from both parents rather than frozen as a diff | polars (lazy), mitomap + clinvar lanes |
_rcrs |
RM273: the revised Cambridge Reference Sequence (NC_012920.1, GRCh38's chrM) as a vendored constant, so mitomap_miss_build left-aligns both sides' indels without a fetch |
none |
mitomap_draft |
RM171: the rated misses → variants.csv + studies.csv. Only the rated ones; the photocopies and VUS* are each refused for their own stated reason |
mitomap_miss, compiler draft |
clinvar_build |
[dev]: VCF → snapshot parquet; var_citations.txt → citations/ (+ its own release.json block) |
polars, httpx |
lookup |
authoring lookups — rsID validity/loci, ref/alts + populations, which paper a PMID names. Writes nothing | every client above, compiler hints |
pharmvar |
star-allele definitions + function (Api-Key header, 2 rps) |
httpx, tenacity |
cpic |
allele function, diplotype→phenotype, defining variants (PostgREST) | httpx, tenacity |
pgx |
pass 5: cross-check star-allele tables, write the licence table | the three above |
clinpgx_build |
[dev]: summaryAnnotations.zip → snapshot parquet + pinned LICENSE.txt; refuses the retired clinicalAnnotations.zip by name |
polars, httpx |
clinpgx |
pass 6: evidence-level cross-check over the snapshot (offline) | duckdb (core, not polars) |
gnomad |
live gnomAD GraphQL: batched + paced rsid resolution, frequency, gene constraint | httpx, tenacity |
frequencies |
pass 2: resolution.csv → frequencies.csv (per-ancestry-group AC/AN) |
compiler load_csv_rows, format |
gene_metrics |
pass 3: the module's genes → gene_metrics.csv (snapshot first, live API second) |
duckdb, format |
constraint_build |
[dev] builder: gnomAD constraint TSV → gene-level parquet + release.json |
polars (lazy), httpx |
vrs |
GA4GH VRS allele-id minting onto resolution.csv (substitutions stdlib, indels normalized) |
ga4gh.vrs |
sequences |
reference-sequence access (cached) + the reference-allele check | ga4gh.vrs |
locations |
cache-location resolution — one resolve/default_dir/env_var triple per lane in CACHE_LANES, paired by a test rather than counted here — plus .env and the snapshot root filenames (moved from the compiler) |
platformdirs, python-dotenv |
download |
HuggingFace snapshot download, footer-checked and atomic — one ensure_* per lane whose publish_repo is set, which is the registry's own answer to which lanes may be pulled rather than a second list. A lane with no repo is operator-built; see The caches |
huggingface_hub (lazy) |
cpic_build |
[dev] builder (0.5.1): the whole CPIC PostgREST database → five parquets + release.json |
polars (lazy), cpic |
pharmvar_build |
[dev] builder (0.5.1): /genes → alleles + defining variants. Operator-built, never published |
polars (lazy), pharmvar |
pubmind_build |
[dev] builder (0.7, RM134): the ANNOVAR-distributed PubMind table → one parquet + release.json. Operator-built, and pubmind publish refuses |
polars (lazy), httpx, clin_sig |
mane_build |
[dev] builder (0.7, RM168): one pinned MANE release → three parquets + release.json. Operator-built; there is no mane publish because NCBI states a policy rather than a licence |
polars (lazy), httpx |
strchive |
RM165: the STRchive repeat-locus catalogue — the repeat_alleles.csv band cross-check (offline, reports only) |
format binning, compiler load_csv_rows |
strchive_build |
RM165 builder: download STRchive-loci.json + release.json. Core httpx only, no parquet and no [dev] extra |
httpx |
strchive_draft |
RM165: STRchive → repeat_alleles.csv partial rows — identity only, bands left to a human |
strchive, compiler draft |
ensembl |
live Ensembl: V2 GraphQL → V1 REST fallback, tenacity | httpx, tenacity |
upload |
publisher surface — push a compiled module or a reference snapshot to HF ([dev]) |
huggingface_hub (lazy) |
litvar |
RM167: LitVar2/PubTator3 literature coverage per locus, with the tier that answered (allele node / position node / absent). Reports only; writes no row and no SourceRow |
httpx, tenacity, clingen_allele |
drug_labels |
RM166: five regulators' drug labels — the (gene[, allele], drug) cross-check at two join tiers (offline, reports only, writes no SourceRow) |
duckdb, format pgx, compiler load_csv_rows |
drug_labels_build |
[dev] builder (0.7, RM166): ClinPGx's drugLabels.zip → one parquet + LICENSE.txt + its own release.json, dated from the archive's own CREATED_*.txt |
polars (lazy), httpx |
caches |
The cache registry (CACHE_LANES) and the rebuild adapters behind cache rebuild — one entry per lane carrying its three stages and, for each stage it lacks, the reason as a field. Walked by a test against the *_build modules on disk (RM176) |
— |
cli |
Typer app: the runtime commands (enrich, frequencies, gene-metrics, gene-validity, assertions, enrich-and-compile), the authoring ones (hint, draft-*, the check-* family), upload, cache status/pull/rebuild/prune, vrs mint, and one builder group per CACHE_LANES lane carrying a build_command — a publish exists exactly where the lane names a publish_repo, and pubmind publish is the one that exists in order to refuse with its reason. The per-command flag tables are --help and the dated listing in audit/; a hand-kept copy here would be the second thing to update |
typer |
Rate limits (public APIs)¶
Every live client that can fire more than a handful of requests goes through net.PacingGate
(injectable clock, pace-before-retry). Snapshot / one-shot downloads are listed too so the full
egress surface is in one place. Published is what the service documents; our pace is what
the client actually waits; when the source publishes nothing, the gate is a courtesy, not a claim
that the ceiling is known.
PacingGate.spent is what the gate admitted (S95, RM203). A host metering egress per upstream
had no number to meter: nothing downstream reported the calls actually made, so a proxy charged by the
shape of a request, an upper bound that bills a caller for a call that never happened. Every egressing
client waits on its gate once per attempt, inside its retry loop, so one increment is one upstream
attempt — a 429 retried three times counts three, a snapshot hit counts nothing. Monotonic, bumped
under the slot lock, never reset; a rate is two readings apart. It is the one thing this counter says;
what the call cost stays the client's to know. Five clients keep their gate private as _gate
(pgs, litvar, civic_api, pharmvar, clingen_allele), so there the reading is
client._gate.spent; the rest expose client.gate.
One PacingGate is safe to share across threads, and that is now a stated contract rather than an
accident of who happened to call it (S15). It matters because the injection API asks for sharing:
LookupClients tells callers to hold a client and reuse it — a fresh one per question would discard
exactly this state — so a server running blocking work through a thread pool ends up with several
workers on one gate by following our own advice. Until 0.5.4 wait() read last, slept, then wrote it
with no lock, so two workers could both find the interval elapsed, both skip the sleep, and turn a
published 3/s budget into 6/s — a budget somebody else enforces by blocking the operator's IP. The lock
covers the bookkeeping only: each caller reserves the next free slot and waits for it alone, so N
callers get N slots spaced one interval apart and no worker is blocked by another's sleep.
Single-threaded behaviour is unchanged, and test_net.py proves the spacing on a frozen clock without
really sleeping. It is a pace, not a concurrency limit. If what you need is one request in flight
per service, put a semaphore around your call site; the gate will not give you that.
| Service | Used by | Published budget | Our pace / batching | Auth / identity |
|---|---|---|---|---|
| gnomAD GraphQL | gnomad (resolve, frequencies, live constraint) |
10 req / IP / 60 s | min_request_interval=6.0 (exactly that budget); GraphQL alias batches of 20 (25 worked live; 29 → HTTP 400) ≈ 200 variants/min |
none |
| NCBI E-utilities | eutils → literature + rsID currency |
3 req/s without a key; 10 req/s with NCBI_API_KEY |
1/3 s or 1/10 s from whether the key is present; esummary batches of 200 |
tool=just-dna-enricher; email from JUST_DNA_CONTACT_EMAIL when set; key optional |
| PharmVar | pharmvar / pgx |
2 req/s (OpenAPI “Limitations”) | PHARMVAR_MIN_INTERVAL=0.5 |
Api-Key header from PHARMVAR_API_KEY (personal; never written to a module) |
| Crossref | literature.CrossrefClient (DOI existence) |
polite pool: 10 req/s single-DOI (5 public); concurrency 3 polite / 1 public (Crossref docs, Dec 2025 revision) | min_request_interval=0.1 (10/s — polite single-DOI ceiling) |
User-Agent just-dna-enricher (mailto:…) when JUST_DNA_CONTACT_EMAIL is set → polite pool; omitted rather than invented |
| Europe PMC | literature.EuropePmcClient (OA fulltext + abstracts + per-article licence) |
no durable official figure on the developer pages (community reports vary) | min_request_interval=0.5 (2/s), batches of 25 on search |
none |
| PMC ID converter | literature.PmcIdConverterClient (PMCID → PMID, reporting only) |
no published figure; the service documents a 200-id batch ceiling | min_request_interval=0.5 (2/s), batches of 200 |
tool=just-dna-enricher; email from JUST_DNA_CONTACT_EMAIL when set |
| OLS4 + HGNC | identifiers.OntologyClient |
neither publishes a documented limit | min_request_interval=0.2 (courtesy — GET-per-id, unbatched) |
Accept: application/json |
PGS Catalog REST (www.pgscatalog.org/rest) |
pgs.PgsCatalogClient → identifier currency + the release probe |
none published | _REQUEST_INTERVAL=0.2 (courtesy — GET-per-accession, unbatched; one extra request for /rest/info) |
Accept: application/json |
LitVar2 / PubTator3 (ncbi.nlm.nih.gov/research/litvar2-api) |
litvar → literature coverage (0.7, RM167) |
NCBI asks for no more than 3 req/s without a key | _REQUEST_INTERVAL=0.34 — that budget with a little air in it; a 388-locus corpus sweep ran at it without a refusal |
none |
CIViC GraphQL (civicdb.org/api/graphql) |
civic_api → civic citations, evidence-status currency (0.7, RM160) |
none published | _REQUEST_INTERVAL=0.34 (courtesy) |
none |
ClinGen Allele Registry (reg.clinicalgenome.org/allele) |
clingen_allele → the identity-from-a-name procedure (0.7, RM153) |
none published | _REQUEST_INTERVAL=0.1 (courtesy) |
none |
Ensembl GRCh37 REST (grch37.rest.ensembl.org) |
grch37.diagnose_wrong_build (0.6), hint recover |
Ensembl publishes 15 req/s | _REQUEST_INTERVAL=0.1 — keeps a whole panel inside it while still interactive |
none |
| ClinVar release probe | currency.ClinVarReleaseClient — one bounded read of the VCF's header (HEADER_PROBE_BYTES) per --verify-datasets run (0.7) |
NCBI asks for no particular pacing on the FTP-over-HTTPS mirror | CLINVAR_MIN_INTERVAL=0.5 — this tier paces every service it touches, and one probe per run is inside any budget |
none |
Ensembl REST (rest.ensembl.org) |
ensembl V1 fallback |
15 req/s per IP, ~55 000 / rolling hour; 429 + Retry-After / X-RateLimit-* |
no PacingGate — live path is the last link after cache/snapshot, so volume stays low; tenacity on transport only |
none |
Ensembl GraphQL (beta.ensembl.org) |
ensembl V2 first try |
unpublished (beta) | no PacingGate; 5xx falls through to REST |
none |
CPIC (api.cpicpgx.org) |
cpic / pgx_draft |
unpublished | no PacingGate — coarse PostgREST GETs (gene-scoped), not per-allele loops |
none |
| ClinPGx | clinpgx / clinpgx_draft |
n/a at runtime | offline snapshot only for the check/draft path — no live poll budget | none (live API retired → snapshot) |
seqrepo REST (services.genomicmedlab.org) |
sequences / VRS indel mint |
unpublished | no PacingGate; in-process memo of window reads |
none |
GWAS Catalog REST (www.ebi.ac.uk/gwas/rest/api) |
gwas |
none established — EBI publishes no numeric budget for this API, unlike gnomAD's real and load-bearing 10/60s | DEFAULT_REQUEST_INTERVAL=1.0 — a courtesy, not a transcribed limit. Cost is 1 + 2N per variant (the study/trait facts sit behind _links); measured at 382 requests for one real module, and --no-study-facts drops it to one per variant |
none |
AlphaGenome Atlas (gdmscience.googleapis.com) |
alphagenome expression, alphagenome check |
none published — the Additional Terms bar classes of holder outright rather than metering requests, which is an eligibility bar and not a budget | measured ~1,091 SNVs/s end to end; a gene plus its ±512 kb flanks is ~3.3 M SNVs ≈ 50 min, which the pass prints before the query rather than after | ALPHAGENOME_API_KEY |
| ClinGen dosage TSV | clingen |
n/a (one file) | single download, then local parse | none |
| ClinGen gene-validity CSV | gene_validity |
n/a (one file, ~1 MB) | single download, then local parse | none |
| GenCC submissions CSV | gene_validity |
n/a (one file, ~28 MB) | single download, then local parse; the client's timeout is 180 s because one response is the whole export | none |
| ACMG SF list | acmg |
n/a | one HTML GET (~75 KB) or a local --sf-list workbook |
none |
| Hugging Face Hub | download (snapshots), upload (modules / references) |
5-minute fixed windows, three buckets — see below | no custom gate; huggingface_hub handles 429 via RateLimit / RateLimit-Policy headers (smart retry in 1.2+) |
HF_TOKEN / hf auth login — anonymous shares a per-IP pool; a free token is the usual fix |
Hugging Face Hub tiers (as of Sep 2025)¶
Quotas are per 5-minute window. Snapshot provisioning (ensure_*) and publisher upload both hit
API (listing / commits) and Resolvers (/resolve/ file bytes). Pages is the website and is
not on our path. Numbers marked * are subject to change with platform health.
| Plan | API | Resolvers | Pages |
|---|---|---|---|
| Anonymous (per IP) | 500 * | 3,000 * | 100 * |
| Free user | 1,000 * | 5,000 * | 200 * |
| PRO | 2,500 | 12,000 | 400 |
| Team org | 3,000 | 20,000 | 400 |
| Enterprise | 6,000 | 50,000 | 600 |
| Enterprise Plus | 10,000 | 100,000 | 1,000 |
| Enterprise Plus + org IP ranges | 100,000 | 500,000 | 10,000 |
| Academia Hub org | 3,000 | 20,000 | 400 |
Org limits apply per member, not shared. Source of truth:
huggingface.co/docs/hub/rate-limits. Always pass
HF_TOKEN for snapshot download when possible — anonymous traffic is the usual cause of a “stuck”
ensure_snapshot (the client is sleeping on a 429, not hanging).
Rules that follow from the table¶
- Pace before retry.
tenacitybacks off on transport / 429, but a blind retry spends the same budget that caused the 429; every gated client waits first. - Reuse clients. A
PacingGateis per-client state — constructing a fresh gnomAD / eutils client per lookup throws the interval away (lookupholds them for that reason). - Batch where the API allows it. gnomAD aliases (20) and NCBI esummary (200) exist because the published ceilings make one-id-per-request unusable; PharmVar is the opposite — 2 rps forces gene-scoped endpoints, never per-allele.
- No response cache for the live clients. NCBI / PharmVar / gnomAD GraphQL / Crossref / Europe PMC
/ OLS4 / HGNC / PGS Catalog / LitVar2 / the CIViC API / the ClinGen Allele Registry / GRCh37 Ensembl /
the ClinVar release probe / live Ensembl are paced only. Persistence is the authored sidecars
(
resolution.csv,frequencies.csv, …) and the published snapshotscache pullfetches (theCACHE_LANESrows with a repo — see "The caches") — delete a sidecar to force a refetch. Note which entries in that list are also the licence-gated ones: CPIC and PharmVar are paced-only and forbid sale, so they are the two RM38 gives a snapshot (see On a host, or in a service below). - A shared IP shares one budget. Every figure above is per-IP or per-token, never per caller, so a hosted deployment multiplies its users onto one allowance rather than getting one each. This is a separate reason from licensing for reaching a source through a snapshot, and it applies to ungated sources too — it just happens that the gated ones are where both reasons land at once.
--offlineclamps to local caches / sidecars; it does not invent a budget for a live API.- The retry ceiling is a floor a deployment can raise —
$JUST_DNA_HTTP_RETRY_ATTEMPTS(RM42). Three attempts is right for the audience the CLI was written for: a person who would rather see a failure in ten seconds than wait out a flapping upstream. It is wrong for the other shape the 0.5 tiering created — a server runningenrich()inside an unattended publish, where giving up on a transient 502 costs the publisher a whole re-upload rather than ten seconds. Two callers wanting opposite things from one constant is a knob, and it was an import-time decorator argument with no parameter, so a consumer's only route was to walk the package and reassignpolicy.stop.
net.attempt_floor(n) resolves per call. It raises each client to at least the configured value
and never lowers one, so the deliberate per-client tuning survives — gnomAD and eutils sit at 4
because their budgets are tightest, and a value that set every client would flatten that. Safe to
raise precisely because every gated client paces before it retries: an extra attempt spends a slot
of the published budget rather than bursting past it. Only a bare stop_after_attempt is replaced —
a composed stop_after_attempt(3) | stop_after_delay(60) means both, and raising one term would
silently change a policy whose author meant the conjunction. None is composed today — and the number
of policies is deliberately not stated here: net.py deleted "the nine" from its own comments after
the tree carried twelve, because a count in prose is a registry nothing iterates.
The Atlas client reads it too since RM280, and it had to lose grpc's retry to do so. Its channel
used to retry under the vendored grpc_service_config.json, which grpc caps at five attempts, the
number upstream already asks for, so the variable could not raise it. A tenacity layer on top
would have made 5 × N attempts instead of a floor. atlas_client.channel_options() now switches grpc's
retry off (grpc.enable_retries = 0, and no retryPolicy in the config it hands the channel), and
AtlasClient._call retries under attempt_floor(5) with upstream's backoff. What it retries is
derived from _translate: every status that becomes AtlasUnavailable, which adds INTERNAL to
upstream's three. Its pacing gate has a zero interval, on a measurement (RM307): no throttle
appeared up to 21 calls/s (ALPHAGENOME_ATLAS.md § 6.7), so the
"paces before it retries" reasoning holds trivially. The gate is still waited on once per attempt,
so spent meters Atlas calls, and connect(gate=…) takes a shared one with a real interval.
The author's overlay, read but never written (RM136, 0.7)¶
The compiler applies overrides.csv before any check reads a row — a check must report on what the
module asserts. Until 0.7 this tier did not, so an author who corrected a resolution.csv cell
through the overlay went on being told the same finding on every run, with no way to clear it and
nothing saying the correction had been recorded and honoured one tier over.
Three rules, and each has a refused alternative worth knowing:
- Read-only. The enricher never writes through the overlay. An overlay row is the author's answer to a difference, never this tier's (RM83).
- Input reads only.
licensing.overlaid_input_rowsis called where a pass reads a derived table as an input to something else —frequencies,assertions, andidentifiers' gene-locus check all readresolution.csvthat way. Every merge baseline stays raw: a pass that reads its own output file to merge against it writes that file back, and post-overlay rows would bake the correction into the derived table. Read the file you write, and write what you read. - Per field.
licensing.overlay_answers(spec_dir, table)returns the(subject, field)pairs anupdatenames. A finding is answered when the overlay corrects the very cell the finding is about, so a coordinate correction silences the coordinate check and leaves an unrelatedclin_sigfinding standing. Per row was cheaper and was refused — it would silence findings the author never looked at, which is the silent-suppress hole the overlay's own design calls its worst case.
There is no second implementation of the overlay, deliberately: compiler.load_overlay is public
for this, and overlaid_input_rows calls the same apply_overrides the compiler calls. Two copies
would drift on the normalization seam that produced a silent P7 break in this feature's first week, and
a test compares the helper's output against apply_overrides directly so a future copy would show.
Answered is not agreed. The rsid↔coordinate check is the one wired to this today, and an answered
pair leaves disagreements while staying inside subjects, counted by PairCheck.answered and
reported as one INFO line naming what happened. The comparison ran and found a difference; removing it
from the denominator would publish a cleaner module than there is. Only update counts — an insert
answers no prior finding, and a suppress is already reported in its own right by RM131.
Wiring a further check to the answered set is deliberate rather than automatic, because which cells
does this comparison read is a per-check fact (resolver._COORDINATE_FIELDS is the first one written
down) and guessing it wrong silences a finding nobody answered.
A second answered-set reader arrived with RM151, and it reads the overlay the other way round.
licensing.overlay_answered_subjects(spec_dir, table) returns the (subject, member) pairs the
overlay names — every operation and every field, which is the opposite rule from overlay_answers
above and is right for the opposite direction. overlay_answers decides whether a finding may be
silenced, so it insists the overlay touched the very cell the finding is about; anything looser
silences findings the author never looked at. overlay_answered_subjects raises a finding, and what
goes stale there is the reason — mandatory on every overlay row whatever the row does — so a
per-field rule would have to name a column the reason does not live in. A finding raised too widely
costs a reader one line; one silenced too widely costs them the finding.
What --offline promises, and the one axis the two readings differ on (RM220)¶
--offline means no egress. Every pass that takes it turns into a no-op with a warning rather
than a failure where it has no snapshot to fall back on, and the warning says which of the two it is
(@unreachable-not-absent — nobody-asked is not asked-and-absent).
An injected client is where the tier held two readings, and the axis is the source's licence.
| reading | passes | why |
|---|---|---|
offline outranks an injected client |
pgx, enrich, frequencies, expression |
the source is licence-gated. A live client under a flag documented as making no egress is exactly the loophole RM38 closed — test_pgx_licensing.py asserts it by name, and the AlphaGenome Atlas joined that column in RM220 because its Additional Terms bar classes of holder outright |
| an injected client wins | gwas |
the GWAS Catalog is ungated, and handing over a transport you already hold is not egress. Stated in enrich_gwas's own docstring, deliberately |
Neither is wrong and the difference was nowhere written down, which is the defect RM220 actually
repaired: expression had gwas's shape against pgx's situation, so an injected client fetched
from a gated source under --offline. @flag-means-same is the rule, and this table is what makes
the flag mean the same thing given the licence, rather than differing by which module you happened
to call.
gwas keeps its behaviour on purpose. Changing a contract its docstring states, for the one
ungated source, is a decision rather than a repair — recorded here so the next reader meets a written
rule instead of an apparent inconsistency.
Injecting text is a different thing entirely and is not part of this: clingen,
gene_validity and acmg accept an already-downloaded export, which cannot egress by construction.
That is the documented "inject what you already hold" escape.
enrich() — the resolver chain¶
enrich(spec_dir, *, mode="best_effort", offline=False, ensembl_cache=None,
clinvar_cache=None, use_clinvar=True, use_gnomad=True, download=True,
genome_build=None, write=True, mint_vrs=True,
verify_ref=True, verify_clinsig=True, verify_rsids=True, verify_datasets=True,
keep_par_twin=False, rederive=False, keep_staging=False, progress=None,
resolver=None, gnomad_client=None, grch37_client=None,
release_probes=None) -> EnrichmentResult
genome_build=None means "read the module's declaration" (spec_genome_build), not "assume
GRCh38". It defaulted to the literal "GRCh38" and no caller ever passed anything else, which made every
genome_build == "GRCh38" gate below — and the warning saying a non-GRCh38 module resolves nothing —
unreachable: a genome_build: GRCh37 module was resolved against GRCh38 Ensembl and the GRCh38
coordinate written into its resolution.csv under the label GRCh38, silently. Enrichment is
GRCh38-bound (RM15), so for any other build it now warns, runs no link, and records no lookup
result — not even not_found, which would claim the source was asked. Authored coordinates are still
transcribed verbatim, under the module's own build. An explicit value stays the inject-only override.
Reads every table that can ask for a coordinate — variants.csv, pharm_variants.csv and
haplotypes.csv — computes which rows still need work (need_pos = rsid but no coord; need_rsid =
coord but no rsid), runs a first-hit-wins chain, and writes/merges resolution.csv (sorted by
(variant_key, locus_index)), stamping each row's source/status. The chain:
- Existing / human rows — a
resolution.csvalready beside the spec is authoritative and never clobbered; avariant_keyit already covers is skipped by the chain. - Local cache (offline) —
locations.resolve_ensembl_referencelocates a cache;resolver.lookup_lociruns the DuckDB lookups (rsid → [loci],pos → rsid). Sourcecache. - HF snapshot — if no cache is present and not
offline/download=False,download.ensure_snapshotfetches the parquet slice (populates the cache read in step 2). The snapshot is a static slice of popular rsIDs, not a canonical reference — same pains as any source (incompleteness, versions, reachability), so a miss falls through. - ClinVar cache (offline,
use_clinvar) — for the variants the Ensembl cache/snapshot missed,clinvar.lookup_lociruns the same DuckDB lookups over the ClinVar snapshot (locations.resolve_clinvar_reference, ordownload.ensure_clinvar_snapshotwhen absent and online). Sourceclinvar. It sits after the Ensembl cache on purpose:altsis a resolution fact (it flows intoweights.parquet→artifact.digest), and ClinVar carries only its submitted alleles while Ensembl carries every dbSNP allele — so a variant both caches know keeps the Ensemblaltsandsource="cache", and no already-compiled module'sartifact.digestmoves. ClinVar is a complementary reference (4.4M clinically-curated records, 1.54M with no rsid) that makes an offline clinical enrich possible without provisioning the 14 GB dbSNP cache. - Live Ensembl — for rsIDs every cache missed,
ensembl.EnsemblResolver.resolve_rsid(V2 GraphQL → V1 REST fallback). Sourcesensembl-graphql/ensembl-rest. It has four outcomes: loci, an answered[],(None, None)for could-not-ask, and(None, source)for an answer whose one-sided indels could not be anchored (RM268) — see Live Ensembl below, and note that neitherNoneleaves anyresolution.csvrow at all. - Live gnomAD (
use_gnomad) — the last link, for whatever nothing else could resolve.gnomad.GnomadClient.resolve_rsidsbatches the leftovers. Sourcegnomad. It goes last for exactly the reason ClinVar goes after the Ensembl cache: gnomAD reports only the alleles observed in gnomAD, not every allele dbSNP knows, so promoting it would narrow some already-compiled module'saltsand move itsartifact.digest. Last place makes the link strictly additive — it can only ever add variants nothing else had. A failure here is logged and skipped, never fatal.
After the chain settles, vrs.mint_resolution_rows stamps vrs_id/vrs_spec onto every mintable row
(mint_vrs=True). An existing id is never overwritten.
The run is a transaction (0.7, RM128)¶
enrich() used to persist nothing until its tail: a single if write: block at the bottom, everything
above it in memory, so a run killed at minute 29 had written zero bytes and the thirty minutes were
gone (S66). The obvious repair — checkpoint the table as it goes — trades away a property somebody may
be relying on, that a refused strict run leaves the module exactly as it was. The trade is
unnecessary. The run is a transaction, which keeps the promise absolutely and recovers the work.
- Staged answers, in the target's own directory. As each live link answers, the raw answer
(
rsid → [locus], plus which link said so) is written to.<name>.staging/answers.csvbeside the table —transaction.ResolutionJournal,transaction.staging_dir_for. Same directory is the correctness condition, not a convenience:os.replaceis atomic only within one filesystem, andshutil.moveacross a partition degrades to copy-then-delete. Staging beside the target makes a cross-device move structurally impossible rather than merely avoided, and a test asserts the sibling relationship rather than trusting the prose. - What is staged is the answer, never the row. Everything downstream of an answer — the
allele-aware hosting filter, the pseudoautosomal selection,
locus_index, the minted ids — recomputes on the resumed run, so a flag that changed between the kill and the resume changes the table exactly as it would have on a first run, and the journal cannot carry a stale derivation. What it cannot survive is a differentgenome_build, which is checked and, on a mismatch, discarded whole with a warning. Only positive answers are staged: a request that failed is unchecked rather than absent, and freezing that into the journal would make a transient outage permanent. - Seeded between the caches and the live links, so a snapshot provisioned between the kill and the resume still wins the variant it would have won on a first run — and the live request is not made twice. That placement is what makes a resumed run produces the table an uninterrupted run produces a property rather than a hope; it is the item's P7 obligation and it has a test.
- A staged answer is honoured only if the link that produced it would run this time. The seeding
reads the same two booleans that decide whether the live Ensembl and gnomAD blocks execute, so a
--no-gnomador--offlineresume drops what that link had already answered instead of stamping asource="gnomad"row a first run with those flags could never have written.altsis inRESOLUTION_FACT_FIELDS, so honouring a switched-off link would move the compiled digest too. The dropped answers are logged, and their subjects go back to the chain. - The gate commits. Every refusal (
refmismatches, withdrawn rsIDs, stale rsIDs, unresolved keys) raises before the write block, so a refusedstrictrun changes nothing is now a written promise rather than an accident of statement order. The test asserts it on the bytes on disk of a pre-existing table, not on a return value. The staged answers are left behind by a refusal, which is the point: the next run resumes. keep_staging=True/--keep-stagingleaves the staged answers after a successful commit, for debugging, and logs where they are. The default removes them.- Not mode-conditional.
write=Truemeaning "at the end" understrictand "as we go" underbest_effortwas refused in advance — a flag must mean the same thing in every function that takes one. Under a transaction it does, because committing is the only write.write=Falsestages nothing and takes no lock: with nothing written there is no window to exclude.
The licence row is part of the table's commit, in every pass (S98, RM231). Eight passes wrote
their data table and then merged the SourceRow; a scaffold's placeholder licensing.csv made the
merge refuse after alphagenome expression had already landed 12,003 non-commercial rows, and the
compile gate — keyed on the licence table and nothing else — then passed them as unrestricted. Now
each writer takes the merge as before_commit, which layout.atomic_writer runs after the temp file
is fsynced and before the rename: a merge that refuses removes the temp, a table that fails to
serialize never reaches the merge, and neither file exists without the other. Every pass also reads
the table before its fetch (licensing.require_sources_file), so a placeholder fails in a second
rather than after a whole-gene query. The one residual — the table's own rename failing after the
row's has returned — leaves a licence row for data that never arrived, and the error says so. The
drafting surfaces record their row after the compiler's draft append, the same gap one layer over,
and the guard names them as RM228's to close.
The advisory lock, and how it degrades¶
The transaction does not close the concurrency window. Two runs can each stage and each commit, last writer winning over a merge with neither knowing — the reported incident is the sharp form, where a client-side kill did not stop the worker, a zombie run reached the write and overwrote a restored 330-row table with 162 rows, and the module then validated, closed and compiled green. Nothing downstream could see it: the three branches that deliberately write no row for an unanswerable subject make a shorter table indistinguishable from a module whose author resolved less, and those branches are correct.
So the whole read-modify-write window is held under fcntl.flock on the spec directory's own
descriptor (transaction.spec_lock). No lockfile, because one left behind by exactly the kill this
item is about would block every subsequent run — a worse unattended failure than the one it prevents —
and the staleness rule that would fix it is a clock. flock dies with the process, so there is nothing
to expire. The lock is non-blocking: a second run refuses with a pinned message rather than waiting
half an hour behind a zombie, and the refusal is accurate by construction since the lock is only ever
held by a live process.
spec_lock is the first thing enrich() does, ahead of every loader, so a spec_dir that is not
an openable directory meets it before anything else. It refuses with EnrichmentError rather than
letting os.open's own exception past the pass that owns the call.
The degradation is documented rather than silent. Without fcntl (a non-POSIX platform), or on a
filesystem that refuses the call — a network mount may answer ENOLCK/EOPNOTSUPP, or emulate flock
in a way that excludes nothing — the run proceeds and logs Advisory locking is unavailable for …, so
this run is NOT excluded from a concurrent one. flock is untested here on the network filesystems
a consumer may use; serialize enrichment of one spec directory yourself where that matters. Both
degradation branches are exercised by tests, because an unreached branch is not an API.
progress — (done, total) over subjects¶
progress: Callable[[int, int], None] | None = None. Two integers, no object, no event vocabulary to
keep working forever. The unit is argued rather than guessed, because P3 keeps a leaf working:
- The incident is an idle timeout. Both reported runs died at 1800 s with essentially every variant resolved, so what the caller needs first is a keepalive with monotonic progress — which rules out phases, since a 29-minute phase emits nothing and the timeout fires anyway.
totalmust be known up front for the number to mean anything to a caller rendering it. The subject count is; the link count is not, since it depends on what resolution finds.- Subjects are the only unit the author's mental model already has. Links are an implementation
detail of the batched resolver, and publishing one would make a refactor of
resolver.pya contract change.
(0, total) is reported before any work starts. done is the size of a set that only grows, so
monotonicity is structural rather than promised, and the assembly loop touches every subject, so the
last report is always (total, total). The callback's exceptions are not swallowed — under the
transaction an abort there leaves the staged answers exactly as a kill does.
rederive — the drift canary, performed (0.7, RM83's residue)¶
An ordinary run gap-fills and never re-asks about a recorded subject, so a source that silently
revised an answer moves no fetched_at, no fact signature and no digest: MODULE_LIFECYCLE § 5.1's
canary was an instrument that could not fire. rederive=True / --rederive re-asks every subject
and reports which recorded ones came back different, as EnrichmentResult.rederived.
It composes with the transaction rather than adding machinery: the recorded table is already in memory
and the fresh one has not been committed, so both sides exist at the commit boundary and the comparison
is free. The comparison is over RESOLUTION_FACT_FIELDS only — read off the registry rather than
restated — because the provenance columns move on every run by design.
Noneis not[].Nonesays nobody re-derived;[]says every recorded subject was re-asked and every one still answers the same. Only the second is a clean bill, and only a real difference is printed — an empty comparison must not announce a zero as though it were evidence.- A recorded subject this run could not ask about keeps its recorded rows. Otherwise re-deriving
would be a way to shorten the table, which is the reported incident wearing a new flag; an offline
--rederivewould replace a full table with an empty one. Answered-and-absent is an answer and does replace (it writes anot_foundrow); could-not-ask is not. The carry-forward warns, naming the subjects it kept. - A re-derivation resumes only another re-derivation. After a gap-filling run commits, its staged
answers are exactly what produced the recorded table — so a later
--rederiveseeded from them would compare that table against its own provenance and report a clean bill for precisely the subjects it was asked to re-check. The journal records which run wrote each row, a--rederiverun skips the gap-filling ones and says how many it skipped, and the reverse direction is allowed because an answer a re-derivation obtained is still an answer.--keep-stagingis the door this closes: without the rule, a file kept for debugging silences the canary. - The honest limit.
rm resolution.csvplus a re-run re-derives just as correctly and reports nothing: it destroys the old values before the fresh ones arrive, so nothing holds both sides. - What is not built, and each looks obvious from the headline: a
--refreshcommand, a diffs file or table, a proposed table beside the current one, a pass that applies the newer value to an authored or curator-set cell (that destroys the evidence of the upstream change — still the rule), and re-asking every subject on every run (it would put the full resolution time on every pass to buy drift detection nobody asked to run continuously). - Put the cheap question first.
--verify-datasetsbelow asks the same has the world moved about the release label and costs one request per source, so it is what tells you whether re-asking every subject is worth the run. A module whose recorded release is still the one its source publishes has had no upstream revision to detect — and one whose release has moved is exactly the module worth re-deriving.
verify_datasets — has your source published since? (0.7, RM85)¶
SourceRow.dataset has recorded which release a module's rows came from since RM4, and two things read
it — the tautology skip, and licensing.withdraw_stale_dataset when a module ends up mixing two.
Neither answers the question an author who has forgotten, or a curator who inherited the module,
actually has: has ClinVar published since this was drafted? currency.check_dataset_currency is that
comparison, run at the end of enrich() and attested as dataset_currency.
It is a comparison and not a column. The column-shaped repair — a field saying what this module was
made from and what would age it — was refused one table over in RM71 on the grounds that it restates
dataset and then rots where dataset is maintained. This reads sources.csv and writes nothing at
all; repairing a stale label is a re-draft, which is an author's decision and a different command.
- Three states, and the third is the item.
DatasetCurrency.behindisTrue(the source has published since, with both labels named),False(still current), orNone— nobody could ask.Noneis neverFalse.--offlineis where that bites: an offline run makes no request, so every recorded release isunchecked, and a check that reported a clean bill for a source it never reached would be S4's defect wearing the badge of the mechanism S4 built. - The denominator is what was compared.
subjectscounts the legs asked and answered comparably; an unreachable or unaskable source is named in the record'sdetailinstead of being counted, because counting it would publish coverage of a fraction the record does not state. With no leg settled the pass records a skip rather thanran(0, 0). - Comparability is tri-state too.
clinvar_dataset_labelhas two forms —clinvar_2026-08-25from the VCF header andclinvar_sha256:…for a snapshot whose VCF stated no date. They name one release space and equality across the two forms means nothing, so a digest against a stated date is uncomparable (no_reference), not behind. - Severity follows the mode, over
behindalone.strictrefuses on a superseded release and never on an unchecked one: an unreachable source and an--offlinerun both leave every leg unchecked, and escalating those would make--offline --strictimpossible forever over something no author can edit — theunreachable_rsidsrule, which warns in both modes and escalates in neither. - One probe ships, deliberately: ClinVar. It is the source the item is about and the only one this
tier can ask for a release label in the namespace it already records —
currency.ClinVarReleaseClientstreams the live VCF, reads##fileDate=through the same readerclinvar_builduses on a downloaded file, and abandons the stream after 256 kB rather than asking for a byte range a server may ignore. One definition of the label, because two spellings would not fail: they would simply never match, and the check would report every ClinVar module as behind forever. Every other source is honestlyunsupported.currency.PROBE_SOURCESis the set, derived fromdefault_probesrather than restated beside it;release_probes=injects a wider registry, and the injected one — neverPROBE_SOURCES— decides who can be asked. - One request per source, not per row.
sources.csvis keyed(source, layer)and "what does ClinVar publish now" has one answer whatever a module used it for, so the two layers share it and can never be told different things. - It cannot agree with itself. The rows it reads are the ones on disk before this run's commit, and
the licence rows
enrich()writes are at theresolutionlayer with nodatasetat all — so nothing here is comparing the module against something this same run derived. That is the trap next door:--rederiveseeded from its own staged answers would have reported a clean bill for exactly the subjects it was re-checking.
Which tables ask for a coordinate¶
The chain was never variant-specific; only its input was. Until 0.5 it read variants.csv alone, so a
PGx module — which by design carries none (one CSV = one concern) — enriched to an empty
resolution.csv and shipped with no coordinates at all. enrich._collect_subjects normalizes every
eligible row to a _Subject and feeds it through the unchanged chain, caches, ordering and back-fill:
| Table | Identity | Allele constraint fed to hosting_verdict |
|---|---|---|
variants.csv |
frozen variant_key (with alts) |
genotype |
pharm_variants.csv |
variant_key property (without alts) |
genotype (optional — None keeps every locus) |
haplotypes.csv |
derived the same way, without alts |
the defining allele |
heteroplasmy.csv (0.5.3) |
variant_key property, with alts — it mints one exactly as VariantRow does |
none: a measurement band over a locus is not a claim about a genotype |
A HaplotypeRow reuses the same membership predicate rather than a parallel one: its defining
allele is the one-allele form of the question a genotype asks of two. Subjects are deduped by
variant_key with variants.csv first, so when two tables name one variant the SNP row wins — it
is the only one carrying alts, a resolution fact, and letting a PGx row win would move an
already-compiled module's artifact.digest. The PGx tables key without alts deliberately: a
pharm annotation or haplotype junction matches a variant at chrom:start:ref regardless of allele.
heteroplasmy.csv joined the list in 0.5.3, and it is the one row here that is build-dependent.
Its coordinates are optional exactly as the PGx ones are, so an rsid-authored heteroplasmy module
resolved to nothing at all — the same gap PGx had before 0.5, found from the other end when the
compiler started reporting which tables a VCF cannot join (COMPILER.md § Scope). Because its
variant_key carries alts, it can mint a ga4gh:VA.…, so that load passes the module's
genome_build where the two PGx loads rightly do not — the RM36 trap, one call site further on.
What the enricher resolves is not what the compiler applies. These subjects all land in
resolution.csv, and the compiler applies that table to variants.csv only; a PGx or heteroplasmy
table is materialized verbatim. So enriching a PGx module is still worth doing — the table records the
coordinates, and a consumer can join them itself — but the parquet keeps the author's nulls until
RM43.
Multi-allelic snapshot rows¶
The Ensembl snapshot stores a multi-allelic site as one row whose alt is pipe-joined (A|C|T),
while live Ensembl, ClinVar and gnomAD all emit comma-separated lists. resolver._snapshot_alleles
normalizes at that one boundary so a locus dict's alts has a single canonical shape.
This is load-bearing, not tidying. The hosting predicate splits alleles on commas, so an un-normalized A|C|T
collapsed into one opaque "allele", no genotype was ever a subset of {ref} ∪ alts, and the
allele-aware filter dropped every locus — a cache-resolved rs4244285 with the ordinary genotype
A/G resolved to not_found. The reverse back-fill had the mirror bug, comparing an authored alt
against the whole joined cell with !=. Both are pinned by tests that fail on the pre-fix code.
A located but unusable cache (a stale snapshot, or a parquet some other tool left in the cache directory) is treated as a miss, with a warning — one optional link's bad data must not sink an enrichment the other links can still complete.
--offline clamps the chain to the local caches (steps 2 and 4 — Ensembl and ClinVar; guaranteed
zero egress). Every filled row records the link that won in source, so the compiler can surface
resolution_sources in the manifest.
Modes. best_effort fills what it can and records the rest as status="not_found" rows (a warning,
not a failure). strict raises EnrichmentError unless every in-scope variant resolves to a position —
the network analogue of the compiler's strict=True — and unless every authored ref agrees with
the reference sequence. The reference check is raised first, deliberately: a row contradicting the
genome is a worse diagnosis than a row the chain could not find, so it should be the error the author
sees. EnrichmentResult carries rows,
unresolved (variant_keys with no position), sources, mode, and a fully_resolved property.
A not_found row is written only where a source actually answered (S20, 0.5.4). enrich() used to
write ResolutionRow(status="not_found", source="ensembl") for a request that failed, stating in the
injected table that Ensembl was asked and does not have the rsID. It now writes nothing for that key —
the key stays unresolved, so strict still refuses and best_effort still warns, but nothing claims
a source said no — and EnrichmentResult.unreachable_rsids names them. That list is deliberately
separate from unresolved, which says a key has no position and is silent about why: a key that failed
because egress broke and one with genuinely no locus to find are the same entry there, and only one of
them is worth re-running. Empty under --offline, since nothing was asked. It warns in both modes
and strict does not escalate it — no authored edit clears a failed request (P5), and the not_covered
and VRS-coverage findings are the same class. The argument was already four lines below in the same
function, where the non-GRCh38 branch declines to write not_found for precisely this reason.
A not_found row also says something false when the source ANSWERED and the alleles did not match
(S85, 0.7). That is a fourth state, and until it was named the row was byte-identical to a genuine
absence: enrich asked, the source returned loci, the allele-aware filter rejected every one, and the
row went out as status="not_found" — this source has no record of your rsID — about an rsID the
source demonstrably has. The reported case is the common one: a paper's supplementary published against
GRCh37/hg19 spells the submitted strand, so its G/A meets GRCh38's C/T and five subjects of a
64-variant module read as unknown to Ensembl when Ensembl returns all five immediately.
EnrichmentResult.allele_mismatches carries it, as AlleleMismatch(rsid, genotype, loci, offered,
strand_flip) — the shape ref_mismatches and stale_rsids already have. It completes the family the
RM98 round started: unreachable_rsids means the request was made and failed, unconsulted_rsids that
nobody looked, unresolved that a key has no position while staying silent about why, and this one that
the asking succeeded and the answer did not match. strand_flip is the single cause this tier can
settle from the allele strings alone, and it is False for not established rather than established
otherwise — an allele that cannot be complemented withholds it.
The row itself is unchanged, deliberately, and that is not a half-fix. Adding a
VALID_RESOLUTION_STATUS member would change a wire vocabulary every reader of a published
resolution.csv shares; dropping the row would move resolution_signature, since variant_key and
rsid are RESOLUTION_FACT_FIELDS while status is provenance and is not. The row was never the
untruth — it is honestly unresolved either way — only the reason it gave, so what moved is the reason.
The run warns once, in both modes, naming the rsIDs and saying the source has them: an author who greps
the artifact for not_found is the person this finding exists for.
The per-locus warning stopped asserting a cause it had not established. hosting_verdict returns a
confident False from two arms — a substitution/MNV locus (no flank, so no spelling freedom) and an
event length the locus does not offer — and the message gave the second arm's reason for both, so a
strand-flipped SNV was reported as "The event sizes differ, which re-anchoring cannot change" about two
1 bp substitutions. compiler.resolution.contradiction_reason is now the False twin of
undecided_reason and supplies which, the same repair that function received for None's five causes.
EnrichmentResult.vrs is the MintResult this call computed (RM40, 0.5.1). It carries exactly the
two counters compile_module later stamps into manifest.compilation.vrs_alleles /
vrs_alleles_identified, plus unmintable_reasons — the grouped-by-reason breakdown that is the
actionable half, where "no refget table for build 'GRCh37'" and "needs the reference sequence"
live. It used to be logged and dropped, so a consumer reading coverage before a compile — which is
what a publish dry run is — had to re-implement the counting, and get two non-obvious rules right to
agree with the manifest a publish would produce: count per ALT slot (vrs_id is a parallel array of
alts), and treat an absent cell as len(alts) unnamed slots rather than zero, or a table where
nothing minted reports flawless coverage out of a denominator of nothing. None when mint_vrs=False
— never a coverage of zero.
Resolution is allele-aware in both directions¶
A rsID is a position/multi-allelic dbSNP tag, so one id routinely names several genuinely different records. Both directions of resolution therefore match on the allele, not just the position — and the forward direction only caught up in this round, which is why the asymmetry is worth naming.
Forward (rsid→loci). A candidate locus whose {ref} ∪ alts cannot host the module's authored
genotype is reported and left out of resolution.csv. In the committed HBB example
rs281864532 names G>GT, GT>G and GTT>G at one position, and rs613985 names records at two
positions 254 bp apart; the authored genotype says which are meant. Recording the rest would hand the
compiler a locus it can only drop — and a dropped locus makes the compile unreproducible from the
injected table, which --strict refuses. This selects, it does not repair: no authored value is
touched, and every skipped record is logged. (The compiler applies the same predicate as a safety net
for hand-authored or stale tables — resolution.hosting_verdict is shared, so the two cannot drift.)
A row that authors both halves is recorded from the reference, not copied (S104, RM251)¶
A subject carrying rsid and chrom/start needs no resolution, and until 0.7.2 that was the
whole of what happened to it: the verbatim branch wrote the authored coordinate into resolution.csv
under source="authored", with no ref, no alts, and therefore no VRS id. Yet the same run had
already looked the rsID up — the rsid↔coordinate check rides its rsIDs in the Ensembl cache batch —
and the answer was discarded. Every CPIC-drafted haplotypes.csv has this shape (allele_definitions
gives an rsID, a position and, since the gene.chr join, a chromosome), so every such module compiled
with "VRS allele identity covers 0/N" and nothing the author could do about it, and the compiler's
_verify compared the module against a photocopy of itself.
Now such a row takes the forward branch when the reference knows its rsID: the loci go through the
same allele-aware filter and the same PAR rule, and are recorded with the link's source, ref and
alts, so a VRS id mints and the compiler holds two independent values. The authored coordinate is
untouched — it is the row's identity, the compiler keeps it, and its positional fill only completes
the cells the author left empty (ref, for the CPIC shape). Where the two disagree, that is the
finding: rsid_coordinate_agreement records it here, and resolve_from_table._verify warns in
best_effort and refuses in strict, exactly as the mishap matrix already said for an authored
coordinate that contradicts the table. Three cases still write the authored row: an rsID the snapshot
does not know (no link is asked live for a pair, so there is no answer and no negative to invent), a
reference whose every locus is rejected by the allele filter (the mismatch is reported, the row keeps
its coordinate), and a run with no Ensembl snapshot at all. A table written before this change keeps
its authored rows — merge-not-clobber covers them — so delete resolution.csv or run --rederive
to record the reference's answer.
Reverse (position→rsid) back-fill is allele-aware¶
A coordinate-only variant (rsid None, coordinate authored) can have an rsid back-filled from the
reference — but the lookup is allele-aware: it matches the exact allele (chrom, start, ref, alt)
via the shared resolver._lookup_rsid_candidates, not just (chrom, start, ref). This matters because
an rsid is a position/multi-allelic-level dbSNP tag, not a per-allele one — e.g. rs33922842 at one
HBB locus tags C>A (pathogenic), C>G (benign) and C>T (uncertain). An allele-blind match would
let an un-rs'd insertion inherit a co-located SNV's rsid. So:
- 0 allele-exact candidates →
rsidstaysnull,source="authored"(the coordinate is the identity; never guess a label). - 1 → attach it (
status="resolved"). - ≥2 for the same allele (a genuine dbSNP merge) → deterministic pick in
rsid,status="ambiguous", and the full candidate list inResolutionRow.rsid_alternates— recorded, never silently chosen.
A pseudoautosomal locus is one place, recorded as the X spelling (RM32)¶
PAR1 and PAR2 are shared between X and Y, so dbSNP maps one rsID to both contigs and the expansion above
would emit two rows for one finding. enrich() keeps the X spelling and reports the Y twin it left
out; --keep-par-twin (keep_par_twin=True) records both, for a consumer whose reference is not
analysis-set masked.
This records the sources' convention, not the consumer's reference — which is what makes it this
tier's call rather than a data-agnostic violation (P2 makes the enricher the only tier permitted to hold a
source convention). Probed 2026-08-04: ClinVar holds no variant in either PAR on Y (all 677 of its Y
records lie outside them), gnomAD v4 excludes the Y PAR from its callset (region(chrom:"X",
640000-641500) serves 880 variants; the same interval on Y serves none), and the ClinGen Allele Registry
does mint a separate Y allele id but leaves that record a bare dbSNP cross-reference. Only Ensembl/dbSNP
reports both contigs. There is therefore no external place identity to adopt — the Registry minting
two CA ids is what closed that direction — and none is invented here.
Three properties worth knowing:
- The pairing is an offset, not an equality.
vrs.par_partner(format tier, stdlib) maps PAR1 at offset 0 and PAR2 at 98,813,480, from the interval table. PAR1 shares coordinates between the contigs and PAR2 does not, so a "same base on X and Y" shortcut would pass a PAR1 module and silently fail a PAR2 one. - Allele agreement is required. A twin is dropped only when the partner position carries the same
ref/alts. Partner coordinates say "same place"; they do not say "same variant", and a same-place different-allele pair is a real finding that survives whole. - The verdict is per locus.
XGruns out of PAR1 andSPRY3runs into PAR2, so a gene- or module-scoped policy would misclassify half of either.reference_examples/par_boundary/is that case, and it demonstrates the round-trip fixed point — which is exactly why this belongs here and a--parcompiler flag would be P7-illegal:resolution.csvtravels with the module, a flag would not.
Relatedly, the frozen identity now carries the allele: base.derive_variant_key keys a coordinate
variant as chrom:start:ref:alts (normalized) when an alt is present, so distinct alleles at one locus
are distinct identities. Together these make the compiler's compile → reverse → compile a full
fixpoint (artifact.digest, content_signature, and the provisional resolution_signature). See
SCHEMAS.md (derive_variant_key, ResolutionRow.rsid_alternates) and
reference_examples/pathogenic_clinvar/README.md (the ClinVar dogfood these fixes came from).
Live Ensembl — V2 GraphQL + V1 REST fallback (ensembl.py)¶
EnsemblResolver.resolve_rsid(rsid) -> (list[locus] | None, source | None) tries two backends in order:
- V2 — beta GraphQL (
_graphql_rsid,DEFAULT_GRAPHQL_ENDPOINT): the endpoint + variant-query shape leeched from ensembl-mcp,eliot.start_actionswapped for stdliblogging. - V1 — legacy REST (
_rest_rsid,DEFAULT_REST_ENDPOINT=rest.ensembl.org/variation/{species}/{rsid}): newly written for this repo (it did not exist in ensembl-mcp); parses GRCh38mappingsinto loci.
resolve_rsid falls back V2→V1 on a status in _FALLBACK_STATUS = {500, 502, 503, 504} (or a
GraphQL/transport error, or an empty V2 result). tenacity (@retry, exponential jitter, 3 attempts,
on httpx.TransportError/timeout) wraps each backend independently, so an endpoint is retried before
the fallback triggers.
Three outcomes, because two of them used to be one (S20, 0.5.4). A non-empty list is an answer; []
is also an answer — Ensembl was reached and has no GRCh38 locus for this rsID — and None means
Ensembl could not be asked at all, so its answer is unchecked rather than empty. Fusing the last two
into ([], None) made a failed request read as a definite negative, and loci: [] beside "live
Ensembl has no GRCh38 locus for it either" is exactly the fingerprint of a fabricated identifier: a
consumer checking which rsIDs in a machine-written document were real put two published variants
(rs6567160, a long-standing MC4R BMI locus, and rs13010010) in the fabricated pile, and caught it
only because five-of-seven succeeding looked more like flaky egress than a 30%-honest document. This is
the tri-state rule the rest of the tree keeps — an unreachable source reports unknown, never the
negative.
Two boundaries. A 4xx is an answer, not a failure: Ensembl 400s on rsIDs it cannot resolve
(rs3216883, which dbSNP reports as merged), so only a 5xx, a transport error or a timeout returns
None. And an answered-empty carries its source, so hint.checked records ensembl-rest when
Ensembl was reached and said nothing — the old code's only trace of that case was a missing element in
a set, which is unreadable in practice.
A one-sided indel is anchored before it becomes a locus, and a fourth outcome withholds it (RM268,
2026-09-27). REST spells an insertion -/C with start = end + 1 and a one-sided deletion
AGTAAG/- over [start, end]; _loci_from_rest used to copy both through, so rs8176719 resolved
to 9:133257522 ref='-'. It now prefixes every allele with the GRCh38 base at start - 1
(clingen_allele.anchor_indel, reading through the private EnsemblResolver._read_base, by default
a lazy SequenceProxy), giving 9:133257521 T>TC. An allele string with no - is already
anchored and reads nothing. When the base cannot be read the locus is withheld, and an answer whose
every locus was withheld returns (None, "ensembl-rest"): unchecked like a failed request, but
enrich names it in its own warning (a structured field is RM272, minor), and lookup gives it its
own finding, because the request did not fail. The GraphQL leg's convention is unprobed, so a one-sided
node there answers [] and the rsID goes to REST.
A non-base allele is not a locus, at either rung (RM271, 2026-09-27). REST serves the literals
dbSNP_novariation and dbSNP_variant as an allele_string; the VCF-dump snapshot serves an empty
alt, <.>, IUPAC placeholders (<R>, <Y>), N runs and N-masked chrY refs. One private predicate,
ensembl._placeable_alleles, runs at every read of both (REST, GraphQL, and the four snapshot reads in
resolver): a ref outside ^[ACGT]+$ withholds the locus, an alt outside it is dropped, and a locus
with no alt left is withheld. REST with nothing left answers the ordinary [] (it is permanent, so no
re-run helps), and the snapshot names such rsIDs in one warning instead of rsid_unresolved.
checked carries labels only, and snapshots is where a path lives (S93, RM205). The set used to
be mixed — ensembl-rest for the live leg beside an absolute path for a snapshot, and the
unreadable-snapshot finding interpolated the same path — so a host serving lookup_variant over HTTP
mapped every known path back to a lane name, inside prose too, an audit that has to be repeated on every
field added. Now checked holds the lane's name (ensembl, clinvar) or the live source, the finding
reads ensembl snapshot unreadable: …, and VariantHint.snapshots maps each label to the path it
resolved to for every snapshot the lookup opened or tried to (ensembl, clinvar, pubmind) — the
one field a host drops to keep its layout private. duckdb's own first line may still name a file; that
is upstream's sentence and stays as evidence.
Honest caveat (bare rsID). The beta variation GraphQL wants a composite
region:pos:rsidid, so a bare rsID typically won't resolve through V2 and falls through to V1 REST, which does the real work — mirroring ensembl-mcp itself. Today V1 is the workhorse; V2 is wired, retried, and first in line but mostly hands off. Adding a REST→composite-id step (a follow-up) would make V2 a genuine first responder. Endpoints and the human GRCh38 genome id are configurable viaEnsemblSettings.
The caches¶
Adding a lane? CACHE_SURFACE.md is the checklist, and every line in it names the test that asserts it — written after a lane shipped with a publish_repo no command could reach.
Parquet snapshots, one base directory, one rule: locate, never download — except where you ask.
(The count used to be stated here and went stale twice while the table below grew; a number in prose
is a registry nothing iterates. Since RM176 the roster is caches.CACHE_LANES — a walked registry
with one entry per lane, compared by a test against the *_build modules on disk, so this table is a
rendering of it rather than a third copy. Three lanes were missing from the previous list and nothing
could have noticed: cache status reported nine caches on a machine that has twelve.)
The Override column below is CacheLane.env_var rendered (S89, RM184), and the constants behind it
live in locations as <LANE>_CACHE_VAR beside CACHE_BASE_VAR, the one variable no lane owns. A
consumer generating a .env.template, clearing its test environment, or auditing which caches were
provisioned by variable derives the whole set as {lane.env_var for lane in CACHE_LANES} |
{CACHE_BASE_VAR} rather than keeping the fourteen names by hand.
The status half is caches.lane_status() (S91, RM204), one LaneStatus per lane in registry
order, and cache status renders it — a consumer serving the same answer over HTTP reads the function
rather than re-writing the loop, which is two projections of one registry. Its state is
three-valued: present, absent, and occupied — the place the lane looks exists, is non-empty
and holds no snapshot (a build that failed after its downloads, a payload deleted beside its
release.json, a stray .part, a foreign parquet). That third state used to print as absent,
which sends an operator to run a pull that prepare is going to refuse, since provisioning never
deletes. looked_in says which directory the verdict is about (the lane's env_var if set, else the
default), release is what the snapshot names, and release_unreadable is the present-and-unreadable
release.json — a provenance failure, not a data failure.
A lane declares its size and a present one measures it (S97, RM229). CacheLane.approx_mb is an
order of magnitude in whole megabytes, measured on a provisioned box on 2026-09-11 and rounded up (1
means at most a megabyte; Ensembl is ~15 000, the AVI lane ~30 000), so a first-run offer can decide
does this fit, is it worth asking without du-ing a box that already has it; None means nobody has
measured, which a caller reports as unknown rather than guessing. A test re-measures every lane
present on the machine it runs on and refuses a declared number more than an order of magnitude off,
which is what keeps the field from becoming the hand-kept table it replaces. LaneStatus.size_bytes
is the measured size of a lane you already hold, and cache status prints it. A derived lane's
price is its parents': provisioning_closure(lane) is the lane plus its transitive parents in
registry order — mitomap_miss is under a megabyte and its closure is a ClinVar download.
Every live source this tier reaches has (or can have) a local copy, and the whole reason is in the rate
table above: a shared IP shares one budget. An author on their own machine can go live for everything;
a host cannot, and for the three licence-gated sources it should not (see On a host, or in a
service below). Pre-provisioning is therefore a deployment step, not an optimization.
| Cache | Subdir | Override | ensure_* |
Published at | Serves |
|---|---|---|---|---|---|
| Ensembl | ensembl_variations/ |
$JUST_DNA_ENSEMBL_CACHE |
ensure_snapshot |
just-dna-seq/ensembl_variations |
rsID → coordinate (enrich) |
| ClinVar | clinvar/ |
$JUST_DNA_CLINVAR_CACHE |
ensure_clinvar_snapshot |
just-dna-seq/clinvar |
records + citations/ (enrich, draft-panel) |
| gnomAD constraint | gnomad_constraint/ |
$JUST_DNA_GNOMAD_CONSTRAINT_CACHE |
ensure_constraint_snapshot |
just-dna-seq/gnomad_constraint |
v4.1 gene constraint (gene-metrics) |
| ClinPGx 🔒 | clinpgx/ |
$JUST_DNA_CLINPGX_CACHE |
ensure_clinpgx_snapshot |
just-dna-seq/clinpgx |
clinical annotations (clinpgx check) |
| CPIC 🔒 | cpic/ |
$JUST_DNA_CPIC_CACHE |
ensure_cpic_snapshot |
just-dna-seq/cpic |
alleles / diplotypes / recommendations (pgx, draft) |
| PharmVar 🔒 | pharmvar/ |
$JUST_DNA_PHARMVAR_CACHE |
none, by design | never published | star alleles (pgx) |
| PubMind ❓ | pubmind/ |
$JUST_DNA_PUBMIND_CACHE |
none, by design | never published | literature-derived verdicts (RM134) |
| CIViC ✅ | civic/ |
$JUST_DNA_CIVIC_CACHE |
ensure_civic_snapshot |
just-dna-seq/civic — CC0, repo not created yet |
curated cancer interpretations, direction axis only (RM152); --submitted widens the basis from the release's own VCF (RM169) |
| MANE ❓ | mane/ |
$JUST_DNA_MANE_CACHE |
none, by design | never published | the transcript numbering frame — summary, changed Select accessions, excluded genes (RM168) |
| ClinPGx drug labels 🔒 | drug_labels/ |
$JUST_DNA_DRUG_LABELS_CACHE |
ensure_drug_labels_snapshot |
just-dna-seq/clinpgx_drug_labels — repo not created yet |
what five regulators say about a gene/drug pair (clinpgx check-labels) |
| STRchive ✅ | strchive/ |
$JUST_DNA_STRCHIVE_CACHE |
ensure_strchive_snapshot |
just-dna-seq/strchive — MIT, repo not created yet |
repeat-locus bands (check-repeat-bands, draft-repeats) |
| ACMG SF ❓ | acmg_sf/ |
$JUST_DNA_ACMG_CACHE |
none, by design | never published | the secondary-findings list (check-acmg) |
| MITOMAP ✅ | mitomap/ |
$JUST_DNA_MITOMAP_CACHE |
ensure_mitomap_snapshot |
just-dna-seq/mitomap — CC BY 3.0, repo not created yet |
curated mtDNA variants, both mmutation and rtmutation (RM171) |
| MITOMAP miss ⛓ | mitomap_miss/ |
$JUST_DNA_MITOMAP_MISS_CACHE |
none — derived | not published, deliberately | what MITOMAP publishes and the ClinVar cache does not (draft-panel --source mitomap-miss) |
| AlphaGenome AVI 🔑 | alphagenome_avi/ |
$JUST_DNA_ALPHAGENOME_AVI_CACHE |
ensure_alphagenome_avi_snapshot |
just-dna-seq/alphagenome_avi — Permissive Use (RM195), repo not created yet |
variant-impact scores for 8.8 B SNVs, Int32×10⁵ plus a 466 KB knot table (alphagenome build --input, RM191/RM198) |
🔒 = licence-gated (commercial_use=False). ❓ = terms unestablished (commercial_use=None), which
is a different state and not a weaker one: unknown is not permissive. ⛓ = derived: this lane has
parents rather than a download. 🔑 = the acquisition itself is gated, which is a fourth thing again
(RM191): the AVI artifact is 88.5 GB behind a sign-in whose eligibility clause bars classes of holder
— "the AlphaGenome Services aren't available for any commercial entity, even if conducting
non-commercial work" — so there is nothing for a rebuild adapter to fetch and the operator supplies
the file under their own acceptance (@acquisition-gate-is-not-a-read-gate). It is the second lane
with no rebuild adapter and the first with a builder anyway: having a builder module and being
rebuildable stopped being the same property here. Since RM198 it is also pullable without being
buildable — this tier may not fetch the gated source, but the 34 GB re-encoding is Permissive-Use
output (RM195) and redistribution=True records the reading that publishing it openly falls inside
prohibition 1's carve-out, so an operator who cannot download the artifact can still provision the
lane. Its snapshot is the first that is incomplete without a root-level file: PHRED is not
stored, so avi_knots.parquet travels beside data/ and locations.SNAPSHOT_ROOT_FILENAMES is what
makes the publisher carry it. The three 🔒 rows are RM38, new in 0.5.1; the PubMind
❓ row is RM134 and the MANE one is RM168; the two MITOMAP rows are RM171.
The derived lane, and the parents field (RM171)¶
mitomap_miss is the registry's first lane that fetches nothing. Its acquire stage is both parents
being on disk — the MITOMAP snapshot and the ClinVar one — and its build is an exact
(start, ref, alt) join on chrMT, upper-cased on both sides, with no position-level fallback: a
hit at the same position on a different allele is a different allele, and collapsing onto it would
hide a real increment or invent one where the two sources anchor an indel differently.
CacheLane.parents is a tuple of lane names, empty for every lane that acquires its own bytes. It is a
correctness fact — which digests get recorded, what rebuild refuses on — and it is also a cost
fact, which its first reader did not expect (S97): the increment is under a megabyte built, and on a
blank box its honest price is its parents, a ClinVar download among them. provisioning_closure
walks it transitively for exactly that sum. Three things read it and only one is the join:
rebuild_laneguards on it. A child whose parents are not on disk isbuilt=Nonenaming which one, with the command that would provision it. The two wrong answers are both silent: aFalsefiles another lane's absence as this lane breaking, and an empty increment is the strongest possible claim about MITOMAP — it publishes nothing ClinVar lacks — derived from a comparison that never ran.- The registry order is load-bearing, because
cache preparewalks it top to bottom. A parent must precede its child, and a test asserts it rather than a comment. cache rebuildhands the child the parents it just cut, underout/<parent>/, so a fresh set is internally consistent.--only mitomap_missnames none and falls back to the resolvers, which is the right answer for a run with no rebuild behind it.
The child pins both parents in its release.json, which is what makes a ClinVar rebuild without a
child rebuild detectable rather than silent — mitomap_miss_build.stale_parents re-reads both and
names the one that moved. A parent that is gone is deliberately not reported as one that moved:
"provision the parent" and "rebuild the child" are different instructions.
It is not published, and the reason is a fourth one. Both parents are redistributable — ClinVar is public domain, MITOMAP is CC BY 3.0 — so nothing in the licensing bars it. What bars it is what the artifact is: a pulled copy would carry a currency check its holder cannot run, having neither parent. Rebuilding it locally is seconds against caches the machine already has, and it cannot be stale by construction.
The two ❓ rows are unestablished for different reasons, and the third column differs with them.
PubMind's bytes carry no stated terms at all; MANE's carry a policy — NCBI places no restrictions on
use or distribution and, in the same paragraph, declines to grant unrestricted permission. Neither is
permission, so both caches are operator-built, but PharmVar's and PubMind's absent ensure_* records a
refusal, CIViC's records a gap, and MANE's records an unestablished question nobody has asked NCBI.
One base, so a single just-dna-lite deployment's cache serves all of them. Each subdir sits under
$JUST_DNA_PIPELINES_CACHE_DIR, or platformdirs' user cache for just-dna-pipelines
(~/.cache/just-dna-pipelines on Linux) when that is unset. Precedence per cache is explicit argument
→ its own $JUST_DNA_*_CACHE → the base, and a .env beside the working directory is loaded
automatically (locations.load_env, walking up from CWD). Every resolver returns None rather than
guessing when nothing is there.
That load writes into os.environ, and it is a library path rather than a CLI one — pass
load_dotenv_file=False if that is not wanted. load_dotenv sets every variable the file holds, not
only the cache ones, so a process that merely asks where the ClinVar cache is inherits whatever
credentials sit in the nearest .env above its working directory. override=False means a variable
already present is kept — which reads as the safe direction and has one sharp edge: deleting a
variable is what lets the file supply it, so a test isolating itself with del os.environ[...] (or
monkeypatch.delenv) is un-isolated by the next resolve. Set it empty instead; that is the same rule
the enricher's own tests follow. Every resolver and every default_*_cache_dir takes
load_dotenv_file, and since 0.6.3 (in the tree, uncut) passing False really does reach the load —
before that it reached none of the six, because the default directory is computed as an argument and
loaded the file on its own way in (S39). The credential paths export nothing since RM301 (S124).
Every client, the PharmVar guard, the AlphaGenome key, the HuggingFace token and the retry floor read
their one variable through locations.env_value: the process environment first, then the nearest
.env, with load_env's precedence and without writing into os.environ. So constructing a client no
longer hands its host the rest of the file. Only the cache resolvers above still export, and they take
the switch. RM102 carried the whole question and closed on 2026-08-21 as a decision not to act. It
was reopened as RM301 when a host's own configuration record turned out to be the boundary its
trigger named.
Inside a cache the layout is fixed, because four parties have to agree on it — builder writes, publisher uploads, provisioner fetches, reader queries — and every past disagreement was silent:
<base>/<subdir>/
data/*.parquet # the records; the readers glob exactly this
citations/*.parquet # optional sidecar, a SIBLING of data/ (ClinVar only)
release.json # which release this is — what reference_sha256 pins against (RM4)
LICENSE.txt # the terms, for a snapshot that ships its own (ClinPGx and its labels)
Two caches hold no parquet at all, and that is a second layout rather than a defect in this one:
ACMG's list is acmg_sf.csv and STRchive's catalogue is STRchive-loci.json, each at the snapshot
root beside release.json. Both are read whole rather than queried, and a parquet for eighty-one rows
would be ceremony. This is why they had no resolver until RM176 — every predicate in locations tested
for data/*.parquet — and why plan_reference_snapshot takes a payload filename from its caller
rather than assuming one shape for every lane.
Pre-caching the published snapshots from HuggingFace¶
just-dna-enricher cache status # what is present, where, which release
just-dna-enricher cache pull # every published lane with no `--use` gate
just-dna-enricher cache pull --use non-commercial # …and the gated ones you may hold
just-dna-enricher cache pull --only clinvar --only cpic --use non-commercial
cache pull is re-runnable and cheap: a complete cache is trusted without touching the network, and
only an empty or corrupt one refetches. One snapshot failing does not sink the rest — each reports its
own line, and the command exits 1 if any failed.
Four things worth knowing before you run it on a server:
- Set
HF_TOKEN. Anonymous traffic shares a per-IP pool (500 API / 3,000 resolver calls per 5-minute window); a free token roughly doubles it. The usual symptom of not having one is anensure_*that looks hung — the client is sleeping on a 429, not stuck. A.envbeside the working directory is enough: both provisioners and the publisher callload_env()beforeget_token(), which otherwise reads only the real environment and~/.cache/huggingface/token. --useis required for the gated pair, and it is not ceremony. ClinPGx and CPIC forbid sale, and under a data-usage policy the terms are accepted when the data is taken — so downloading is the act being gated.unstatedskips them with a reason,commercialrefuses,non-commercialproceeds. Same three states as everywhere else; the tool will not assert a purpose for you.- Five of the repos have to exist first, and a missing one is not an error.
just-dna-seq/cpicandjust-dna-seq/clinpgxare new with 0.5.1;civic,strchiveandclinpgx_drug_labelswith RM176 — none of the five has been published yet. Each provisioner raisesSnapshotNotPublishedwhen the repo or itsdata/does not exist, andcache pullprints that in yellow without counting it as a failure: nobody-published is the same third state as nobody-asked, and the command that provisions a deployment must not exit 1 on a fresh machine because a snapshot has never been uploaded. A download that breaks is still a failure and still exits 1. Build and publish them once (below), or point at a locally built directory. - A published dataset accumulates. Each
ensure_*fetches only the files its own snapshot is made of, because the ClinVar repo still carries a 159 MBclinvar.parquetfrom the single-file era whose columns are raw VCF INFO fields. The readers globdata/*.parquet, so one foreign file puts two schemas under one DuckDB relation and every query dies onReferenced column "clin_sig" not found.
If you prefer the Hub CLI, the layout is plain and the same files are all there is:
hf download just-dna-seq/clinvar --repo-type dataset \
--include 'data/clinvar-*.parquet' 'citations/*.parquet' 'release.json' \
--local-dir "$JUST_DNA_PIPELINES_CACHE_DIR/clinvar"
Note the --include: data/*.parquet alone would drag in the stale flat file above. In Python it is
from just_dna_enricher.download import ensure_clinvar_snapshot; ensure_clinvar_snapshot(), which does
the filtering, the footer check and the atomic rename for you.
Building the ones that are not (fully) published¶
(Another count that went stale as the table grew — CIViC's builder, documented in its own section below, was already a fifth. The list here is the roster.)
# CPIC — open and unauthenticated, so this is about a host's shared budget, not access.
# Every builder writes data/repro/<lane>/ unless --out names somewhere else (RM177).
just-dna-enricher cpic build --use non-commercial # 132 genes, ~120k rows, ~256 KB
just-dna-enricher cpic publish data/repro/cpic --repo <org>/cpic # optional; redistribution is granted
# ClinPGx — the bulk archive; its LICENSE.txt is extracted and travels with the parquet.
just-dna-enricher clinpgx build --use non-commercial
just-dna-enricher clinpgx publish data/repro/clinpgx --repo <org>/clinpgx
# PharmVar — needs YOUR key, and there is no publish command.
PHARMVAR_API_KEY=… just-dna-enricher pharmvar build --use non-commercial
# PubMind — no key and no `--use` flag; `pubmind publish` exists and refuses (see below).
just-dna-enricher pubmind build --download # or --table hg38_pubmind_db.txt.gz
# MANE — no key, no `--use`, no publish command. Discovers the newest version from current/ and
# then pins it to release_<version>/ before fetching anything.
just-dna-enricher mane build --download # or --release 1.5 to pin by hand
# The regulator drug labels — a SECOND ClinPGx archive with its own cadence, so its own cache and
# its own repo. Same CC BY-SA terms as the annotation lane, so publishable on the same grounds.
just-dna-enricher clinpgx build-labels --use non-commercial # → data/repro/drug_labels
just-dna-enricher clinpgx publish-labels data/repro/drug_labels
# STRchive — MIT, so publishable. Pin a release: the default branch moves, and only a pinned build
# gets a `dataset` label the comparison can name.
just-dna-enricher strchive build --release v2.26.0
just-dna-enricher strchive publish data/repro/strchive
# ACMG SF — nothing is fetched. The workbook is ACMG/Elsevier supplementary material, so you supply
# your own copy; there is no publish command, because nothing grants redistribution of those bytes.
just-dna-enricher acmg build ./acmg_sf_v3.3.xlsx # → data/repro/acmg_sf
cache prepare — the whole set, by whichever route each lane has (RM176)¶
just-dna-enricher cache prepare --use non-commercial # the deployment's one provisioning step
just-dna-enricher cache prepare --only mane --only acmg # or a subset
cache pull fetches the published snapshots and stops, and five lanes are not published for
recorded reasons — PharmVar's personal key, PubMind's absent terms, NCBI's policy over MANE, ACMG's
supplementary material, and mitomap_miss, which is derived from two parents on the pulling machine
and would only pin somebody else's ClinVar if it travelled. A machine that only pulled is therefore
five caches short, and the checks that
read them skip themselves with no_reference. prepare runs each lane by the route it has: pull if
published, build if not.
- The route is a property of the lane, never a flag. A published lane pulls, because building it
would spend an operator's bandwidth re-deriving bytes somebody already made; an unpublished one
builds, because that is the only route there will ever be.
PrepareOutcome.routerecords which answered —present,pulled,built— because a deployment auditing its own caches has to tell a snapshot it fetched from one it made, andrelease.jsonnames the release but not the route. - A present cache is left alone, exactly as
cache pullleaves one alone, so this is idempotent. Re-cutting a snapshot that already exists iscache rebuild, which writes somewhere else on purpose — a build straight into a live cache is visible half-done to anything reading it, and unlike a truncated download there is no footer check to catch it because the file is real. - A built lane is staged and moved. It builds into
<cache>.incomingand renames across only once the build finished; a failure removes the staging directory rather than leaving a half snapshot a latercache statuswould call present. - In Python:
caches.prepare_caches(...), which the command calls.caches.rebuild_caches(...)is its sibling. Both return one outcome per lane in registry order, so a caller can zip againstCACHE_LANESwithout matching on names.
Every bulk download goes through one body, and a failed one is a failed lane (RM187). Each
builder streams its source's file through net.stream_to_file, which is atomic (.part, renamed only
on success), retries a transport failure — a connection cut mid-body, the failure that produced
this item — and translates everything else into the lane's own error type. A status error is not
retried, because a 404 from a mistyped release tag is the same 404 four times over.
That last part is what a caller notices. Before RM187 four of the eleven downloads raised httpx's
own exception, and each lane adapter catches its own builder's type — so a truncated download was a
traceback rather than a built=False, and it escaped rebuild_lane too, aborting every lane after it.
If you catch httpx.HTTPError around one of these builders, catch the lane's error instead; the
httpx exception is kept as __cause__. A retry re-fetches from byte zero: there is no resume, so a
190 MB download that dies at 180 MB costs the full 190 MB again.
Two refusals a payload-less directory earns (2026-09-09). resolve() answering None means the
lane's directory holds no payload, not that it is absent. A directory that exists with none — a
build that failed after its downloads, a payload deleted by hand beside its release.json — is
refused on the build route before any build is spent, naming the directory and cache prune,
because provisioning never deletes and the rename onto it would have raised Directory not empty
out of the whole command. And a derived lane judges each parent by that same resolver: an empty
out/<parent>/ left by a download cut mid-body is a missing parent (the child could not run),
never a present one whose join then fails and files the child as FAILED. prepare_caches isolates
each lane the way cache pull always has, so one lane raising is that lane's FAILED line and the
rest still run and print.
One endpoint over every builder (cache rebuild, RM176)¶
Thirteen builders today — eleven when RM176 shipped, plus mitomap and mitomap_miss from RM171 the
next day — three stages each — acquire, build, publish — and until RM176 the only way to run them all
was one command per lane, each with its own flag shape. cache rebuild is the single endpoint, and
it calls the same download_*/build_* functions the per-lane commands call, so there is one
conversion algorithm with two callers rather than two that have to agree. The per-lane commands stay:
they offer the local-file inputs an operator holds, which a rebuild pass by definition does not.
just-dna-enricher cache rebuild --use non-commercial # → data/caches/<lane>/
just-dna-enricher cache rebuild --only strchive --only mane --pin mane=1.5
just-dna-enricher cache rebuild --only acmg --source acmg=./acmg_sf_v3.3.xlsx
just-dna-enricher cache rebuild --publish --dry-run # rehearse the uploads
scripts/rebuild-caches.sh data/caches # the driver, all lanes
just-dna-enricher cache rebuild --out /srv/just-dna/caches-$(date +%F) # a deployment's own path
Three things about it are worth reading before a deployment runs it nightly.
- Every lane builds into
<base>/<lane>/, never in place over a resolved cache. A rebuild takes minutes, and a short parquet still has aPAR1footer — so anenrichreading a half-written snapshot mid-flight sees a real but incomplete table and no resolver can catch it. Moving the result into the live caches is a separate, deliberate step. - Every
--outdefault is underdata/, and none of them is written by hand (RM177).cache rebuildtakesdata/caches/; every builder takesdata/repro/<lane>/fromlocations.repro_out. That is not cosmetic: run from a checkout,--out ./cpicdrops an untracked snapshot directory in the repository root, which is exactly the statecivic reproduceneeded its own.gitignoreline to paper over. The rule was prose for a release and nine defaults drifted past it —civic,clinvar,pubmind,gnomad_constraint,mane,strchive,acmg_sf,mitomap,mitomap_miss— with four more builders requiring--outand no default at all, which is how the docs came to tell an operator to write--out ./clinpgxinto the root. It is derived now and an AST walk over the CLI asserts it, so the tenth builder inherits the rule instead of repeating the defect. A deployment passes its own absolute path and none of this applies. - The outcome is three-valued, and not run is not a failure. ACMG needs a workbook that is
Elsevier supplementary material, PharmVar a personal key, CIViC a release date to pin, Ensembl is
built by just-dna-pipelines, and a derived lane whose parents are not on disk (
mitomap_miss, RM171) names the parent it lacks rather than reporting an empty result. Each prints its own reason — taken from the registry field, not composed here — and the exit code counts only real failures, so a nightly run does not alarm on four lanes behaving exactly as designed. - Credentials come from the
.envtoo, and until RM176's follow-up two of them did not.$PHARMVAR_API_KEYand$HF_TOKENare read at the point they are used (@credential-where-read), throughenv_valuesince RM301, so a workspace that keeps them in a.envbeside the working directory needs nothing exported and the rest of the file is not exported either. An exported variable still wins, as it did underload_env'soverride=False. The two that were missing it failed in opposite directions: the PharmVar guard reados.environin front of a builder that does load the file, so the lane reported "no key" and never built on the machine most likely to have one;_hf_apirefused a publish outright. export FOO=is stronger thanunset FOO, and both messages now say so.override=Falsekeeps a variable that is present, and an empty string is present — so a shell that ran a snippet whose placeholder was edited out reports no key for the rest of the session while the.envholds a working one. Absent and exported-empty are named separately, with different remedies: add it, orunsetit. The same trap the tier's own tests exploit deliberately, from the other side.- The ACMG lane uses the checkout's own workbook when nobody names one.
assets/acmg_sf_v*.xlsxtravels with this repository, so a rebuild run from a checkout needs no--source— it is found by walking up from the working directory and, for an editable install run elsewhere, relative to the package. A pinned filename is deliberately not used: the list is versioned, v3.2 became v3.3 in June 2025, and a constant naming one version stops finding the asset the day the next lands — silently, in the direction that falls back to scraping NCBI's stale page. Two matching workbooks are reported rather than ordered, because "the highest version" is an ordering nobody defined and filename sort is not it.assets/is not in the wheel, so apip installstill supplies its own copy and gets the operator-supplied message: the workbook is ACMG/Elsevier supplementary material, and shipping it in a published package is exactly the redistribution question this lane's registry entry records as unestablished. - A
--sourcepath is checked before the run starts. Typer'sexists=Truecannot reach a value embedded in alane=valuestring, so a mistyped one used to travel to the lane's builder and surface as a bare[Errno 2]— andacmgis last in the registry, so that was after everything else had downloaded.~is expanded, because no shell expands it inside an assignment. --source lane=pathis the offline off-switch for clinvar, constraint, clinpgx, drug_labels, pubmind, strchive and mitomap (apg_dumpyou already hold; the snapshot is then honestly unlabelled), and the only route for acmg.maneandcivicrefuse it: each takes three input files, two of three is not a build for either, and a flag that can supply one would be a flag that cannot do its job.
The CIViC adapter fetches the three TSVs and no VCF, which is civic build's own default. RM169
made --submitted opt-in precisely because the release VCF widens the status basis — it admits
submitted-but-not-accepted evidence — so a rebuild that fetched it unconditionally would build a
different snapshot from the same release than the per-lane command does. Two callers, one release, two
artifacts is the fork this endpoint exists to prevent; widening the basis is a curation decision, and
civic build --release <date> --submitted is where it is taken. The rebuild's outcome line prints
status_basis so the two are distinguishable at a glance.
PharmVar's not run is decided before the request, not from the failure. The service returns an
identical 401 for an absent, a malformed and an unrecognised key, and PharmVarError is flat, so an
adapter reading the message would be parsing prose. $PHARMVAR_API_KEY unset is the designed third
state — the key is personal and non-transferable, so a machine without one is a machine PharmVar never
meant to serve — while a key that is configured and then fails is asked-and-failed and counts.
Then point at them, or move them under the base directory so the default resolvers find them:
export JUST_DNA_CPIC_CACHE=data/caches/cpic
export JUST_DNA_PHARMVAR_CACHE=data/caches/pharmvar
just-dna-enricher pgx spec/ --offline --use non-commercial # zero egress, both legs answered
Why PharmVar has no publish command, and will not get one. Its bulk data is pulled under a key its
terms §2 make personal and non-transferable, and no axis SourceTerms records covers passing that
on — redistribution=True describes the CC BY-SA grant over the content, not a clause about the
account. An unestablished permission is never a permission, the same None ≠ False rule that
governs share_alike and commercial_use. So the snapshot is operator-built and inject-only, its
release.json says so ("redistributable": false), and the build command prints the same warning.
Why PubMind has no publish command either, and why its reason is a different one (RM134). PharmVar's
bytes arrive under terms that bar passing them on; PubMind's arrive under no stated terms at all. Two
things are licensed and only one clearly: CHOP's LICENSE.md covers the software (academic,
non-commercial, tech-transfer for anything else), the paper is CC BY-NC-ND 4.0, and the
ANNOVAR-redistributed coordinate table — the only per-variant channel there is — publishes nothing. Under
the house rule that is unknown, not permissive, so PUBMIND_TERMS records None on every axis and the
snapshot is operator-built and inject-only.
pubmind publish therefore exists in order to refuse, rather than being absent: a missing command
reads as an oversight somebody will helpfully add. It exits non-zero and names the reason, the PharmVar
precedent, and what would lift it — an answer in writing from WGLab and CHOP's Office of Technology
Transfer, which is the only thing that can.
Unknown terms warn; they never gate, and that is why pubmind build has no --use flag.
taints_commercial_use requires commercial_use is False, so a module carrying PubMind values compiles,
lands pubmind in manifest.sources.unknown_terms_sources, and drives the module-wide verdict to None
— undetermined, never permitted. check_declared_use(PUBMIND_TERMS, …) returns a skip reason for every
declaration, so wiring the flag in the way pharmvar build does would refuse every build; a flag feeding
a gate that never gates is a flag that does nothing. What the terms would gate is publishing a module
carrying those bytes, which is RM27's undesigned redistribution axis rather than this source's problem.
Snapshot-first, live second, --offline first-only¶
Every pass follows one of two shapes, and which one it follows depends on whether a live route exists at all:
| Pass | With a snapshot | Without one, online | Without one, --offline |
|---|---|---|---|
enrich |
cache | provision, else live Ensembl/ClinVar/gnomAD | cache only |
gene-metrics |
snapshot (v4.1) | provision, else live API (v2.1.1 — dataset says which) |
snapshot only |
pgx |
snapshot | live PharmVar/CPIC | skipped, with a reason |
draft |
snapshot | live CPIC | skipped, with a reason |
clinpgx check |
snapshot | provision (no live route exists — the API was retired) | skipped, with a reason |
dosage, literature, frequencies |
— | live | no-op + warning (no snapshot exists) |
The asymmetry in the middle column is deliberate. clinpgx provisions automatically because there is
no live fallback to degrade to; pgx and draft fall back to live because there is one, and pulling a
whole database to answer one gene would be the wrong default for an author on a laptop. Neither adds a
second flag — --offline is the switch, and an explicit --*-cache / --snapshot path is the
inject-only escape hatch, never second-guessed.
Which route actually answered is recorded, not implied: PgxResult.routes reports snapshot or
live per source, and a snapshot stamps its own release into SourceRow.dataset (cpic_snapshot_<12
hex>), exactly as the two gnomAD constraint routes already distinguish themselves. A consumer must be
able to tell a pinned file from a live API, because the two can differ by a release.
A pass that could run neither way is a third state, never a silent pass: PgxResult.skipped_offline
and ClinGenResult.skipped_offline carry the reason, distinct both from "ran and found nothing" and
from a failure.
The one cache that is not in the table¶
The ACMG snapshot used to be listed here, and the reason given was that it is "a single small CSV an author points at, not a shared reference". That reasoning was wrong in the way a design note is wrong when nobody re-checks it: an author points at it because nothing could find it otherwise, and the consequence was that the flagless path fell through to scraping NCBI's page — which serves v3.2 while ACMG published v3.3 in June 2025, so a correctly authored row came back reported as wrong. It has a resolver and a roster row since RM176.
- No response cache for the live clients. NCBI, PharmVar's live path, gnomAD GraphQL, Crossref,
Europe PMC, OLS4, HGNC and live Ensembl are paced only. Persistence is the authored sidecars
(
resolution.csv,frequencies.csv, …) — delete a sidecar to force a refetch, becauseenrich()treats an existing one as authoritative and merges into it rather than clobbering it.
When something is wrong with a cache¶
| Symptom | Cause | Fix |
|---|---|---|
Referenced column "clin_sig" not found |
a foreign parquet in data/ (stale layout, or an old builder) |
cache status, then move the file aside and rebuild |
| "present but not queryable" | as above, or a truncated download | remove the file; cache pull refetches it |
cache status says occupied |
the directory the lane looks in is non-empty and holds no snapshot — a failed build, a stray .part, a foreign file |
move it aside (or cache prune --only <lane> for a retired file); prepare refuses to build over it, by design |
ensure_* appears to hang |
anonymous 429 backoff | set HF_TOKEN |
| a pass says a source was "skipped: --offline and no built snapshot" | exactly what it says | cache pull, or <source> build |
repository not found for cpic/clinpgx |
nobody has published that snapshot yet | build it locally and point $JUST_DNA_*_CACHE at it |
The drafting scaffold — what every provider shares, and what it is allowed to differ on (RM228)¶
Seven *_draft.py providers turn a snapshot into authored rows, and they grew one at a time. By 0.7
each had its own copy of the same four decisions, and the copies had drifted: two providers recorded a
release label and never withdrew a stale one, one imported another's private _MATCH_ON across
modules, and one consumed pydantic's rendered error message as an API. drafting.py is the mechanism
those seven were approximating.
The split the scaffold enforces. A provider's skip rule is two rules that were being written as one hand-kept list:
| derived or declared | why | |
|---|---|---|
| the model's requirement | derived, always | VariantRow's rule is one fact; a provider restating it is a copy that drifts. pgx_draft once restated "no rsID and no position" where HaplotypeRow wants rsID or chrom+start, and draft --gene CYP2C9 died on an unhandled pydantic error |
| the source's precondition | declared, with a reason | a true fact about that snapshot — ClinPGx carries no coordinate, MITOMAP publishes no rsIDs. Legitimate, provider-specific, and useless to a reader without the reason |
Mashed together nobody can tell them apart, which is the state mitomap_draft's clin_sig clause was
in: it reads like an identity requirement, it is not one, and the reason it is nonetheless correct (a
rated_miss carries one by construction, and the guard buys a named refusal instead of a raw
ValidationError) sat three lines below in a comment. SourcePrecondition.reason is a field, so the
two halves cannot merge again.
The identity verdict comes from constructing the model, not from authoring_requirements. That
was measured rather than assumed. authoring_requirements("variants.csv") answers
any_of: [['rsid'], ['chrom','start']], a grammar that cannot express VariantRow's third clause —
ref/alts require chrom and start. A guard built on it accepts {"rsid": "rs1", "alts": "G"},
which the model refuses and a compile would refuse, so the partial coordinate rides through
(@identity-whole-or-none). authoring_requirements still answers the human-readable which cells
are missing; it is not the verdict. Because the probe pre-fills every non-identity field with values
the model accepts, any ValidationError reaching it is an identity refusal — no message parsing.
What the registry holds, so no module keeps a private copy: match_on, the table, the kind
(projection — it later re-reads a column it drafted — or judgement), the precondition, and for a
projection the checked columns. provenance's DRAFT_PROJECTIONS is derived from it;
pubmind's projection identity genuinely differs from its match_on (the snapshot has no rsID
column) and that is a field with a required reason, refused at construction if absent.
What providers still differ on, deliberately, because unifying either would change behaviour:
the covered predicate — each provider's own reading of "this run contributed something" — and the
stale-label wording, since a published warning is an API and two providers ship two sentences for
this finding (@warning-text-is-api).
Adding a provider. Register it in DRAFT_PROVIDERS, derive the identity verdict through
drafting.skip_reason, and write provenance through record_draft_provenance. Two guards make this
inherited rather than remembered: test_drafting_scaffold.py asserts the registry equals the
*_draft.py modules on disk, and walks each module's AST to refuse a hand-listed identity column or a
direct merge_sources_file / withdraw_stale_dataset / record_source_terms call.
One thing that looks open here and is not. clinvar_draft and pubmind_draft write their licence
row on any non-dry run rather than gating on covered, which is the shape RM222 found wrong in
civic_draft. Both are nonetheless correct: each returns early — "nothing matched; no rows drafted" —
before the write, so the property holds upstream of the gate.
test_draft_licence_row_needs_coverage.py pins that early return, because it is what actually holds
the rule and nothing else asserted it.
CIViC — the direction axis, and a source whose coordinates are all on the wrong build (RM152)¶
just-dna-enricher civic build --release 01-Aug-2026 --out data/caches/civic
just-dna-enricher draft-panel spec/ --gene VHL --source civic --civic-cache data/caches/civic
It writes direction, never clin_sig, and the inversion is the finding. CIViC was proposed as a
second clinical-significance concordance authority. Measured, its germline subset carries five
ACMG-tier calls and zero benign-class, so discordant is unsayable about anything and the check
would read single or concordant by construction. What it does carry is 1,458 germline rows on
Predisposition/Protectiveness — this format's direction (risk/protective) — where the NA
count is 0, against 812 on clin_sig. Every number is in
CIVIC_SURVEY.md.
The build reads the dated bulk TSVs, not the GraphQL API, and the two are different sources. Only
the download side has dated releases, and a snapshot that cannot name its input cannot be reproduced.
The cost is that the TSV is accepted-only at 4,903 rows while the API defaults to NON_REJECTED at
11,518 — 2.35× apart, declared by neither — so release.json records status_basis and a count
from one surface must never be compared with a count from the other.
Three input files are required and the third is not padding: without
MolecularProfileSummaries.tsv, a profile naming two variants and a profile naming none are one
blurred drop reason instead of two different facts.
Coordinates are GRCh37 or absent, never GRCh38 — and nothing is lifted over. 2,189 gene variants
are GRCh37, 2,433 carry no build, two are GRCh38. The snapshot places rows by reading the identity
CIViC publishes beside the coordinate: an rsID, or a GRCh38 RefSeq accession inside ClinVar's HGVS.
That is RM48's rule applied rather than circumvented — and better than RM48 hoped, because a
published rs-number is an independent value ordinary resolution cross-examines, where a recovered
one would have resolution verify Ensembl against Ensembl. Scoring an accession needs the
per-chromosome map: NC_000001.11 is GRCh38 while NC_000002.11 is GRCh37. What carries neither
identifier is dropped by count, and RM153 is what to do about it.
Identity comes from the registry where CIViC states none, and never from a liftover. A row whose
only route is a ClinGen allele_registry_id is kept in the snapshot as identity_derivation="caid"
— a route to an identity rather than an identity — and draft-panel --source civic walks that route:
the registry returns an rs-number (preferred, because Ensembl then verifies it and two authorities
make the check real) or a GRCh38 coordinate. One-sided indels are anchored VCF/Picard-style: the
registry states an insertion with an empty reference allele and a deletion with an empty allele, and
prefixing both sides with the reference base before the event gives a VCF row at the registry's own
position. That position is HGVS's 3′-most one, so inside a repeat the row is then left-aligned
against the reference (sequences._left_align, RM273: rs72613567 anchors as 4:87310241 A>AA and is
written 87310240 T>TA); a repeat longer than the read window withholds as an unreadable anchor. That takes recovery from 48% of the direction set to 82%. --offline withholds those rows as
caid_unresolved — unplaced, not unplaceable — and an unreadable anchor withholds under its own
reason rather than being guessed at. Liftover was measured and refused: its ceiling is 13 rows, and
the one precise event in the class lifts exactly to the wrong allele.
An identity CIViC states in a variant's name is read out of it (RM159). 53 variants in the
01-Aug-2026 release carry nothing in the identifier columns, and for most of them the identity was
published all along in the name itself — N150fs (c.448delA), IVS2+1G>A, D1709N. Thirty-three
were resolved by hand and ship as the constant civic_identities.CIVIC_NAME_IDENTITIES, emitted with
identity_derivation="curated_name" — a member of its own, because rsid/grch38_hgvs mean "the
source stated this in the column for it" and a consumer must be able to exclude the difference without
re-deriving it. Coverage goes 237/290 → 270/290 of variants.
They are data rather than a draft-time lookup because four of the 33 needed a judgement no lookup
makes: a legacy IVS2 name whose structural conversion lands on the wrong exon (both readings being
real registered alleles 9 kb apart), a name pairing a missense protein label with a synonymous cDNA
change, a protein consequence standing over an intronic allele, and an rs-number that is
position-level where two alleles spell the same substitution. So the answers are shipped, the
procedure is written down in docs/probes/CIVIC_IDENTITY_PROTOCOL.md, and the build stays offline.
The one identity not adopted is TP53 R72P: it resolves to the reference allele
(g.7676154G=), and ref == alt is not a variant row.
Each curated identity is keyed to the exact name it was read from, and lands in one of four
counted states in release.json — applied, superseded (CIViC now publishes one of its own; the
source always wins, and this is also the cheapest currency signal there is), renamed, absent. The
four sum to the table, asserted as an equality. allele_registry_id stays CIViC's verbatim cell: the
CAIDs the resolution went through are provenance on the table, never written into the source's column.
civic reproduce then cross-examines every placed coordinate against the GRCh38 reference through
refget — 57 of 57, 0 mismatches, up from 24 before the adoption.
The unreviewed majority is readable without leaving the dated release (RM169). CIViC publishes
<date>-civic_accepted_and_submitted.vcf beside the TSVs, so --submitted widens the basis with no
API read and no loss of reproducibility: 507 rows on 270 variants → 1,149 on 397, a rebuild still
byte-identical, and the refget cross-check going from 57 coordinates to 129, 0 mismatches. Every
row carries evidence_status — CIViC's own word, unconverted — and release.json records
status_basis, status_counts, vcf_evidence and unjoinable_submitted.
The VCF is joined onto the TSVs, never substituted for them. A VCF record needs a POS, so the file cannot carry a variant with no GRCh37 coordinate — and 52 of the 54 variants it drops are exactly the class whose identity had to be read out of its name (RM159). Reading it as the row source would discard the hardest-won half of the snapshot.
VariantSummaries.tsv is accepted-only too, which is why identity_derivation="vcf_csq" exists:
112 of the 127 variants the submitted evidence introduces have no row in it, so their identity comes
from the same CSQ entry, through the same parsers, on the same published identifiers (57 by CAID, 40
by rs-number, 14 by a GRCh38 accession, 1 by both). The member names the file rather than the
route, because the route is already visible in the row's own cells. Nothing is placed from the VCF's
own position: it is GRCh37 and lifting it stays refused.
A refutation is kept and its direction withheld. Does Not Support removes a claim without
establishing the opposite one, so the row keeps its raw words and states no direction — an unknown
is withheld, never negated. The drafter reports how many it held back and why.
And an authored direction beside a refutation is a finding, on two surfaces (RM170). Withholding
the refuting row was never the whole answer: a variant CIViC supports and also rebuts still gets a
risk row drafted — correctly, because the provider writes what the source said — and until 0.7
nothing then told the author the rebuttal existed. contested_variants cannot see it: that counts a
variant whose camps hold both risk and protective, and a refutation enters no camp at all, so
the counter is correctly 0 on every basis and the muddy variants are invisible to it.
draft-panel --source civicnames them at the point the row is written: the variant, the supporting evidence id, the refuting one with its status, and the snapshot'sstatus_basis. The row is still written.enrichfolds in thepublished_refutationcheck whenever a CIViC snapshot resolves ($JUST_DNA_CIVIC_CACHE), so a hand-authored module that never ran the drafter meets it too. Two finding codes, because they are two sentences:refutation_beside_claim(the source asserts and rebuts) andrefutation_without_claim(the module asserts a direction the source has only denied). They are this pass's own keys and travel in the record'sdetail— deliberately notVALID_WARNING_CODESmembers, which is the compiler's vocabulary and whose guard asserts every member is built by a compiler check. A compile restates them asverification_findings_recorded.
Warns in both modes and escalates in neither — a source disagreeing with itself is not an
authoring error, the same call clin_sig makes. Nothing is repaired: a refutation withholds a claim
rather than establishing its opposite, so there is no opposite value to write even if the tier were
allowed to write one.
The finding keys on the refuting evidence item, then fans out to the rows it touches. CIViC
evidence item 8721 is one statement about the two-variant genotype VHL S183L AND VHL D126N, which
the snapshot writes as two single-variant rows (RM174) — so keying on the variant would report one
rebuttal as two independent ones.
The record states its basis on every run, including the empty one. Every assert-and-refute pair in
CIViC rests on a submitted rebuttal: both accepted refutations in the whole database stand against
nothing at all (probes/CONTRADICTION_CORPORA.md). So on the accepted basis this class is empty by
construction, and findings: 0 without the basis beside it would read as clear water.
A row names the profile its evidence item actually belongs to (RM174). CIViC publishes molecular
profiles as boolean expressions over variants — VHL S183L (c.548C>T) AND VHL D126N (c.376G>A) — and
an evidence item on one of those is a claim about a combination genotype. Reaching the builder
through the VCF, it is fanned out into one row per variant it names, and molecular_profile_id has to
be the variant's own single-variant profile or the row does not join at all. So the item's real
profile rides beside the join key in evidence_molecular_profile_id and
evidence_molecular_profile_name, on every row rather than only on composites — a column that is
null on the common case invites a reader to treat null as "not a composite". A composite is
evidence_molecular_profile_id != molecular_profile_id, derived rather than stored, and
release.json publishes the count as composite_profile_rows.
The name is null on a TSV-sourced row because MolecularProfileSummaries.tsv publishes none, and
filling it from the variant's name would state a profile name the source never wrote. The path
asymmetry is deliberate and audited rather than inherited: a multi-variant profile arriving through
the TSV is dropped as combination_profile and counted, while one arriving through the VCF is kept
and labelled. Making both paths keep it would move an accepted-basis build's numbers, and what the
format can honestly represent about a two-variant claim is a different question — that one is
RM28's, parked, and RM174 gave it its first counted corpus entry.
Every drop is counted, and the somatic majority is the point of that. Roughly three quarters of
CIViC describes tumour tissue no germline genotype can satisfy. A filter whose scope is narrower than
its name is what this adoption was designed against, so input_rows == record_count + sum(dropped) is
asserted as an equality over a walked registry and every reason lands in release.json.
There is no --use flag on civic build. CIViC is CC0 on every axis, so a declared-use gate would
permit every build unconditionally, and a flag feeding a gate that never gates is a flag that does
nothing.
civic citations — the citations no dated file can carry, and the canary on them (RM160)¶
The gap is structural rather than a coverage shortfall. RM169 widened the snapshot as far as a
dated file goes, and the wider basis is a VCF — so a VCF record needs a POS, and CIViC publishes no
GRCh37 coordinate for a variant it names as a class of event or as a legacy notation. The submitted
evidence on those records is published on exactly one surface, the GraphQL API. Ten records' citations
are unreachable from every file the builder reads, and one of them is variant 1955 VHL P71fs
(c.211insT), whose only reachable evidence for the numbering convention its identity turns on is EID
9969 (PMID 12202531, Dollfus 2002, free full text) — submitted, in the API, in no file. The counts are
in probes/CONTRADICTION_CORPORA.md and probes/CIVIC_SURVEY.md.
civic build and civic reproduce are untouched, and that is the whole of the decision. RM160
was settled as shape 3: read SUBMITTED at enrich time, where network reads already live and
reproducibility is never claimed. Hashing an API capture as a build input (shape 1) keeps the word
reproducible while changing what it is reproducible against, and a second API-built parquet (shape 2)
was dissolved by RM169. The read is one request per variant by construction — evidenceItems
takes a single variantId — which is why it fits a command an author runs over their own module and
would not fit a builder pinning a dated release. Batching it into civic build is the first repair
anyone proposes and it is exactly the bargain this shape refused.
Three routes reach a CIViC variant id, and the third exists because the first two miss the class
this is for. The snapshot's own coordinate join goes through clinical.comparison_plan, the same
resolved-(chrom, start, ref, alt) route the refutation leg and the ClinVar cross-check use, so every
CIViC question in this tier is asked about the same alleles. The curated name-identity table (RM159)
is how a variant whose identity CIViC publishes only inside its name is reachable at all. Neither
reaches 1955 — it is one of the two records that round could not resolve — so --variant-id N
asks about an id directly and writes a module-level citation row, which StudyRow has permitted
since RM47.
A recovered citation carries the same identity its variant row got — rsID where there is one, the
coordinate otherwise, never both. That is clinvar_draft's rule, learned from the compiler's orphan
check on the first real panel, and it is load-bearing twice here: a row carrying both has a different
match_on signature from the rsID-only rows draft-panel --source civic writes, so the two lanes
would each append a row under one (variant_key, pmid) and the compiler would call them duplicates.
The canary joins on StudyRow.variant_key for the same reason — the model's own derivation over those
cells, rather than a tuple of raw columns this lane spells its own way.
A recovered citation lands in studies.csv and nowhere else. literature.csv is the derived
article table, filled by the literature command from the PMIDs studies.csv names, and an article
row nothing cites is dropped from the artifact. Writing one here would either duplicate that pass or
add a row the compiler discards; drafting the citing row and letting literature fill the article is
the pairing that works in both directions. Several evidence items citing one paper are one row —
five of CIViC's items cite PMID 17661816 on variant 844 — because (variant_key, pmid) is the grain.
status rides as confidence/confidence_unit, unconverted (civic_evidence_status). CIViC's
own instrument named rather than translated into a house grade, so an accepted row and a submitted row
are not the same row once both are in the file. Where the live items behind one paper disagree, the
confidence is withheld rather than picked. Rejected evidence is not drafted at all: status:
ALL returns items CIViC's editors threw out, and a paper whose every item is rejected is content the
source repudiated — counted under rejected_by_source, never silently dropped. A citationId on a
non-PubMed source (ASCO and ASH abstracts) is a real id in another namespace and withholds rather than
becoming a pmid.
The pin is on the SourceRow, and the layer is literature. Every drafted row records when the
API was asked (fetched_at) and on what basis (dataset = civic_api:status=ALL) — not a timestamp in
each conclusion, which would fold the moment of a read into content_signature. (civic,
annotation) stays civic_draft's slot: a second surface of an already-declared source may not claim
the lane's row, and a civic_api source would publish a route as a licensed body. literature is
what this pass really supplies and is orphan-check exempt. Merge-not-clobber means the pin records the
ask that first put a recovered citation into the module, which is a floor on not asked since — and
the canary is what closes the gap.
evidence_status_currency is the canary, and it is what makes drafting from a live read honest.
enrich() re-asks CIViC about every citation the module records from this lane and reports what has
moved: a status accepted or rejected since the draft (civic_evidence_status_moved), or a citation
CIViC has added since (civic_citation_added). Two codes because they are two remedies. It warns in
both modes and escalates in neither — a source re-curating its own evidence is not an authoring
error — and it is deliberately not dataset_currency: that one asks which release a table came
from, this one asks whether a per-item judgement has moved, and the two currency findings stay apart.
Its four skip reasons are cleared by four different things, and none of them is a pass:
| skip | what it means |
|---|---|
nothing_to_check |
the module records no citation from this lane — a civic_draft row carries no confidence_unit, so it is not re-asked against itself |
offline |
the run had no egress. Never ran, findings=0: nobody-asked and the-source-has-nothing-more are different facts |
no_reference |
no recorded citation could be mapped back to a variant id this run — no snapshot, or rows that ground the module rather than a variant. Published with its count rather than silently passed |
The two sets the canary reads are deliberately different widths. Which citations have a recorded
status to compare is the narrow one — rows carrying confidence_unit = civic_evidence_status, so a
civic_draft row is never re-asked against itself and a row whose confidence was withheld is not a
subject. Which papers the module already cites is the wide one, over every studies.csv row
whatever wrote it, because the added-a-citation arm asks whether there is a row for this paper at all;
keying that on the narrow set would report a citation sitting in the file as one CIViC has added,
every lap (@a-set-that-silences-is-narrower-than-one-that-raises, taken the other way round).
| unreachable | CIViC answered for none of the variants asked about |
Cache internals — locations, resolver, download¶
locations.py(moved fromjust_dna_compiler.cache) —resolve_ensembl_referencelocates a usable reference by precedence (explicit arg →$JUST_DNA_ENSEMBL_CACHE→$JUST_DNA_PIPELINES_CACHE_DIR→ platformdirs), anddefault_ensembl_cache_dir/load_env. It never downloads — location only.resolver.py(moved fromjust_dna_compiler.resolver) — the DuckDB engine:_connect/_view_over_parquetover a.duckdbfile or adata/*.parquetdir,resolve_variants(fill/expand/ verify with the one-to-manyORDER BY id, chrom, start, refexpansion), and the publiclookup_locithe enricher and (until 1.0) the compiler's deprecated path share so they never drift.resolver.probe_table— a batch lookup must HASH its probe, and this is why a panel used to never finish (0.5.2). DuckDB cannot fold a disjunction of equality conjunctions into a hash probe, soWHERE (chrom=? AND start=? AND ref=? AND alt=?) OR …is evaluated against every row of the reference: cost grows withalleles × rows, quadratically in the module. A 297-gene panel ran two hours at 12% CPU with no I/O and looked like a deadlock. Measured on the 4,431,781-record snapshot, same 5,000 alleles, same connection: 88 s OR-chained against 0.21 s joined against a temp table. Three call sites moved (clinvar.lookup_clin_sig,resolver._lookup_rsid_candidates, andclinvar.select_by_gene, which is single-column and becamegene IN (…)— 20.9 s → 6.6 s, sinceINis pushed into the parquet reader and an OR-chain is not)._lookup_positions_by_rsidandcitations_foralready usedINand were left alone.
Two things about it that must not be "simplified". The probe rows are rendered as SQL literals,
escaped the way _connect escapes the parquet path, because DuckDB's Python parameter binding is
where the remaining time goes — same query, same data: literals 0.21 s, a composite-key
IN (?, …) 1.04 s, a parameterized UNNEST(?::VARCHAR[]) 3.51 s, executemany 8.6 s.
Parameterizing it back gives up most of the win, so re-measure before changing it. And every caller
keeps its own ORDER BY: a join reorders nothing by itself, and emitted row order is digest-visible
(Principle 7). A regression guard lives in test_query_shapes.py — it asserts the plan contains a
hash join (no clock involved) and, separately, times both shapes in one process so a slow runner
moves both numbers together.
- ClinVar cache location — locations.resolve_clinvar_reference mirrors the Ensembl ladder
(explicit arg → $JUST_DNA_CLINVAR_CACHE → $JUST_DNA_PIPELINES_CACHE_DIR/platformdirs, under a
clinvar/ subdir), also never downloading. The ClinVar snapshot ships as parquet only (no prebuilt
.duckdb).
- One resolver body, six callers. _resolve_parquet_cache(explicit, env_var, default_dir) is that
ladder, written once: it was copied per snapshot and the copies had already drifted (ClinVar's silently
lacked the bare-.parquet case its constraint sibling had). accept_bare_file=True is the one
difference, and only the single-file constraint snapshot wants it. _cache_dir(subdir) reads
$JUST_DNA_PIPELINES_CACHE_DIR at call time, not at import, so a .env loaded by load_env can
still change the answer — and since 0.5.2 it calls load_env() itself, which is the fix for a
family of "the cache is right there" reports. _resolve_parquet_cache loads the environment inside
itself, but each resolve_*_reference passes default_*_cache_dir() as an argument, evaluated
before the call: with the base set only in .env, the first resolve in a process computed its
default from platformdirs and returned None, and every later one was correct. That asymmetry
produced three separate bug reports — cache pull writing into ~/.cache while cache status
looked in the configured directory and called the snapshot absent moments after a successful pull,
draft-panel --offline refusing with "no ClinVar snapshot found" for a snapshot cache status
reported present, and a test module whose first skip-guard silently skipped. One load, six resolvers,
both CLI paths; override=False, so a real environment variable still wins.
The same argument-position evaluation made load_dotenv_file=False a knob that did nothing, in all
six resolvers, until 0.6.3 (S39). The load above was unconditional, so it ran before the resolver
had looked at its own flag — and a caller passing False still had the whole .env written into
os.environ. The flag is now threaded through _cache_dir and the six default_*_cache_dir
helpers rather than the load being removed: the unconditional load is the repair described above, and
the True path is byte-identical. test_locations.py pins both directions and walks the two
families asserting each takes the parameter, so a seventh snapshot's resolver cannot quietly reopen
it. Same shape as the watcher's ${BRANCH:-main}: a knob's disabling value is its own case and needs
its own probe, not a reading.
- locations.read_release(reference) — a snapshot's release.json as a dict, or None when it is
absent or unreadable. Written by every builder and, until 0.5.2, read by nothing but cache status;
it is what lets enrich() compare a module's panel: pin against the snapshot in front of it. None
for both absence and corruption on purpose: a caller must not be able to mistake "this snapshot does
not say" for a release id.
- download.py — ensure_snapshot, ensure_clinvar_snapshot and ensure_constraint_snapshot pull
the parquet slice from the HF datasets (just-dna-seq/ensembl_variations / just-dna-seq/clinvar /
just-dna-seq/gnomad_constraint) via one shared footer-checked/atomic body. A complete parquet
begins/ends with the PAR1 magic; downloads go to a .part temp and rename only after the footer
verifies, and a corrupt/truncated file is removed and refetched rather than skipped forever.
huggingface_hub is a guarded lazy import — a missing wheel fails with a clear diagnosis pointing
at the install or the --*-cache flag. Every pass that wants a snapshot provisions through these:
enrich for Ensembl + ClinVar, gene_metrics for constraint. A published dataset accumulates, so
each ensure_* fetches only the files its snapshot is made of — clinvar-*.parquet and not the
159 MB clinvar.parquet the repo still carries from the single-file era, whose columns are the raw VCF
INFO fields. The reader globs data/*.parquet, so importing that one file would put two schemas under
one DuckDB relation and every query would die on Referenced column "clin_sig" not found. A foreign
file already in a local cache is reported, never deleted, and the message names it and the fix.
release.json comes down with the data, so a provisioned snapshot can state its own release — it is
what a drafted module's recorded dataset is derived from (RM4), and a cache that cannot state its
release is one a drafted module cannot name. A repo without one still provisions; absence is not an error.
LICENSE.txt rides along on the same rule, and it did not until 0.5.1. upload's allow-patterns
were data/*.parquet, citations/*.parquet and release.json, so publishing a share-alike snapshot
silently dropped the one file the pinned-licence design exists for — clinpgx_build extracts ClinPGx's
terms out of the very archive the data came from precisely so a holder of the snapshot can read what
governs the bytes, and license_sha256 pins nothing for someone who never received them. Both halves
are fixed: the publisher sends it and the provisioner fetches it. Absence stays normal (only ClinPGx
ships one).
gnomAD v4.1 — three roles, one endpoint (gnomad.py)¶
gnomAD enters in three deliberately different kinds of role: a resolution link (above), an allele-frequency pass, and a gene-constraint pass. One GraphQL endpoint serves all three.
Rate limiting is the design constraint. gnomAD allows 10 requests per IP per 60 seconds, so a
request per variant is unusable. Everything is batched by GraphQL field aliasing: probed live, batches
of 20 and 25 succeeded and 29 returned HTTP 400, so batch_size=20 with a 6 second pacing gate —
exactly the stated budget, about 200 variants/minute. Pacing comes before retry: tenacity retries
transport errors, timeouts and 429s, but a blind retry spends the same budget that caused the 429.
Partial failures never sink a batch. GraphQL puts per-alias errors in errors[] and still returns
data for the rest (a probed 20-alias batch came back resolved=17, errors=3), so every parser reads
both halves. An error with no alias path is different — that is our broken query, and it raises.
A multi-allelic rsID cannot be looked up by rsID. variant(rsid: "rs334") answers only
"Multiple variants found, query using variant ID to select one." — and rs334 is sickle-cell, exactly
the kind of variant a module carries. Those rsIDs are collected and retried through variant_search,
which returns every matching variant id. The frequency pass never meets the problem: it keys on the
already-resolved chrom-pos-ref-alt.
Pass 2 — allele frequency (frequencies.py, online only)¶
enrich_frequencies(spec_dir, *, mode, offline, populations, dataset, write, client) reads the
coordinates in resolution.csv and writes frequencies.csv: one row per (allele, ancestry group)
carrying AC/AN. Existing rows are authoritative and merged, never clobbered.
Three things the raw payload does that the pass undoes, all visible in the committed recording:
- it carries sex splits (
nfe_XX) beside ancestry groups — a second axis (Principle 5), dropped; - it lists
XX/XYtwice — deduplicated; - it names the whole-dataset row with a bare empty id — mapped to
global.
Server order is not preserved (it is not promised, and the table must be byte-stable): rows are sorted
by (variant_key, alt, population_order). Per-population af is not exposed by the API at all —
frequency is AC/AN and we compute it, which is why the CSV stores integers and allele_frequency is
derived. faf95 is a single value with a named owning group, so it lands on that group's row only.
This is the first online-only link in the whole chain, and it will stay that way: the v4.1 sites
VCFs are 58 GB (exomes) and 742 GB (genomes), so there is no slice to ship. --offline makes the pass
a no-op with a warning rather than a failure — and that is not a reproducibility hole, because once
frequencies.csv is written it is the pin, and every later compile reads it offline.
status has three members, and the third exists because gnomAD's callset has a hole.
VALID_FREQUENCY_STATUS is {resolved, not_found, not_covered}:
resolved— the source served counts.not_found— the source was asked and has no such allele. A fact about a locus it does cover.not_covered— the source does not cover the locus, so it has no answer and none can be inferred.
The third was added after the pass was found writing not_found for a Y pseudoautosomal locus, whose
comment claimed the row was a fact ("gnomAD was asked and does not have this allele"). It is not: gnomAD
hard-masks the Y PAR — those bases duplicate the X PAR — so it never looked, and before PAR selection
landed a one-to-many expansion handed this pass ten such loci per SHOX panel. That is the None ≠ False
rule: an unknown may not be recorded as a negative. gnomad.covers_locus decides it (three-valued, and
the source convention lives there while the PAR geometry stays in vrs), and such a locus is now not
queried at all — the request would spend a slot of a 10-per-minute budget to learn nothing, and asking is
what produced the false absence.
not_covered rather than unchecked, which is this codebase's word for a question that was never put
(acmg.py: the row named no gene, the list could not be reached). This is the stronger statement that the
source's scope excludes the locus. FrequencyResult reports them in uncovered, kept apart from
missing, and they are outside the strict gate on purpose: a locus gnomAD cannot cover is perfectly
reproducible, so refusing would make a pseudoautosomal module uncompilable for a reason no authored edit
could fix. (FrequencyRow.status also gained a validator here — until 0.5 it was free text on a fact
table.)
Pass 3 — gene constraint (gene_metrics.py, offline capable)¶
enrich_gene_metrics(spec_dir, *, mode, offline, constraint_cache, dataset, download, write, client)
takes the module's gene symbols (deduplicated in first-occurrence order) and writes
gene_metrics.csv: pLI, LOEUF, missense Z and friends, one row per gene. Snapshot first, live API second.
Since RM157 that is every authored table carrying gene, not variants.csv alone. module_genes
is not a report — it is the scope of this pass, of gene_validity and of the ClinGen dosage pass,
all three of which call it — so a module whose genes live in its PGx tables had all three quietly do
nothing: no rows, no findings, no line saying a question had not been put. Measured on this repo's own
corpus, where cyp2c19_star_alleles, apoe_epsilon, cyp2c9_warfarin_grch37 and hfe_compound_het
returned an empty gene set while naming CYP2C19, APOE, CYP2C9, VKORC1, CYP4F2 and HFE. The set is
derived from the same registry walk the identifier roster uses, so a table kind that gains the column
joins by existing; a table that will not parse refuses here rather than being routed to not_read,
because half a scope is a silently narrowed one. pgx._GENE_TABLES is deliberately not this roster: it
is the pair whose presence decides whether the star-allele cross-check applies at all.
This is the one gnomAD role that works with zero egress, and the difference from frequency is purely size: gene-level constraint is one row per gene, single-digit MB as parquet.
The snapshot is provisioned, not merely hoped for. With no local snapshot and not offline, the pass
calls download.ensure_constraint_snapshot before it considers the API — the same shape enrich() uses
for the Ensembl and ClinVar snapshots, with --offline as the only switch (there is deliberately no
second flag). That wiring was missing until 0.5: ensure_constraint_snapshot had existed since the
download body was generalized and had no caller, so a plain install fell straight through to the live
API and quietly recorded v2.1.1 numbers — then warned about the release difference for a snapshot it
had never tried to fetch. A provisioning failure degrades to the API rather than sinking the pass (HF has
gone dark mid-demo), and the warning names the consequence: older numbers, not no numbers.
The two routes are different releases, and the table says so. Checked against both: for BRCA1 the bulk v4.1 file gives pLI 1.55e-34 / LOEUF 0.885 / mis_z 2.338, while the live API gives 5.52e-38 / 0.928 / 1.734 — same gene, same MANE transcript. The live
gnomad_constraintfield serves v2.1.1 constraint; v4.1 ships only in the bulk file. So a row records which release it came from:gnomad_v4.1_constraintfrom the snapshot,gnomad_v2.1.1_constraintfrom the API, and the fallback logs a warning.datasetis inside the fact set precisely so these cannot be confused — labelling both as v4.1 would record a false fact and make the fact-hash claim two different tables identical.
gnomAD constraint snapshot (constraint_build.py, [dev])¶
download_constraint_tsv + build_snapshot reduce gnomAD's per-transcript constraint TSV to the
gene-level parquet the pass reads. The source is 95.5 MB (not the 4.2 MB some aggregator pages
claim, which comes from an explicitly illustrative demo listing) at release/4.1/, not
release/v4.1/ — anonymous bucket listing is disabled on both the GCS and S3 mirrors, so the path came
from the UCSC track makedoc and was verified directly.
The row pick is load-bearing, not cosmetic. The TSV is per-transcript, 55 columns, and mixes RefSeq
with Ensembl rows for the same gene — and both carry mane_select=true:
A1BG 1 NM_130786.4 canonical=true mane_select=true <- RefSeq
A1BG ENSG00000121410 ENST00000263100 canonical=true mane_select=true <- Ensembl
A1BG ENSG00000121410 ENST00000600966 canonical=false mane_select=false
BRCA1 and MYH7 show the same shape, so this is the file's general structure. A naive "first
mane_select row wins" returns whichever the file happens to list first — for A1BG the RefSeq row,
whose gene_id is the bare NCBI id 1, useless as a stable identity. The rule is mane_select AND
an ENSG-shaped gene_id, falling back to canonical on ENSG, else the gene is dropped as
unresolved rather than guessed. The test feeds the real rows in both orders and demands the same
answer, which is precisely what the naive implementation fails.
The reference-allele check (sequences.py)¶
verify_reference_alleles(rows, *, sequences, offline) compares each row's ref against the actual
bases at its coordinate and returns the disagreements. It runs inside enrich() by default
(--verify-ref/--no-verify-ref) and its findings land on EnrichmentResult.ref_mismatches.
Why it has to exist. A VRS allele id is built from which sequence, which interval, and what
replaces it. The reference allele is not a component — the refget accession plus the interval
already determine it, since sequence[start:end] has exactly one answer. That is correct and
deliberate (a content-addressed identity must be a function of the allele, not of the claim about it,
or two records for one allele could get two ids and the whole scheme collapses). But it means minting
never looks at the authored ref, and VCF's free consistency check — which catches liftover slips,
off-by-ones and wrong-assembly errors — is gone. VCF can afford REF because its CHROM is a name
rather than a digest, so REF is genuinely load-bearing there. VRS traded that redundancy for
canonicality. This check buys it back on the one tier that has the sequence to do it with.
Two failure modes, and the claimed length decides which. The actual bases are always read at the claimed length, so the two cannot be told apart by comparing lengths:
- a single-base claim — absorbed.
11:5227002 C>Aand the trueT>Amint the same id, so the minted identity is still correct and nothing downstream could ever notice the bad row. Only this check reveals it. - a multi-base claim — corrupting. The claimed length sets the interval, so a wrong
refmakes the id span the wrong bases and name an event the author did not intend.RefMismatch.distorts_the_allele_idreports which case a finding is.
RefMismatch.shift is the field to surface first, and it names a third cause. A mismatch does not
only mean the ref cell is wrong; far more often in the wild the position is wrong and ref was
right all along for the variant the author meant. shift is the offset, in bases, at which the
authored ref is the reference sequence — +1 meaning the variant sits one base right of the
authored start, which is exactly what subtracting one from a 1-based VCF POS produces. It is
None when no neighbour explains it, so nothing is claimed; diagnosis renders whichever cause was
established and is the grouping key for a run's summary. The distinction is not cosmetic. Reporting a
shifted coordinate as a bad ref sends the author to the wrong column while leaving them to wonder
why their coordinate validated against dbSNP — and a shifted row always distorts the allele id
whatever its length, because the id is minted at the authored position, so the compiler's VRS pass
recomputes the same wrong id and reports it verified.
It reports; it never repairs. The row is left exactly as authored. Rewriting it would destroy the evidence that something upstream is wrong and silence a problem the author needs to decide about.
Rows with no coordinate, and rows whose ref is not plain ACGT (a symbolic allele, RM5), are not
checked — abstaining beats inventing a verdict. Reads are cached by (accession, start, end), so a
module asking about one locus repeatedly costs one round trip, and the same SequenceProxy is shared
with indel minting so a run builds one proxy in total. Needs sequence access, so --offline skips it:
a check that cannot run is not a check that passed, and the run says so rather than implying success.
The old assembly: rs-number recovery and a wrong-build diagnosis (grch37.py, RM48)¶
An author curating from older literature has hg19/GRCh37 coordinates and the module must be GRCh38. Nothing in these four packages converts, so the conversion happens off-tool and lands as an ordinary authored coordinate with no provenance at all. The compiler refuses the coordinates that are provably impossible; this module answers the ones that are merely wrong.
No chain file, no provisioned asset, no new licence. The roadmap's stated blocker was that
recovering an rs-number needs "either an hg19-keyed dbSNP surface or a chain file … i.e. the whole
snapshot apparatus for one authoring convenience". Probed 2026-08-13 and false: Ensembl runs a
permanent GRCh37 REST service at grch37.rest.ensembl.org with the same API shape, serving both
dbSNP variants (/overlap/region) and reference bases (/sequence/region). The same request that
answers "which rs-numbers sit here on GRCh37" also discriminates the builds outright —
7:140453135..140453137 is CAC on GRCh37 and GTT on GRCh38.
Recovery, never liftover, and the reporter argued their own request down. If the paper gives an
rs-number, liftover is unnecessary and strictly worse: authoring the rs-number produces the
independent second value resolution._verify cross-examines. So liftover is only reachable where there
is no rs-number and only an old coordinate — and in exactly that case the lifted coordinate becomes the
row's sole identity with nothing to check it against, a generator of unverifiable-by-construction
identities. That is the hazard class behind this tree's 3,038-row off-by-one, where a content-addressed
id was a correct digest of the wrong input and every offline gate passed, --strict included.
recover_rsid(chrom, start, *, ref, alts, client, offline) answers with one of four outcomes.
Three of them — recovered / ambiguous / none — are the ones pyliftover fuses, reporting "no
result" both for a position that maps nowhere and for one that maps to several. The fourth is
unchecked, on the other axis: S20 established in this same resolution path that an unreachable source
is unchecked rather than absent. A 4xx is an answer (the service 400s on an unknown contig and on a
position past the end of one); only a 5xx, a transport error or a timeout is unchecked. --offline
reports skipped_offline, never a pass.
The match is anchored on the authored position: a candidate must start exactly there, carry the
authored ref, and contain every authored alt. Anchoring is what keeps it honest — at
7:140453136 seven features overlap the base, two merely span it, one is an HGMD record with no
rs-number, and four dbSNP records genuinely start there, which is why a position-only query is
ambiguous rather than under-specified. The consequence to know: an indel authored in VCF's padded
spelling (POS on the base before the event) will not match Ensembl's unpadded record and comes back
none with that said, rather than wrong.
diagnose_wrong_build(mismatches, *, client, offline) runs only over rows the reference-allele check
already rejected, which is the whole cost control — a module whose refs agree makes no request here.
verify_reference_alleles skips any build refget_accession has no table for, so a RefMismatch only
ever exists for a GRCh38 module, which makes "the other assembly" always GRCh37 rather than a parameter.
Three tiers of evidence, and the message says which one it has:
| Tier | Evidence | What it licenses |
|---|---|---|
single_base_match |
one authored base equals the GRCh37 base there | suggestive only — one base in four agrees by chance, and VCF 4.4 §1.6.1.4 requires an ambiguous reference base to be reduced to the first alphabetically, so an authored A may be a lossily reduced R |
multi_base_match |
several consecutive bases agree | chance does not explain it |
dbsnp_corroborated |
the bases agree and GRCh37 dbSNP records a variant starting there | the strongest, and the only one that names the rs-number to author instead |
The two strong tiers supersede the ±1 neighbour reading, and that is not decoration. On the real
HFE pair — 6:26093141 and 6:26091179, authored from the GRCh37 literature into a GRCh38 module —
_read_with_neighbours reports "coordinate shifted 1 base to the right" for both, confidently and
wrongly: the true variants are 228 and 411 bases away, and a neighbouring base equal to the authored
ref is a one-in-four event. Two explanations printed side by side with nothing to order them is the
shape this codebase keeps fixing, so the summary says which wins. A single-base match does not
supersede a shift — both rest on one agreeing base, and ordering them would invent a verdict.
BuildDiagnosisResult.not_checked carries the reason when the pass did not run (skipped_offline,
no_ref_mismatches) and is None exactly when it did, for the same reason clin_sig_not_checked
exists: an empty list otherwise says both "asked, and nothing points at another build" and "never
asked". The diagnosis travels inside the strict refusal rather than beside it, because a strict
run raises and returns nothing, so a result-object-only answer would be visible to the mode that does
not need it and invisible to the one whose whole output is that sentence.
It writes nothing. just-dna-enricher hint recover --chrom 7 --start 140453136 --ref A --alts T
reports rs113488022 as an advisory Alteration with applied=False and refusal="identity_bearing"
— the sharpest refusal in the table, because an rs-number is the row's identity and a machine filling
one performs an identity migration by network lookup with no authored edit anywhere. Several candidates
are reported and never picked. The author types the rs-number into variants.csv and drops the old
coordinate; a later enrich places it on whichever build the module declares, and resolution.csv's
source column records which link answered. That is where provenance goes — never into an ordinary
authored coordinate.
GA4GH VRS allele identity (vrs.py)¶
mint_resolution_rows(rows, *, minter, offline, source_ids) stamps vrs_id/vrs_spec onto resolved
rows by three routes:
- stdlib — a substitution, via the format tier's
derive_vrs_allele_id. Zero egress, zero heavy dependency, and byte-identical to whatga4gh.vrsand the live gnomAD API produce. - normalized — an indel/MNV, justified against the reference through the seqrepo REST data proxy.
Needs the network, so
--offlineskips it. - null — an indel offline, an unreachable sequence service, an off-assembly contig, or an allele
that names no sequence at all. A missing id is honest; an unjustified one would be a
ga4gh:VA.…that looks interoperable and is not.
A source's own id (gnomAD serves one) is cross-checked, never trusted over the minted value — the point of a content-addressed identity is that it does not depend on which sources happened to answer.
An allele that names no sequence is a permanent class, and it used to be a crash. RM5's symbolic
alleles (<DEL:4977>), RM58's . and RM59's * are legal cells the mint pass routed to indel
normalization like any other non-substitution — and ga4gh.vrs's LiteralSequenceExpression accepts
^[A-Z*\-]*$, so the first two raised an unhandled pydantic.ValidationError that killed a whole
enrich run over one row, while * passed the pattern and would have been handed a content-addressed
id for a state that is not a sequence. vrs._sequence_free_reason decides the class once, for both
mint and why_not, so the verdict and the sentence cannot drift: the row keeps no id, the run
carries on (the UnsupportedBuildError guard beside it is the same shape), and the reason says the
id is unreachable on any run rather than offering the --offline re-run that used to be the
crash. One constant string per class, because unmintable_reasons groups on it. Still a warning in
both modes — no authored edit makes a symbolic allele mintable (P5).
Publisher surface — module upload (upload.py, [dev])¶
The author/publisher half of the enricher's HF use (snapshot download is a runtime path; module
upload is for republishing, e.g. the Gen-I v1-port recreation). Extracted from
just_dna_pipelines.v1_port.publish so lite has a canonical home to adopt.
ensure_repo(repo_id, token=None) # create-or-update (create_repo exist_ok=True)
plan_upload(module_dir, name, repo_id=None) -> UploadPlan # dry-run; validates artifacts present
upload_module(module_dir, name, repo_id=None, token=None, commit_message=None) -> UploadPlan
plan_reference_snapshot(snapshot_dir, repo_id=None) -> SnapshotPlan # dry-run for a reference snapshot
publish_reference_snapshot(snapshot_dir, repo_id=None, token=None, commit_message=None) -> SnapshotPlan
upload_module uploads every parquet the compiled artifact carries + manifest.json + optional
logo and readme to datasets/<repo>/data/<name>/ (default repo just-dna-seq/annotators), matching
just-dna-lite's discovery layout, and the same files again to data/<name>/v<version>/ (RM84 —
see below). The parquet half of the allowlist is just_dna_compiler.compiler.ARTIFACT_PARQUETS,
imported rather than restated — see what publishes, and what refuses below for why that is a rule
rather than a tidiness. publish_reference_snapshot uploads a built data/*.parquet + its
parquet sidecars + release.json to the root of a dataset repo (default just-dna-seq/clinvar),
matching the download.ensure_*_snapshot layout. Both go through ensure_repo — one
create-or-update-then-upload pathway (create_repo was added here; the origin v1_port.publish assumed
the repo pre-existed).
The sidecar was the gap, and it made downloaded snapshots second-class. ClinVar's
citations/was built and published nowhere, so a consumer who provisioned the snapshot had no PMIDs while one who built it did — anddraft-panelcannot produce a compilable module without them, becausestudies.csvis mandatory. The layout lives once inlocations(SNAPSHOT_DATA_DIRNAME/SNAPSHOT_SIDECAR_DIRNAMES/CITATIONS_DIRNAME/RELEASE_FILENAME) because four parties have to agree on those names — builder, publisher, provisioner, reader — and every disagreement so far has been silent. A sidecar stays a sibling ofdata/: the readers globdata/*.parquet, so a two-column citations table inside it is the same poisoning a stale single-fileclinvar.parquetcauses. Absence is normal (only ClinVar has one, only afterclinvar citations), so neither end treats it as an error. Each needs a write token (hf auth loginorHF_TOKEN) — a missing one raisesPermissionError;huggingface_hubis a guarded lazy import.
What publishes, and what refuses (RM89)¶
Until 0.6 the publisher demanded weights/annotations/studies.parquet and uploaded those three and
nothing else. Both halves were written when a module meant a SNP core; RM2 made the SNP core optional
in 0.4 and the constants stayed. Measured against the sixteen reference examples on 2026-08-17: seven
could not be published at all, and eight of the remaining nine published an artifact whose
manifest.artifact.files attests parquets that were never uploaded — so the artifact.digest in the
manifest cannot be reproduced from what arrived. Only grch37_build, a bare SNP core with no sidecar
and no 0.4-family table, was correct. Reported as S35 by just-dna-lite, who found the sources.parquet
half of it from the other end: their report footer renders "Not stated" for a module's licence terms
because the table carrying them was dropped at upload.
- The allowlist is derived, never hand-kept.
_ALLOW_PATTERNSisjust_dna_compiler.compiler.ARTIFACT_PARQUETSplusmanifest.json, the two logo spellings andREADME_CANDIDATES. A new table kind therefore reaches the publisher in the commit that adds it. The hand-kept version of this list is the same defect as a hand-keptfieldnames—@fieldnames-from-modelone tier further out, and the consumer named the property they most wanted preserved as "adding a family becomes one edit that discovery and the publisher learn together". - Three positive rules replace the required triple, ordered most specific first so a refusal names the actual fault:
| rule | refuses | why it is not the old rule |
|---|---|---|
the plan carries every file manifest.artifact.files attests |
a deleted parquet; an allowlist that has fallen behind the compiler | the artifact's own attestation is the comparator, so it needs no list of its own |
weights.parquet never travels alone |
a half-finished SNP-core compile | scoped to the weights-led shape — a pharm_variants-led module legitimately has neither companion |
at least one lead parquet (LEAD_PARQUETS: weights + the nine 0.4 families) |
a directory of manifest.json + README with no data |
this is what discovery actually probes to decide a directory is a module |
- The first rule is a self-check as much as a module check, which is why it is worth having on top of a derived allowlist: it compares what would be sent against what the artifact says it contains, so the two can never drift apart silently again.
- An unreadable or absent
manifest.jsonwithholds, it does not refuse. Same tri-state asversion_unknown_reasonbelow, and the same reason: a manifest that cannot say what the artifact contains has said unknown. A directory with no manifest is still publishable, exactly as before. - Nothing that published before stops publishing. The change is a widening on both axes — more module shapes accepted, more files sent — and the only new refusals are the two shapes that were never a publishable module (weights alone; no lead table at all).
A module is published twice, and the second path is the one that can name a release (RM84)¶
On the discovery path there is no version in the path, no manifest fetch and no digest check, so a republished module keeps the same URL and a cached copy shadows it — the only invalidation a consumer has is keyed on its own package version, which makes the identity of "the module changed" a property of the reader. The half this tier owns is the layout, and it is written twice now:
| path | what it means | when it is written |
|---|---|---|
data/<name>/ |
latest — the deployed path, unchanged in meaning, and what discovery scans | always |
data/<name>/v<version>/ |
this release — a subdirectory inside the flat path, not a sibling | when the manifest states a version |
v<version>verbatim, never a barevN.v0.6.0, notv0. A bare major segment throws the rest away, so two patch releases of one module would collide at one path — the defect being fixed rather than a smaller version of it.- A module with no version gets the flat path alone, and says which of four reasons applies.
Identity.versionisstr | Noneand stays null unlessmodule_spec.yamlstates a canonical SemVer (the registry stamps one on publish), so most locally-compiled modules have none.UploadPlancarriesversioned_path_in_repo=Noneplusversion_unknown_reason— nomanifest.json, not readable as JSON, noidentity.version, or notMAJOR.MINOR.PATCH— so a caller reads the reason as a field and--dry-runprints it before anything is sent. Adata/<name>/vNone/directory is never constructed. - One field is read out of the JSON, not the whole manifest through
read_manifest. Validating the fullModuleManifestto reach one string lets an unrelated defect — a hyphen inidentity.name, anicon_setoutside the vocabulary — withhold a version that is right there and legible, and the refusal then names neither field nor value. The SemVer gate isidentity.is_valid_version, the same predicateIdentity.version's own validator calls, which is also what keeps a stray/out of a path segment. - It never refuses.
manifest.jsonis in the allow-patterns and not in the required set, so a directory without one has always been publishable; RM84 is not a licence to tighten what the publisher accepts. An unreadable manifest is a withheld version, not a failed upload. - Two commits, not one — and the docs say so rather than implying an atomicity the code lacks.
upload_foldercommits per call, so a reader can briefly see the flat path refreshed while the versioned copy is not there yet. The flat path goes first, because it is the one anything reads today. If the second call fails, the first has already landed: latest is the new release and the versioned copy is absent until a re-run, which is idempotent. Making it atomic meanscreate_commitover an explicit operation list, a different shape from the allow-pattern plumbing every publish here uses; it was not worth holding the fix for, and if the window ever matters that is the change to make. - A versioned path is only as stable as the author's
version:is. Nothing reads the remote before writing, so re-publishing without bumping the version overwrotedata/<name>/v<version>/with different bytes — the same overwrite the flat path has always done, but under a name that invites caching. Closed in 0.6.1 (RM88): the versioned path refuses unless--force. Before either write, the publishedmanifest.jsonat that path is read and itsartifact.digestcompared with this module's; a different digest raisesPublishCollisionError, and the CLI reports it asALREADY PUBLISHEDrather thanUPLOAD FAILED, because the module is fine and the remote is what disagrees. Four things about the gate that are decisions rather than details: - Identical bytes are not a collision. The comparator is the digest, not presence — a presence check would refuse exactly the re-run this section documents as the recovery when the second commit fails.
- The flat path is not guarded. It means latest; overwriting it is what it is for, and the whole point of the versioned copy is that it is the one that does not move.
- It fails open. An unreadable published manifest is an unknown, and nothing established a
collision, so the publish proceeds with a warning. Failing closed would make a network flake
demand
--forceand train an author to pass it by default. - A recompile under a newer compiler trips it, correctly: P4 scopes byte-reproducibility to a
fixed
compiler_version, so the path really would come to hold different bytes than the ones it was published with. The refusal says so, because "but I changed nothing" is the first thing its first user will think. upload_folderadds and replaces; it never removes — so a republish leaves a union of two releases. A recompile that stops emitting a table leaves the previous release's parquet at the path beside a manifest that does not attest it, and this happens on the flat path, every time. The format's answer is that an unattested file is not part of the module —manifest.artifact.filessays which parquets are, andartifact.digestis a Merkle root over exactly those — so a manifest-first reader never sees it. What stops that being true is the reader: MODULE_LIFECYCLE § 6.8 records that the discovery path fetches no manifest and probes named files, so there a fossil parquet is indistinguishable from a live one and the module reads as the wrong kind.delete_patternswas considered and declined: it cleans nothing on a module nobody republishes, does nothing for a consumer that probes rather than reads, and is one wildcard away from dangerous — HuggingFace filters those patterns withfnmatch, whose*crosses path separators, so a single*.parquetin the allowlist would delete every archived version's parquets. The allowlist is literal basenames today and the archive survives by that accident. The fix that closes it is the reader's, and it is asked explicitly in INTEGRATION_0_6 § 2.8.-
Nothing already published moves. The flat path keeps being written, so every module published under the old layout stays exactly where it is and keeps resolving. That is the argument for writing both rather than migrating.
-
Nothing prunes the versioned copies. The dual write leaves one full artifact set per release in the collection forever, which a consumer mirroring it pays for. Raised by just-dna-lite in S35 as a consequence rather than an objection — it does not affect discovery — and recorded here rather than fixed, because a retention policy is the collection owner's decision and not the publisher's.
Both questions above were answered by just-dna-lite on 2026-08-17 (S35), and the segment spelling is settled:
v<version>verbatim stays. 1. Their scan matches onlyv-plus-integer (^v(\d+)$, compared withint()), sov1.0.0does not match andv10would sort underv9— but the correction that decides it is that the fallback lives only in their generic fsspec branch. HuggingFace has its own discovery branch with no version fallback at all, so on this path no spelling is read today and the segment cannot be chosen to suit one. A barevNwould collide two patch releases at one path and buy nothing, so verbatim it stays. Both halves of their fix are theirs, and they are unscheduled: teach the HF branch a versioned fallback, and replace the regex withjust_dna_format.identity.Version, which already parses and orders. 2. No, and by construction. Both of their discovery branches callfs.lsat exactly one level, neverfs.find, a**glob or a recursive listing, and their probe asksfs.existson named files rather than listing the directory — so a nesteddata/<name>/v<version>/is never enumerated. Verified in their tree by search rather than assumed.Until their discovery half lands, the versioned directory is written-but-not-yet-read: the flat path is unchanged and remains latest, so following it is still correct.
ClinVar reference snapshot (clinvar_build.py, [dev])¶
ClinVar is a second, complementary reference beside the Ensembl snapshot: ~4.4M clinically-curated
GRCh38 records (~200 MB gz), 1.54M of them carrying no rsid, so a clinical module can enrich offline
without provisioning the 14 GB dbSNP cache. Build ([dev], polars) and download (core) split the same
way the Ensembl snapshot does:
flowchart LR
ncbi["NCBI clinvar.vcf.gz GRCh38"] --> build["clinvar_build.py (dev)"]
build --> pq["clinvar/data/*.parquet + release.json"]
pq --> hf["HF datasets/just-dna-seq/clinvar"]
hf --> cache["local cache"]
cache --> link["clinvar.py lookup_loci (core)"]
link --> chain["enrich() chain"]
chain --> res["resolution.csv (source=clinvar)"]
res --> comp["compiler: weights.parquet"]
build_snapshot(vcf, out_dir) turns the NCBI ClinVar GRCh38 VCF into the per-chromosome parquet
snapshot the clinvar link reads (out_dir/data/clinvar-chr{N}.parquet, same layout as the Ensembl
snapshot so one DuckDB view shape serves both). One row per ACGT ALT allele (symbolic/structural and
>50 bp alleles are skipped and counted), with the columns:
chrom, start, ref, alt, rsid, variation_id, allele_id, gene, genes, clin_sig, clin_sig_raw,
review_status, review_stars, condition, molecular_consequence, variant_type, origin
The resolver link reads only chrom/start/ref/alt; the rest is annotation the parquet carries for RM4
(it never enters resolution.csv — orthogonal axes, P5). clin_sig is folded into vocab.VALID_CLIN_SIG
by an explicit severity order (a multi-valued CLNSIG picks the most severe, splitting on |///,
so Pathogenic,_low_penetrance is recognised) while clin_sig_raw keeps the verbatim CLNSIG
(lossless, auditable). Since 0.7 the fold itself lives in clin_sig.py, not here — see below.
A release.json records provenance (clinvar_file_date from the VCF ##fileDate,
source_url, source_sha256, record_count, built_at, builder_version) — the values
clinvar.clinvar_dataset_label turns into the dataset a drafted module's licence row records (RM4),
which is what the clinical cross-check reads back to know it would be comparing a value against itself.
Coordinate convention — no shift. start is the 1-based VCF POS, passed through unchanged; the
Ensembl snapshot uses the same convention, so a variant resolved by either reference lands on the same
coordinate. (ResolutionRow.start's field doc was corrected from "0-based" to 1-based accordingly.)
The parquet is byte-reproducible across rebuilds (rows sorted per chromosome; only release.json's
built_at varies). polars is a [dev], guarded import — the runtime clinvar link is polars-free.
download_clinvar_vcf streams the NCBI VCF with the core httpx (atomic .part rename, sha256 while
streaming). VCF-parsing idioms are leeched from just-dna-lite's v1_port.clinvar.
One significance normalizer, shared (clin_sig.py, 0.7 — RM134)¶
normalize_clin_sig(raw) is the single raw-token → VALID_CLIN_SIG fold, and it moved out of
clinvar_build the moment a second source started reporting a significance. The reason is not tidiness:
a concordance check's entire output is a comparison of two normalized calls, so two hand-written maps
make any drift between them read as a disagreement between the authorities rather than between our own
tables. The module imports VALID_CLIN_SIG and nothing else, so a runtime pass reads it without the
[dev] extra the builders need.
Two defects were fixed on the way out, and neither is visible from ClinVar's side. The map's keys
are underscored because that is how ClinVar spells CLNSIG; PubMind spells the same concepts with
spaces, so two of its six tokens fell through to other:
| Token | before | after |
|---|---|---|
Uncertain significance |
other |
uncertain_significance |
Conflicting |
other |
conflicting |
ClinVar's own Uncertain_significance and Conflicting_classifications_of_pathogenicity mapped
correctly all along, so the two sources would have been reported as disagreeing where they agree — on
the largest disagreeing class in the measured corpus join. The repair is a whitespace→underscore step
in the tokenizer, an identity on every existing key, plus a bare conflicting key. Both registries
are asserted as equalities at import: set(CLIN_SIG_SEVERITY) == VALID_CLIN_SIG and
set(CLIN_SIG_MAP.values()) == VALID_CLIN_SIG.
A composite is still resolved by severity, and PubMind gets ClinVar's answer. Benign/Likely benign
folds to likely_benign, exactly as ClinVar's Benign/Likely_benign does. PUBMIND_ASSESSMENT.md
wrote that mapping as benign while it still proposed a second map; teaching one spelling a different
answer from the other is the drift the single map exists to remove, so the assessment's line is
superseded here rather than implemented.
One more tightening rode along: a whitespace-only or token-less value is not_provided rather than
other. "The source states no classification" and "the source stated something we do not model" are
different answers, and only the second is a disagreement. No ClinVar CLNSIG takes that shape, so
nothing built to date moves.
The clinical cross-check (clinical.py, offline)¶
verify_clin_sig(variants, resolution_rows, *, reference) compares each authored clin_sig against
the ClinVar snapshot's own and returns the disagreements. It runs inside enrich() by default
(--verify-clinsig/--no-verify-clinsig) and lands on EnrichmentResult.clin_sig_conflicts.
Tier note, because the obvious reading puts it in the wrong place. This check needs no network at all — the snapshot is local — yet it belongs here rather than in the compiler. The boundary is not online-vs-offline but does the check need a reference: the compiler is inject-only by charter and holds no ClinVar. Reading "offline ⇒ compiler" would put it in the tier that cannot host it.
Allele-exact, never rsID-level, and the snapshot itself shows why: rs334 at 11:5227002 carries
T>A as pathogenic (2 stars) and T>G as likely_benign (1 star). One rsID, one locus, two
opposite calls. Comparing by rsID would report a module that is simply right. The allele the annotation
is about is taken from effect_allele when set, otherwise from the genotype allele that is not the
reference; when neither pins it down, the comparison falls back to the whole locus and reports only if
no record there supports the authored call.
What counts as a disagreement is coarser than the vocabulary: pathogenic vs likely_pathogenic
is a difference of confidence inside one conclusion, not a conflict, and anything paired with
uncertain_significance/conflicting/not_provided is not a conflict either — ClinVar has no opinion
to disagree with. Opposed calls (pathogenic-class vs benign-class) are the finding worth acting on and
are flagged as such.
A module drafted from this very snapshot is not checked, and the run says so (0.5.2). Where the
clin_sig came out of draft_gene_panel, the comparison is a value against itself: a consumer
measured 27.1 s with the check on and 2.6 s with it off on a 7,818-row panel, byte-identical output,
and 0 conflicts either way — necessarily 0. That zero is the problem rather than the cost: it looks
like evidence and is none.
A pass that contributes nothing records no terms (RM142, 0.7). Every fact pass writes the licence
row for the source it consulted — that rule is unchanged and is what the compile gate reads. What it
means is this module uses this source, so the row belongs to a pass that actually put a row in a
table, never to one that merely ran. The family already worked that way by construction:
gene_metrics, frequencies, assertions and gene_validity pass {row.source for row in out} to
record_source_terms, so an empty pass records nothing. clingen.py built a fixed row and wrote it
unconditionally, and a single-variant module on a gene ClinGen does not curate therefore shipped a
ClinGen obligation and warned about a licence disagreement it had no part in.
The predicate is what this run covered — not the table's contents, which include what an earlier
run merged in and already recorded, and not the absence of missing genes, which would drop the
declaration from any module carrying one uncurated gene beside a curated one. The compiler cannot check
this: its orphan warning exempts the annotation layer deliberately (RM46), so only the pass knows.
ClinGenResult.source_row is still returned whatever happened, because the terms of what was consulted
are a real fact with a different meaning.
The marker is machine-written, not authored (RM4, 0.6). clinvar_draft stamps the release it
copied the rows out of into the dataset column of the clinvar/annotation row it already had to
write in the licence table — clinvar_2026-06-27, from clinvar.clinvar_dataset_label, which prefers
release.json's clinvar_file_date and falls back to its source_sha256. clinical.tautology_reason
recomputes that same label from the snapshot in hand and compares. Both sides call the one function,
so the writer and the reader cannot drift apart — and this drift would be silent, since a disagreement
about the label does not fail, it just never matches.
Widening a panel from a newer snapshot withdraws the label rather than re-writing it.
merge_sources_csv is never-clobber so a curator's hand-written terms survive a re-run, and dataset
inherited that protection the moment RM4 made it load-bearing — leaving the row naming the older
release while half the rows came from a newer one, in the column manifest.sources publishes.
licensing.withdraw_stale_dataset blanks it instead, and only when rows were actually added: a module
carrying two releases has no single release to name, so the honest value is unknown, and an empty
dataset skips nothing. The terms on the row are untouched. Re-labelling to the newer release was the
other candidate and it is the same false claim pointing the other way.
It keys on dataset rather than on the module's panel: block because the claim is provenance —
these rows came from this snapshot — and the tool that copied them is the authority on it. Asking an
author to maintain a declaration whose only reader is one skip is bureaucracy the enricher exists to
remove. panel: is deprecated in 0.6 and reads nothing here any more; a 0.5 module whose pin
matches gets the check run, which is the safe direction. Only an established match skips: no
licence table, a ClinVar row with no dataset, a different release, or a release.json that cannot be
read all leave the check running. The row must be at the annotation layer — enrich() writes a
second clinvar row at the resolution layer for the coordinates it looked up, and a coordinate is
not a copied clinical call.
The release label is half the question; the drafter's digest is the other half (RM73, 0.6). A
matching release says the rows were copied out of this snapshot. It says nothing about whether they
still are — a cell edited by hand after the draft is no longer a copy of anything, and no module-level
fact can see that. RM4 shipped that hole knowingly and put an expensive per-row audit behind strict
to recover what the marker was blind to.
clinvar_draft now also stamps SourceRow.draft_digest: a hash of variants.csv projected onto
(rsid, chrom, start, ref, alts) → clin_sig, the identity cells and the column this check reads.
clinical.tautology_reason recomputes it. Skipping requires both — this release and an
unmoved digest — and either half missing runs the check in full.
| what the module records | what happens |
|---|---|
no licence row, no dataset, a different release, an unreadable release.json |
the check runs |
| this release, no digest (hand-written label, or a module drafted before 0.6) | the check runs |
| this release and a digest that still matches | skipped as a tautology, in both modes |
this release, digest moved — someone edited a clin_sig |
the check runs, over the whole table |
Same in both modes, and that is the RM4 ladder collapsing. strict used to pay for a per-row
look-up because deciding whether a value was still a copy was the look-up; the digest answers it
offline, so there is nothing left for a mode to switch on. EnrichmentResult.clin_sig_audit and its
copied/authored/no_record split are gone — what remains is clin_sig_comparison, carrying
compared, no_record and the conflicts. strict still does not escalate a conflict into a failure;
that is this check's standing exception and it is unchanged.
Why the digest is scoped to the column and not the row. A ClinVar-drafted module always has
edited rows: genotype is a placeholder the human is required to fill, so a whole-row hash would be
invalidated on every drafted module and the skip would never fire once. Scoped to clin_sig, filling
the stub leaves it alone and editing the call moves it — exactly as sensitive as the question.
The limit that remains, stated rather than left to be found. The digest covers the whole table, so it means no checked value has changed since the drafter last wrote. A row hand-authored before a later re-draft is covered by the new stamp along with the drafted rows, and escapes the check. That is strictly narrower than what it replaces — the module-level marker let every hand-edit escape — and any later edit re-enables the check.
The same mechanism, the same way, for the other two providers. pgx_draft writes
function_status out of CPIC and pgx._function_conflicts compares that column against CPIC;
clinpgx_draft writes evidence_level out of ClinPGx and the ClinPGx check compares that. Both were
tautologies nobody had filed, both publishing a structurally guaranteed findings=0. In enrich_pgx
the skip is per leg: the CPIC leg records tautology while PharmVar, an independent authority,
still runs and the record aggregates both.
The skip carries its reason on EnrichmentResult.clin_sig_not_checked, because an empty
clin_sig_conflicts says two opposite things on its own ("compared everything, nothing disagreed" and
"never compared"), and a consumer reading the first when the second happened has been told a check
passed that was never put. Its values are not_requested (the author's own --no-verify-clinsig),
no_snapshot, unusable_snapshot (present but not queryable — compare_clin_sig returns None rather
than a comparison of zeros), the tautology sentence, or None when the check really ran. Where a human
typed the clin_sig, nothing changes — that is the case this check exists for.
The concordance record (concordance.py, offline) — RM130 + RM134 § B¶
The check above counts its findings and, until 0.7, kept none of them. clin_sig_concordance(variants,
resolution_rows, *, reference, pubmind_reference=None, sources=None, spec_dir=None, checked_at=None,
clinvar_comparison=None) in clinical.py consults every authority and returns a ConcordanceRecord
whose two tables concordance.write_concordance_tables puts beside the spec. enrich() calls it once
per run and commits both tables at the gate, so a refused strict run leaves none behind.
| file | key | carries |
|---|---|---|
clin_sig_concordance.csv |
(variant_key, genotype) |
authority_concordance, authored_position, opposed, and the module's own call |
clin_sig_authority_calls.csv |
(variant_key, genotype, authority) |
that authority's normalized clin_sig, its raw token, and its confidence in its own units |
None rather than two empty tables when no authority could be consulted at all, and the
distinction is the whole tri-state: two empty tables are a claim (nothing here is contested), while
None says the question was never put. A caller must not write a file that says the first when the
second happened, and enrich() writing nothing on None is what leaves an earlier record readable
instead of overwriting it with a comparison nobody made.
The three-way check subsumes the two-way rather than running beside it. With no PubMind snapshot
that authority's call reads unchecked on every subject, and the degenerate case is exactly the
ClinVar-only finding: the same subjects are contested, the same conflicts are logged in the same
words, and no author meets one disagreement twice. What changes is what the record withholds —
authority_concordance reads unchecked rather than single, because one authority speaking while
another was never asked is not corroboration.
Each authority is a leg, and a leg has three states (clinical.AUTHORITY_LEG_STATES):
| state | means | contributes |
|---|---|---|
consulted |
a snapshot answered | that authority's call per subject |
unchecked |
no snapshot, or one present that would not answer | an unchecked call per subject — never an absence, never agreement |
tautological |
this module's rows were drafted out of this very snapshot and have not moved | nothing at all; the reason is on the leg and reported separately |
AUTHORITY_ORDER is (clinvar, pubmind) and it is walked rather than restated: it fixes the order
the detail rows come out in, which matters because the table becomes a parquet whose bytes depend on
row order.
PubMind's several PVIDs over one allele fold to one call, and the camp guard runs first. The
detail table is keyed (variant_key, genotype, authority), so a subject carries one row per
authority and something has to stand for the set. Where the records straddle the pathogenic/benign
line there is no representative call: fold_authority_records answers conflicting, the
vocabulary's own word for it, which sits in the undecided camp and so opposes nothing. Folding by
severity there would silently answer pathogenic — a winner picked by an ordering nobody defined.
Within one camp the fold is the shared normalizer's own severity rule, the same one that resolves
a composite token, so Benign and Likely benign in two rows fold exactly as Benign/Likely benign
does in one cell. Confidence is withheld unless exactly one record stands behind the call: PubMind's
0–3 count is per record, and any function of two of them is an arithmetic nobody defined. The
multiplicity is counted on the record (multi_record_subjects, internally_contested) rather than
discarded.
Not merge-not-clobber, and it is the one derived sidecar that is not. Every other one gap-fills
because a recorded row might carry a curator's judgement; this one carries none, since the judgement
about a contested subject goes in overrides.csv. Merging would be actively wrong: a subject the
archive stopped contesting has to leave the record, because a conflict that stops being reported is
exactly how an author learns the archive caught up with them.
classify_concordance is a pure function over N authority outcomes, and it knows nothing about
ClinVar. That is what makes the arity property testable at the three and five authorities the
vocabularies were designed against rather than at the one the producer reaches today. CLIN_SIG_CAMP
moved here from clinical.py for the neighbouring reason: the record and the two-way check must draw
the opposed-versus-differing line in the same place, and two maps for one distinction is how a drift
in our own code comes to read as a disagreement between two archives.
The tautology skip reaches the record too, and since 0.7 it is decided per LEG. Where a module's
clin_sig column was drafted out of a snapshot it would be compared against and has not moved since,
that comparison is a value against itself and is guaranteed to find nothing — so writing two empty
tables would publish nothing here is contested on no evidence, which is the findings: 0 defect
wearing a new file. But skipping the whole check would be the opposite error once there are two
authorities: a module drafted from ClinVar still gets a real comparison out of PubMind, which copied
nothing, and throwing that away to suppress the hollow half is what enrich_pgx already learned.
So clinical.is_tautological_leg(sources, authority, dataset, spec_dir) states the conjunction once —
the licence row names this release of this source, and the drafter's digest over the
checked column still matches — and every leg calls it with its own label. The checked column comes
from DRAFT_PROJECTIONS, derived rather than restated, so a drafting provider added later says which
cell it copied without this module learning about it. tautology_reason is the ClinVar instantiation
and keeps its shipped wording, because a warning's text is an API. Omitting sources/spec_dir
establishes nothing and the comparison runs: an unknown is never a permission to skip.
The subject list is built from the comparison that already ran, never from a second pass over the snapshot. A subject that produced a conflict is rendered from that conflict's own record, so the verdict and the finding are one fact told twice rather than two computations that can drift — and re-asking would cost the whole comparison again.
ClinSigConflict no longer names its authority in a field name. It carried clinvar: str until
0.7; authority and authority_clin_sig replace it, with clinvar kept as a read-only alias so an
existing caller keeps working. The rename is the point: a finding whose field is named after one
archive costs a rename — major-only work — the moment a second one arrives.
One thing the record makes visible that the two-way check cannot. At a single authority every
reported conflict is opposed by construction, because the check only reports where both sides are
opinionated and their camps differ, and pathogenic/benign are the only two opinionated camps. So
the differing-but-not-opposed case had no producer at N=1 — _clin_sig_detail's second group was
unreachable. It became reachable the moment PubMind arrived: two authorities can disagree with each
other while neither contradicts the module, which is discordant + matches_some and is what the
record was shaped for.
Severity: warning-tier in both modes, escalating in neither (@clinsig-never-escalates), with
more force at two authorities than at one. A disagreement with a literature miner's aggregate over
the field is a statement about that extraction's limits at least as often as about the module — the
measured corpus join agreed 62 % of the time — so discordant is a fact about the field rather than
a defect to gate on. clinical.concordance_sentences produces the warnings, each carrying its
denominator and the authorities that answered; clinical.concordance_notes produces the info-tier
lines for the legs that did not run, because a leg that silently did not run reads as a leg that
found nothing. A run that found nothing contested emits neither: a check that cannot fail reports no
zero, and neither does one that ran and found nothing. The two warning stems a consumer greps, pinned
by test because a warning's text is an API:
{n} of {m} subject(s) put to the authorities are contested, {k} of them on opposed calls
{n} of {m} contested subject(s) are a disagreement between the authorities themselves
The leg notes are info-tier and deduped against the two-way skip's own sentence — the ClinVar
tautology's prose comes out of tautology_reason on both paths, so a drafted module would otherwise
read it twice in one run. Matched on the sentence, not on the skip key, so a reword stays deduped.
Nothing resolves a split, at two authorities or at five. There is no majority, no consensus
call and no resolved winner, and module_spec.yaml's optional authority_precedence: is recorded
and computed with by nothing — it says whose call the curator weighted while deciding, so a
consumer can see the stance, and no tier reads it. Choosing between a declared order and a majority
needs a weighting model this workspace has declined to invent three times.
Has the disagreement you answered moved? (RM151, 0.7)¶
RM117 shipped the first of two signals about an answer an author has already recorded: a subject that has left the record, which the compiler computes offline because the record holds contested subjects only and is rewritten whole. This is the second, and it lives here because it needs the archive's value now against its value at record time — the first of those requires a consultation.
An overrides.csv row against clin_sig_concordance.csv is a judgement about a particular
disagreement: the archive said X, the author says Y, and the reason column explains why. If the
archive later says Z, that reason was written about a value that is no longer there. Nothing in the
record distinguished a justification that still describes the disagreement on file from one that
describes a disagreement since replaced by a different one.
The baseline is the previous run's clin_sig_authority_calls.csv, and it is the only one this
format keeps. That table records what each authority actually said — clin_sig, the verbatim
clin_sig_raw, and the dataset it came from — keyed (variant_key, genotype, authority), so the
comparison is recorded-call against fresh-call per authority. Do not expect this for a table that
records no prior value (@probe-names-the-table): an overlay row against frequencies.csv or
resolution.csv has no recorded baseline at all, so a general the value moved check would be
answerable for one table and silently absent for every other. The finding names its table for that
reason.
It is read before the commit, and that ordering is the whole feature. write_concordance_tables
replaces the file whole, so the previous run's rows exist only until this run commits.
clinical.answered_call_shift is called in the staging phase — the same phase every other product of
the run is computed in, for the unrelated reason that a refused strict run must change nothing — and
an AST guard pins the read above the write, because no assertion over a return value can see statement
order.
A move is therefore observable exactly once, by the run that notices; the run after that compares
against the new baseline and is silent. That is a limit rather than a bug, and it is the honest shape
for an observation: binding an overlay row to the value it justifies is a mechanism this format does
not have, and RM117's three objections to giving outranks a severity consequence all turned on
exactly that missing binding. None of them is an objection to noticing that the value moved.
Three states, and the third is what the check is for. A call recorded on both sides with the
same classification is unchanged, dataset moved or not — a re-released archive saying the same thing
has not moved the disagreement. A call that moved, or that went recorded → no_record (or back), is a
shift. Everything else is withheld and reported as a note rather than as agreement:
no_prior_record (this run is writing the first record, the previous one would not parse, or the
authority was unchecked when the answer was written) and unchecked_now (nobody could ask this
run). Neither ever reads as nothing moved.
A move this tier's own normalizer made is reported apart from the archive's. When both sides
recorded the same verbatim clin_sig_raw and only the normalized member differs, what changed is how
clin_sig.py reads the token — a fact about our code, with nothing for an author to do. Folding it in
with the archive's revisions would accuse a source of a change we made.
A subject that left the record entirely is not this finding. It is overlay_answer_vindicated,
which the compiler reports as good news, and hanging a second and gloomier finding on the same overlay
row is exactly the already firing with the wrong words failure RM117 was.
Warning-tier in both modes, escalating in neither, with more force than the record itself: nothing here is even a disagreement, it is a note that the record an author reasoned over was rewritten underneath their reasoning, and gating on it would make an artifact refuse over an archive's release schedule. Silent on a module with no overlay answers, which is every module today. The two warning stems, pinned by test:
{n} of {m} answered subject(s) rest on an authority call that has moved since the answer was recorded
{n} answered call(s) differ only after this release's own normalization
The wording observes and never adjudicates, and that is a test rather than a convention: a
word-boundary grep refuses correct, wrong, mistaken, vindicated, confirmed and their
siblings in every message. The disagreement you answered is not the one on record now is a statement
about the record; your answer may be wrong is a verdict, and this format does not put a verdict under
a check that cannot see the reasoning.
PubMind as the second authority (pubmind.py, offline) — RM134 § B¶
pubmind.py is the runtime reader over the snapshot pubmind_build writes: duckdb, in the ordinary
install, mirroring clinvar.py's split against clinvar_build.py (@duckdb-vs-polars).
lookup_pubmind_calls(reference, alleles) answers what does PubMind say about this allele, and
pubmind_dataset_label(reference) reads the release label back out of release.json.
There is deliberately no lookup_loci beside it. PubMind's coordinates are PyEnsembl
back-mappings of text an LLM extracted, so they are annotation and never resolution: nothing it
produces may enter resolution.csv, whose authority column is a different word for a different
thing (@source-vs-authority).
The snapshot is located by locations.resolve_pubmind_reference — explicit path →
$JUST_DNA_PUBMIND_CACHE → the default cache directory — and there is no ensure_pubmind_snapshot
to pair with it: the snapshot is operator-built, pubmind publish refuses, and a missing one is
unchecked rather than a download a run should attempt. enrich(pubmind_cache=…) overrides the
lookup, and just-dna-enricher enrich --pubmind-cache <dir> is the same override from the
command line. The flag was owed for a release: § B built the check while cli.py belonged to
the sibling lane, so the second authority shipped reachable only from Python or the
environment variable.
PubMind snapshot (pubmind_build.py, [dev]) — RM134 § A¶
PubMind (Wang & Wang, Nat Commun 2026, doi:10.1038/s41467-026-76834-4) extracts
variant–disease–pathogenicity assertions from 41.7 M abstracts and 5.4 M full texts with an LLM. It is a
source, of the same kind ClinVar is: an authoritative annotation source. Nothing it produces may
enter resolution.csv — its coordinates are PyEnsembl back-mappings of extracted text, and
resolution.csv's authority column is a different word for a different thing.
There is exactly one per-variant channel. The web API has two endpoints, neither takes a variant,
and both state that per-record detail is withheld; the per-record LLM_reasoning and evidence passages
— the genuinely novel part — are reachable only through CHOP's licensed full database. So the input is
the ANNOVAR-redistributed hg38_pubmind_db.txt.gz, whose columns are VCF-style despite the ANNOVAR
packaging: no - alleles, a one-base deletion written 1 1014264 1014265 CC C with the anchor base
retained, and Start the 1-based POS. A join needs no coordinate translation.
build_snapshot(table, out_dir) writes one out_dir/data/pubmind.parquet plus release.json:
The names are unprefixed on purpose — clin_sig, not pubmind_sig. The ClinVar snapshot already
uses them, the source is the file, and one column vocabulary across every snapshot is what lets a check
read N authorities with no per-source mapping. pathogenicity_score is nullable and null means not
computed, never 0.0; confidence is PubMind's own 0–3 evidence-depth count and is deliberately not
normalized against ClinVar's review_stars — different instruments, and folding them into one number
would be three axes in one field.
Two thirds of the file is not a genotypable position, and every dropped row is counted. When PubMind
recovers only a protein change from the text it back-maps through the transcript and writes out every
codon that could encode it. derivation records what survived:
derivation |
What it is |
|---|---|
direct |
the source row was already one base against one base |
codon |
an equal-length block differing at exactly one position, decomposed onto that base |
indel |
a length-changing row, kept but marked — upstream left-normalization is unverified |
and PUBMIND_DROP_REASONS records what did not: off_target_chrom, non_acgt (16 rows in the
2026-08-24 file whose alt is 0 or N), ref_equals_alt (523), no_pvid, unparsable_position,
multi_substitution (a block needing two or three simultaneous changes — a statement about the
protein, not a position), and identical_duplicate (two codon rows decomposing onto one row identical
in every column). The registry is walked, so input_rows == record_count + sum(dropped.values()) is an
equality over it: silent truncation reads as full coverage.
no_pvid earns its place twice over. The PVID is PubMind's record id, and the whole snapshot is
organised around record identity — a verdict with no id cannot be attributed, cannot be deduped against
its twin, and cannot join the multiplicity accounting; carrying one as a null would also merge distinct
records under identical_duplicate, which compares whole rows. A bad cell is different from a bad
row: a malformed pathogenicity_score or confidence withholds that one value as null and counts it
in unparsable_score / unparsable_confidence, keeping the row. NaN and inf count as unparsable
rather than being stored, because float() accepts both and they would otherwise poison every
comparison downstream; a non-integral confidence is withheld rather than truncated, since 2 is a
definite count the source did not state.
A contested coordinate keeps every PVID as its own row, and that is the finding. Consolidation into a
PVID is keyed on the text the model extracted, never on a coordinate, so one physical variant fragments
into many records whose verdicts disagree — at chr6:26092913 (HFE C282Y) the table holds eight, split
across four verdicts, one of which pairs rs1800562 with TMPRSS6 and is exactly the shape
_gene_locus_conflicts catches. Collapsing them would mean choosing a winner by an ordering nobody
defined, which is mode() over an unsorted group. release.json records multi_pvid_keys,
max_pvids_per_key and contested_keys — the last being how many of those coordinates actually
disagree, which is why the multiplicity may not be tidied away.
The parquet is byte-reproducible across rebuilds (rows sorted by chromosome in karyotype order, then
start, ref, alt, pvid); only release.json's built_at varies. release.json also carries the
source's sha256 and its ETag and Last-Modified — all three available, all three recorded, so an
upstream revision becomes a finding rather than a silent change of answer — plus redistributable: false
and the reason. A local --table establishes neither header nor the URL, so all three are recorded
as null: unknown rather than absent, and never a default the build did not establish.
download_pubmind_table translates httpx into PubMindUnavailable, a subclass of
PubMindBuildError, so one except arm catches a moved bulk URL as well as a malformed table while a
caller who needs to tell an outage from bad data still can (@client-exception-contract; the subclass
makes a caller's except order load-bearing). polars is a [dev], guarded import.
MANE snapshot (mane_build.py, [dev]) — RM168¶
MANE (Matched Annotation from NCBI and EMBL-EBI) publishes one agreed transcript per protein-coding
gene, matched base-for-base between a RefSeq and an Ensembl accession. This workspace already leaned on
it before it was code: the CIViC identity protocol pins a numbering frame with
MANE.GRCh38.v1.5.summary.txt.gz, "downloaded once and cited" — a procedure step in a probe document,
where every other reference table here is a cache with a location and a recorded release. The 33 curated
name→identity answers that shipped with that protocol were derived in a frame nothing in the code could
read. This builder is that frame, cached and pinned.
just-dna-enricher mane build --download --out data/caches/mane # discover the newest version, then pin it
just-dna-enricher mane build --release 1.5 --download --out data/caches/mane
just-dna-enricher mane build --summary … --changed … --not-in-mane … [--versions README_versions.txt]
export JUST_DNA_MANE_CACHE=data/caches/mane
MANE is the default, not the answer — and the bound is measured, not asserted. The summary carries
CDKN2A twice for one GeneID: NM_000077.5 / ENST00000304494.10 marked MANE Select beside
NM_058195.4 / ENST00000579755.2 marked MANE Plus Clinical, two CDS numbering frames for one gene.
Plus Clinical is a fraction of a percent of the rows, which is exactly the argument: a case occurring in
a third of a percent of genes is precisely the case a remembered accession hides and a table shows. So
MANE_status is carried as a column and never collapsed — a builder keeping one row per gene would
drop those rows and reintroduce the blind spot the item closes.
RUNX1 is the other half of the bound, and it confirms the protocol rather than fixing it. RUNX1 is a
single row, and the 27-residue RUNX1c/RUNX1b offset the identity protocol derived by translating each
isoform's CDS is not in MANE and cannot be. The table therefore makes the CDKN2A class of problem
visible and is silent on the RUNX1 class, and a pass that treated it as an oracle would be wrong in a
way the file itself cannot warn about. That sentence is in the module docstring, in release.json's
notice, and printed by the build command, because it is the one thing a reader of this lane has to
carry away.
Scope is a transcript-identity aid. Generating c./p. notation is a separately deferred feature
with its own unanswered questions, and nothing here proposes it.
Three files, and the other two are why this is a snapshot rather than a download¶
Table (data/…) |
What it answers |
|---|---|
summary.parquet |
the frame, and a cross-map: NCBI GeneID, Ensembl gene, HGNC id, symbol, both nuc/prot accession pairs, MANE_status, and GRCh38 coordinates on the NC_ accession |
changed_select_accessions.parquet |
every gene whose MANE Select moved, the release it moved from, and Update_Affects_CDS — the numbering-frame axis, stated by the source |
protein_coding_genes_not_in_mane.parquet |
the genes MANE deliberately has no answer for, each with a reason |
The second is the currency check, and taking it in the same pass as the summary is the point:
shipping the cache without the thing that notices it going stale is the defect this item is about. A
MANE Select change that moves the CDS moves every c. and p. derived in that frame; one that does
not, does not — and Update_Affects_CDS says which. Old_MANE_Version spans several releases back, so
each row also says how long that gene's frame had been stable. A gene absent from that table has had
a stable frame, which is a positive statement the cache can make and a memory cannot; RUNX1, CDKN2A
and VHL are all absent from it.
The third is @unreachable-not-absent served by the source. "MANE has no answer for this gene" becomes
distinguishable from "nobody asked", with the reason attached — and pending MANE review is a third
state on its own, neither absent nor decided. The reason vocabulary is derived from the file and
recorded in release.json's excluded_reasons; it is not written down beside it, because a roster
stated in prose is a registry nothing iterates (@registry-completeness). The test asserts the counts
as an equality against the fixture rather than against a list of strings.
Update_Affects_CDS is stored twice, the split clin_sig / clin_sig_raw established: a nullable
boolean a consumer queries, and the source's own token so the mapping stays auditable. A cell that is
neither Yes nor No withholds the boolean and lands in unparsable_update_affects_cds — the source
said something we cannot hold, which is a different finding from the source saying nothing.
unparsable_gene_id and unparsable_coordinate are the same shape for the other two cells that can
be malformed, and every one of them is in release.json.
A bad cell keeps its row; a bad field count does not. Everywhere above, a cell the column cannot
hold withholds its own value and the row survives, because the rest of the row is still the source's
own statement. A ragged row is refused outright (@ragged-csv-row): the cells past the break are
shifted, so a short summary row would enter the numbering frame carrying null accessions and read as
coverage, and a shifted one would put the wrong accession under the right gene. Whole-file damage is a
structural failure, and refusing is what a builder may do about one. _open_table also translates
gzip/decoding failures into ManeBuildError, because the opener is chosen on the .gz suffix and a
file that is named .gz and is not one — a proxy that decompressed the download and kept the name —
otherwise escapes past the CLI's handler as a traceback. EOFError is named beside OSError there on
purpose: a truncated gzip raises it and it is not an OSError.
The version is pinned, and release.json copies rather than restates¶
release_1.5/ back to release_0.5/ all exist, so a pin is a URL rather than a hope. current/ is
read exactly once, for 96 bytes, to discover which version is newest — discover_current_release()
reads its README_versions.txt, and everything is then fetched from release_<version>/. Pinning the
mutable path and hoping is what the item is complaining about.
README_versions.txt publishes MANE Version, NCBI RefSeq Annotation Release and Ensembl Release,
and two of those three appear in no filename, so the builder copies the file into release.json's
versions block instead of parsing a name — parsing one would reconstruct less information than the
source hands over (@probe-the-real-file). It is parsed generically, label by label, so a line MANE
adds travels through. dataset is mane_grch38_v<version>, and it is None when no
README_versions.txt was read: an unknown release, never one reconstructed.
A local build asserts no URL it did not fetch. source_url, source_etag and
source_last_modified are per input and null on a local build, and release_url names the directory
actually fetched from. A README_versions.txt handed over on disk establishes which release the files
claim to be; it does not establish where the bytes came from. release.json is written atomically
(@atomic-sidecar-write) — a truncated one parses as valid JSON that is simply missing keys, and
locations.read_release would believe it.
The two modes do not mix, and the CLI says so rather than picking: --release without --download
is refused (that would be our claim about somebody else's bytes, where --versions is the source's
own), and --versions with --download is refused too, because the download fetches the release's
own copy and would silently overwrite the flag. A flag that is quietly ignored is worse than one that
is turned away.
The source spells its own key two ways — GeneID:1029 in the summary, a bare 1029 in the other
two — so parse_ncbi_gene_id normalizes both into one ncbi_gene_id integer and the test runs it over
the raw cells of both dialects (@one-normalizer-two-spellings). Without that, the cross-table question
has this gene's frame ever moved? is a string comparison that silently never matches. Everything else
is verbatim: MANE_status keeps its spaces, GRCh38_chr stays the NC_ accession, and the version
cells stay strings because 0.91 is a MANE release rather than a number.
The parquets are byte-reproducible across rebuilds (gene id, then accession — a total order, with
unreadable ids sorting last because None cannot be compared to an int); only release.json's
built_at varies. Three schemas share data/, on the CPIC precedent — the glob-union hazard
(@snapshot-layout-locations) belongs to readers that union data/*.parquet into one view, and this
snapshot ships no reader, so anything reading it later reads by filename. resolve_mane_reference
deliberately does not accept a bare .parquet the way the constraint resolver does: one file here
is a snapshot missing the currency check and the negative roster.
No --use flag and no publish command. NCBI states a policy rather than a licence: it "places no
restrictions on the use or distribution of the data contained therein" and, in the same paragraph,
"cannot provide comment or unrestricted permission concerning the use, copying, or distribution of the
information contained in the molecular databases." No restriction imposed is not permission granted, so
MANE_TERMS records None on every axis (@no-named-licence) — and a declared-use gate fed an unknown
skips every build unconditionally, which is a flag that does nothing
(@acquisition-gate-is-not-a-read-gate). MANE is a joint NCBI/EMBL-EBI product and only NCBI's side
was read; nothing here asserts anything about EMBL-EBI's terms for the same tables
(@probe-names-the-table). download_mane_file translates httpx into ManeUnavailable, a subclass
of ManeBuildError, so one except arm catches an unreachable FTP site as well as a malformed file
while a caller who needs to tell them apart still can (@client-exception-contract). polars is a
[dev], guarded import.
One thing this lane does not do: it writes no SourceRow. A builder is not a pass over a spec
directory — nothing here consults MANE on behalf of a module, so there is no module to account for
(@write-the-sourcerow, the converse half). The terms constant exists so that the first pass which does
consult it has one to record.
Note for anyone reaching for MANE next: gene located on mitochondrial genome is a MANE exclusion
class, so the mtDNA lane has no MANE frame by construction. And the summary is adjacent to
identifiers.py's HGNC registry and to the Ensembl lane — worth reading before somebody builds a second
gene cross-map beside it.
Gene–disease validity (gene_validity.py, online only) — RM24¶
enrich_gene_validity(spec_dir, *, source, mode, offline, write, export_text, url) takes the gene
column of variants.csv and writes gene_validity.csv: one row per (gene, disease, mode of
inheritance, submitter). just-dna-enricher gene-validity spec/ [--source clingen|gencc].
The question gene_metrics.csv cannot answer. Constraint says how intolerant of variation a gene
looks in a population sample; ClinGen's dosage rating says whether losing a copy causes disease;
neither says whether variation in this gene causes this disease, which is the claim a clinical
module most often rests on without recording anywhere.
Two submitters, and they are different kinds of thing. ClinGen publishes expert-panel curations — one assertion per (gene, disease, MOI), each from a named Gene Curation Expert Panel working to a numbered SOP. GenCC publishes an aggregate of nineteen submitters, ClinGen among them, plus Orphanet, PanelApp and several laboratories; the same gene–disease pair routinely carries several submitters at different strengths, and that disagreement is the data.
Three things established by reading the real files (2026-08-13 downloads: ClinGen 3,659 rows, GenCC 30,410), each of which decided part of the shape:
- Mode of inheritance is part of the key. 59 (gene, disease) pairs in ClinGen carry two rows
differing only there.
(gene, disease, moi)has zero collisions;(gene, disease)silently keeps one curation and drops the other. submitteris in the key too, or GenCC collapses to one arbitrary opinion per pair — the bare-triple mistake the ClinPGx cross-check paid for once already.- The two vocabularies disagree in spelling and agree in meaning, so both are mapped onto
vocab.VALID_GENE_VALIDITY/VALID_INHERITANCE_MODEat this boundary:DisputedandDisputed Evidenceare one member,ADandAutosomal dominantare one member. A consumer filtering on one spelling would silently miss the other's rows. The submitter's wording survives inclassification_raw, so the mapping stays auditable, and a wording this release does not model is left unset with one aggregated warning naming the distinct values — never one line per row.
A gene the submitter has not curated gets no row, and is reported in missing. That is
clingen.py's rule and its reason: a curating body's silence means nobody has assessed the gene yet,
which is not a fact about the gene, so a not_found row would state one. (The ClinVar pass below goes
the other way, because ClinVar covers the genome — "asked and absent" really is a fact there.)
--offline is a no-op with a warning (skipped_offline), never a failure: neither submitter ships
a snapshot, and both files are small enough to fetch whole. An injected export_text= still wins,
because handing over bytes you already hold is not egress.
Both sources are CC0, so a module using this table stays sellable — GENCC_TERMS joins
CLINGEN_TERMS in licensing.TERMS_BY_SOURCE, and the pass records its SourceRow at the new
gene_validity layer. GenCC's attribution names the contributing sources as well as GenCC, because
crediting only the aggregator credits nobody who did the work.
HPO ships no route, and both reasons came from probing rather than from taste. Its release declares
terms:license https://hpo.jax.org/app/license; that URL answers HTTP 404 with a JavaScript shell, and OBO Foundry records the licence as a bare labelhpowith no SPDX id — so the terms cannot be established from any machine-readable source, and an unestablished permission is not a permission (the PharmVar rule). Recording it withcommercial_use=Nonewould flip every carrying module's manifest verdict to undetermined, which contradicts the sellability this table was designed to keep. Separately, the file the item named —genes_to_phenotype.txt— is gene × HP feature × frequency, a different grain from this table entirely;genes_to_disease.txtfits structurally, but itsassociation_type(MENDELIAN / POLYGENIC / UNKNOWN, 8,288 of 15,944 rows UNKNOWN) is a mechanism class, not an evidence grade, and putting it inclassificationwould overload the axis (P5). The row shape holds an HPO row perfectly well; what is missing is a link this tier may take the data over.
Clinical assertions (assertions.py, offline capable) — RM25¶
enrich_clinical_assertions(spec_dir, *, mode, offline, clinvar_cache, download, write) reads
resolution.csv and the ClinVar snapshot and writes clinical_assertions.csv: one row per (allele,
archive record) carrying the clinical call, ClinVar's own review wording, the 0-to-4 star rating and
the VariationID. just-dna-enricher assertions spec/.
The number this workspace was already computing and discarding.
clinical.ClinSigFinding.confidence rendered the star rating into a warning string and kept nothing;
clinvar_draft.draft_gene_panel used it as a filter (default 2 — multiple submitters, no conflicts)
and kept nothing. So a compiled module flattened a one-star single submission and a practice guideline
to the same clin_sig, and every consumer that wanted the difference re-derived it. A number
recomputed downstream is a place to drift (RM40/RM41, a fourth time).
It records; it does not adjudicate. clinical.verify_clin_sig — comparing the author's
clin_sig against ClinVar's — is untouched, and still warns in both modes on purpose, because
failing would make the format arbitrate a clinical dispute. Escalating that check stays parked
deliberately, and nothing in this pass moves it.
Four mechanics worth keeping straight:
- Structured like
enrich_frequencies, because the input is the same. It consumesresolution.csvrather thanvariants.csv: a clinical record is per allele at a coordinate, and the resolution table is where an rsID has already becomechrom-pos-ref-alt. That also sidesteps the multi-allelic-rsID problem — one rsID at one locus legitimately carries a pathogenic, a benign and an uncertain allele (rs33922842in HBB), so looking clinical significance up by rsID would manufacture disagreements out of ClinVar agreeing with itself. - Snapshot-first and fully offline-capable, unlike the frequency pass: ClinVar ships as a snapshot,
so with one provisioned this pass never touches the network. With none found and not
offlineit callsdownload.ensure_clinvar_snapshot— theenrich()shape,--offlineas the only switch — and with none reachable at all it is a no-op with a warning (skipped_no_snapshot), leaving any existing table as the pin. datasetis the snapshot's own release, read fromrelease.json(the RM38 rule): a consumer must be able to tell a pinned file from whatever happened to be current, and a re-review is only visible against a stated release. A snapshot that cannot state one getsclinvar_unknownrather than a fabricated date.- A coordinate on another build is never queried. ClinVar's lookup key is
(chrom, start, ref, alt)and carries no assembly, so a GRCh37 coordinate is a well-formed query returning a different variant's clinical call under this module's key — the same failure the frequency pass had against gnomAD. Such rows are reported inoff_build, kept apart frommissingbecause nobody asked about them, and are outside thestrictgate: a coordinate on another assembly is reproducibly out of this snapshot's reach, so refusing would make a GRCh37 module uncompilable for a reason no authored edit could fix.
clinvar_build.review_stars is the one place the CLNREVSTAT-to-rating convention lives, and it became
public and tri-state in 0.6: None for a record that states no review status and for a wording
this release does not model, 0 for ClinVar's own "no assertion criteria provided". It used to answer
0 to all three, which files an unread record under the weakest rating available — a claim nobody
made. That is also why ClinicalAssertionRow stores the rating as a column instead of deriving it
from the prose beside it: the derivation is a ClinVar convention, and Principle 2 keeps source
conventions out of the schema tier entirely.
The literature pack (literature.py, online only)¶
Pass 4: a module's citations in, literature.csv out. Three questions of decreasing coverage — does the
citation exist (PubMed esummary), do the identifiers agree (DOI/PMCID arrive in the same response),
and does the quoted passage appear in the article (Europe PMC fulltext, open-access subset only) — plus
the article's own licence, which arrives in the same Europe PMC response.
studies.csv is one citation site of several, and this pass reads every one (RM47, RM132). A
pmid on a binning row grounds the threshold it sits on; one on a pharm_variants.csv row grounds
that row's own drug and genotype claim, which studies.csv structurally cannot, since a study row
keys on (variant_key, pmid) and attaches to the whole variant. A module whose only citations come
from those tables is enriched exactly like one with a studies.csv; a module with none at all is
refused, since the relaxation is about where a citation may live and not about whether one is needed.
Those pointers are read through just_dna_compiler.load_citing_rows / table_citations — public for
the RM41 reason, because the alternatives were importing a private symbol or hand-keeping a second list
of the kinds here, and that list goes stale the next time a model declares a pmid. The compiler's set
is derived (_CITING_TABLE_KINDS is every table kind whose model declares the column), so this
tier gains a new citation site by doing nothing, and a test walks this module's own source to assert no
such roster is kept here. load_binning_rows/binning_citations still exist and still mean the
binning kinds only — a caller asking for those is asking about thresholds.
A citation reaching the pass only from such a row contributes no quote and no authored DOI (neither row
has those columns — provenance_quote deliberately did not follow the pmid to either site), so it
reads as nothing to check rather than as an unretrievable fulltext.
The article's licence is recorded per article, and there is no pubmed row in the licence table
(RM46). The pass writes source="pubmed" into every row it produces and TERMS_BY_SOURCE has no
entry for it — deliberately, and permanently: a literature source's terms are per article, not per
source. PubMed's metadata is one thing; the article belongs to its publisher, and Europe PMC's open
subset spans CC-BY, CC-BY-NC and bronze, so one pubmed row would be right for a module citing only
ids and a false all-clear for one carrying a provenance_quote lifted from a CC-BY-NC article — wrong
in the dangerous direction, since that quote is publisher text in the module's own annotation layer.
Four mechanics:
licenseis stored verbatim as Europe PMC spells it (cc by,cc by-nc,cc by-nc-nd— probed over 100 records on 2026-08-13), andlicensing.article_termsmaps it to the three rights at read time, so a mapping correction reaches rows already written. Same rule ascpic_build.- The licence is independent of
is_open_accessand is not derived from it: PMID 28546431 comes backisOpenAccess: Nwithlicense: cc by, because the flag describes Europe PMC's OA subset while the licence describes the article. - Three orthogonal axes, and
Noneis neverFalse. CC BY-NC forbids sale and expressly allows sharing, which is whyredistributionis its own column; a licence this tier has not read leaves all three null rather than guessing in either direction. - Quoting a non-commercial article warns and never gates. The compiler reports it (reading the
recorded fact, so it still owns no source convention), in both modes, aggregated by licence. It is
the same call as the ClinVar
clin_sigcross-check: refusing would make the format arbitrate a copyright question. And note the merge rule's consequence — rows written before 0.6 carry nolicense, and a re-run will not back-fill them, because merge-not-clobber cannot tell an absent value from a curator's deliberate blank. Deleteliterature.csvto re-derive.
Coverage is partial by nature, and reporting it as a fraction is part of the check. A pass that said
"0 quotes found" for an article it could not read would be describing its own reach as a defect in the
module, so quotes_found is null (not zero) when no fulltext could be read. The denominator counts
only citations that carry an authored quote: one that asks no question was not skipped for lack of an
answer. (That distinction is not hypothetical — it was a real bug, found by running the pass against
reference_examples/pathogenic_clinvar/, whose single citation is open access and quote-free, and
which the first wording therefore described as unretrievable.)
A quote is an attestation, so no tool may write one — and retrieving the fulltext changes what the
check proves. provenance_quote/provenance_regex mean a curator read this passage in this paper,
which is why both are registered in hints.REDUNDANCY_BEARING and in
hints.ATTESTATION_BEARING: the second names the sharper refusal, because filling doi from the
registry that checks it merely spends a comparison, while extracting a passage from a fulltext a tool has
just fetched states something false. The consequence for the check itself is worth being blunt about:
quotes_found is independent evidence only while the author and this pass read the article separately.
Once a machine has retrieved the text, a hit shows the quote pairs with the PMID — still worth
having, since it catches a passage filed against the wrong paper — but no longer that the claim is in the
article, because nothing establishes a human ever looked. (Reported as S11; the map had simply never
learned about the comparison this pass performs.)
Existence and retrievability are different questions, and only the second is affected by a paywall.
PubMed indexes paywalled work like any other, so exists is answered for it — PMID 12345678 is not in
PMC at all and esummary confirms it perfectly well. What a paywall blocks is the fulltext, and two
things close most of that gap:
- The abstract. Europe PMC serves it for non-open-access records, in the same
searchresponse the pass already makes — four of five probed non-OA papers carried one (the exception was a 1994 non-research document). A quote found in the abstract is as conclusive as one found in the body; a quote missing from a 200-word abstract says nothing, so the row recordsquote_sourceand the miss still counts as unchecked. That asymmetry is the whole reason the column exists. - Crossref, for the citations PubMed structurally cannot cover: preprints, books, theses, datasets
have DOIs and no PMID. A probed bioRxiv DOI returns
type: posted-content; a fabricated one 404s. It checks the authored DOI rather than the derived one — the registry's own exists by construction — and a transport failure recordsNone, neverFalse. This is also what makes the 1.0 doi-first flip low-risk: whenpmidbecomes optional, existence checking already works without it.
Google Scholar is not an option and it is worth saying so rather than leaving it as an open idea: it publishes no API, and automated querying violates its terms and is blocked in practice.
Both corrections below came from probing:
- The PMC ID converter is not used by this pass, and the reason is directional.
esummaryalready returnsdoiandpmcinarticleids, and Europe PMC'ssearchreturns them too, so calling the converter for PMID → PMCID is a third request for data already in hand. Worse, it answers a different question: for PMID 12345678 — a real, indexed PubMed record — it repliesstatus: error, "Identifier not found in PMC". Wired in as an existence check it would report every paywalled article as a broken citation. None of that says anything about PMCID → PMID, which is the direction the converter exists for and the one a curator has no other route to — see the PMC id section below. - Europe PMC is not an existence oracle. Asked about three ids where one does not exist, it returns two results and silently omits the third — no error, no marker. PubMed decides existence; Europe PMC decides retrievability.
PubMed and PubMed Central ids are one letter apart, and the outcome used to turn on a space (RM50).
StudyRow.pmid is free-form and validated through spec.extract_pmids, whose pattern is \b(\d{1,8})\b
— so PMC3110566 came back empty (no word boundary between C and a digit) while PMC 3110566 came
back ['3110566'], and 3110566 is a real PMID for an unrelated article, because PubMed ids are
densely allocated. One spelling of one mistake was refused with a message that never said "PMCID"; the
other was accepted as a confident citation of the wrong paper. Three things ship for it, all of them
diagnosis and none of them repair:
- The schema refuses a digit run whose immediate context spells
PMCin any spacing (spec.PMCID_PATTERN,spec.extract_pmcids) and the message names the id it saw rather than the one it wanted. Narrow by construction: a cell carrying both (21551363; PMC3110566) still yields the real PMID and is accepted, so only a cell whose sole numeric content is a PMC id refuses — and that cell previously resolved to another article entirely. literature._pmcid_conflictscatches what the schema cannot see: a cell like21551363 (PMC3110567)carries a real PubMed id, so nothing refuses it, while the two halves name different articles. It costs no request — the PMC id is already in theesummaryarticleidsblock — and it is the_doi_conflictsshape, includingstrictrefusing.lookup_citation(pmcid=…)/hint citation --pmcidresolves the other direction through NCBI's converter and then asks PubMed which paper that is, because a converter that hands back a number and stops is the same existence-is-not-identity failure one registry over. The resolved id comes back as an advisory (applied=False,refusal="redundancy_bearing"): fillingpmidfrom NCBI would makeLiteratureRow.existscompare NCBI with itself, which is the argument already made fordoi. Four outcomes, spelled four ways — resolved, in PMC with no PubMed id, not in PMC, and never answered — because collapsing the last two would render a failed request as a definite negative (S20). The address ispmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/; the long-publishedwww.ncbi.nlm.nih.gov/pmc/utils/idconv/v1.0/301-redirects to it.
On evaluating provenance_regex here. The charter requires a linear-time/ReDoS-safe engine, written
when the match was specified as consumer-side. Here the pattern comes from the module being enriched and
the document from a public archive, on the author's own machine, so the risk is a curator writing a slow
pattern by accident. That is worth a bound rather than a compiled dependency — but the bound must be a
child process, not a thread. The thread version looks correct and is not: re cannot be interrupted,
threads cannot be killed, and the interpreter joins pool threads at exit, so a runaway pattern returns on
schedule and then hangs the process on the way out. A timeout is recorded as not checked, never as
not-found.
Existence is not identity, so a citation lookup says which paper it found (S12, 0.5.4). PMIDs are
densely allocated, so a recalled or invented eight-digit number is usually a real record for a
different article — which means pmid_exists=True could never catch a fabricated citation, and the
surrounding docs had been treating existence as the guard until a consumer's authoring skill had to
retract a rule this surface could not enforce. literature.bibliographic(summary) pulls
title/journal/first_author/year out of the same esummary response that answers existence —
no extra request — and lookup.CitationHint carries them, with an info finding naming the paper and
hint citation --json (which hint variant had and this did not). It is public rather than private
_identifiers-style for the RM41 reason: two tiers read it, and the alternative is a consumer
re-parsing a payload we already hold. Every value is None when the field is absent rather than an
empty string, and year is taken only from a leading four digits of the free-form pubdate
(2017 Nov-Dec), so nothing is invented. No title column on LiteratureRow: that table records
what was checked, not bibliography. Generalize it — when a check answers yes/no about an identifier,
ask whether "yes" could be true of the wrong thing.
Identifier currency (identifiers.py, online)¶
The generalization of COMPILER.md's "is the source stale?" blind spot from datasets to identifiers. A dbSNP merge, an EFO retirement, an HGNC rename and a PGS accession that was never issued all leave a module perfectly well-formed and quietly out of date.
- rsIDs run inside
enrich()(--verify-rsids), because the verdict lands onresolution.csv'srsid_current/rsid_statuscolumns. Three states fromesummary db=snp: live (snp_id== requested,merged_sort='0'), merged (snp_id!= requested,merged_sort='1'), absent ({'uid': …, 'error': 'cannot get document summary'}). - Traits and gene symbols are module-level with no sidecar column to fill, so they get their own
command,
just-dna-enricher check-identifiers. HGNC uses the exactfetch/symbolandfetch/prev_symbolendpoints, neversearch/—search/BRCA1is fuzzy and returns 19 hits includingABRAXAS1. - Gene ↔ locus agreement (0.5.4) rides on the same command and is the first check here about a relationship between two identifiers rather than the currency of one. See below.
- PGS accessions (0.7, RM163) are the fourth registry and ride on the same command, under
--pgs/--no-pgs. They put two questions — does the accession still name a score, and do the two authored cells beside it still match the record — and the command writes the Catalog's per-score terms intolicensing.csvwhile it is there. See below.
NCBI is the oracle, not Ensembl. Ensembl REST resolves some merges (rs77121243 → rs334) and
returns HTTP 400 on others (rs3216883, which dbSNP reports as merged into rs3051860), so Ensembl
alone would misclassify a merged rsID as unresolvable.
absent conflates two opposite meanings and no live endpoint separates them. A withdrawn rsID
(rs11273140, retracted after a clustering error) returns a response byte-identical to a never
assigned one (rs2000000000). For an author these mean opposite things — fix the typo, versus the
variant itself was retracted and the annotation resting on it may be worthless — so the message names
both readings and asserts neither. Guessing "typo" would send an author to fix the wrong thing.
withdrawn is nevertheless a real vocabulary member, and refuses in both modes. Nothing automated
emits it today, which is a limitation of the API rather than of the model, and the member is kept for
two reasons: a curator who has established a retraction can record it in resolution.csv and have the
tooling honour it, and a future source that can tell the two apart starts producing it without a
vocabulary change — which Principle 3 would otherwise make a one-way door. Its severity is deliberately
not absent's: a merged or absent rsID leaves the annotation intact (dated, or unserved), while a
retracted variant may leave it describing nothing, so withdrawn is the one resolution finding fatal
in best_effort too.
Report, never repair, and here that is load-bearing. weights.parquet carries both variant_key
and rsid; for an rsid-authored row they are the same label. Writing the merged-into id back would be
an identity migration performed by a network lookup — reverse would emit the new rsID, the next
compile would key on it, and variant_key would change with no authored edit anywhere. Severity is the
usual ladder (warn / fail in strict), and strict failing is the nudge toward the drift-proof key:
author the coordinate, and the VRS allele id cannot drift at all.
Two true halves, one false row — gene ↔ locus (S24, 0.5.4)¶
variants.csv carries a gene column and nothing compared it to anything. The gene check above asks
HGNC whether a symbol is approved, which is a different question — FTO is approved whatever variant
sits beside it — so a row pairing a real gene with a variant on another chromosome passed everything,
because both halves were individually true and only the relationship was false. Four of one reporter's
seven rows were exactly that: real symbols beside invented rs numbers, which resolve anyway because
dbSNP is dense enough that almost any seven-digit number hits something. Machine-written sources are a
real authoring input now, and this is the shape they fail in — which is also why it belongs beside the
currency checks rather than among them: staleness is the world moving, and this is a claim that was
never true.
check_identifiers reports GeneLocusConflict per row and repairs nothing — which of the two halves is
wrong is not something this tier can know. Four design points:
- Chromosome granularity only, and the stronger version is refused in the code using the reporter's
own argument.
rs1421085sits in an FTO intron and acts on IRX3/IRX5 megabases away, so a row may legitimately name any of the three; an interval check would fire on correct rows until someone switched it off. A test pins that the FTO row stays silent with the variant nowhere near the gene body. Chromosome disagreement has almost no legitimate cause, costs one comparison, and catches the whole fabrication class. - The join is against HGNC's cytoband (
16q12.2→16,mitochondria→MT), and anything unparsed yieldsNonerather than a guess — a guess here becomes a false accusation about a row. Three-valued as everywhere else: an unknown symbol, an unparsed band and a row with no known chromosome all withhold. - Nothing is fetched for the coordinate. For an rsID-only row the chromosome comes from an injected
resolution.csvbeside the spec — the table the compiler already consumes — because a currency check should not depend on a resolver. - A pseudoautosomal gene is exempt.
XGstraddles the PAR1 boundary, so X/Y there is a spelling, not a contradiction (RM32).
IdentifierReport.gene_loci_not_checked carries the reason when the comparison could not run, for the
reason EnrichmentResult.clin_sig_not_checked exists — an empty conflict list otherwise says both
"compared everything, nothing disagreed" and "never compared". The CLI prints it.
The PGS Catalog — the fourth registry, and a licence that is per score (0.7, RM163)¶
pgs.csv is a manifest of PGS Catalog accessions and PgsRow is keyed on pgs_id, so until 0.7
the one authored identifier the format keys a whole table on was the one nothing ever asked its
registry about. check-identifiers now asks, under --pgs/--no-pgs, and it puts two questions
rather than one.
pgs_accession_currency — does the accession still name a score? Three outcomes, and the shape of
the source decides how they are reached.
- The verdict is read off the body, never off the status.
GET /rest/score/PGS000001returns the record;GET /rest/score/PGS999999, a never-assigned accession, returns HTTP 200 with{}; andGET /rest/score/PGSXXXX, which is not an accession at all, returns HTTP 200 with{}too. So the status code carries no existence information, and a 404 from this service means the request went somewhere unexpected rather than that the score is absent. That is the opposite ofOntologyClient's rule, where OLS4's and HGNC's 404 is the answer, and the per-client fact is whytest_client_exception_contract.pykeeps a 404 table rather than a tier-wide rule. - The absence is weighted by the measured base rate, and it is the other way round from dbSNP's.
The accession range is sparsely assigned — roughly a third of it, measured over a random sample —
so an unrecognised accession is overwhelmingly one that was never issued.
@rsid-absent-two-readingsgives typo and withdrawal equal weight because dbSNP's id space is dense and merges are frequent; here the message names the typo reading first and withdrawal as the rarer one. The Catalog publishes no supersession field, which is stated as a limit of the source rather than resolved by guessing. - A malformed accession is settled before any request.
PgsRow's ownPGS<digits>grammar answers it, and asking would have established nothing — the Catalog's reply is byte-identical to a never-assigned one. In practice the schema refuses such a row at load, so the state is reached only by a caller holding an id from somewhere other than a parsed table; the arm exists so the message can say spelling rather than offering a withdrawal that almost certainly did not happen. - A recognised accession says what it found — name, release date, variant count, traits
(
@existence-not-identity). A PGS accession is exactly the shape of identifier where a wrong-but-real id resolves cleanly to somebody else's score.
pgs_metadata_agreement — do the two authored cells beside it still match? training_ancestry
against the score's ancestry_distribution, training_cohort against its samples_training. Reports,
never repairs. Its own vocabulary member rather than a second finding under the line above, because the
subject is a cell and not an id: a module with three accessions can put six cells, settle two and
withhold four, and none of those numbers fits a denominator counted in accessions.
- Only the
devandevalstages are compared. The Catalog publishes three —gwasis the ancestry of the discovery study the effect sizes came from,devthe samples the score was built on,evalthe ones it was evaluated in.training_ancestryexists so a consumer can refuse an out-of-ancestry application, which is a question about where the score was built and tested; foldinggwasin would widen the published set until nothing could disagree. - The two vocabularies do not meet, and the gap is withheld rather than reported.
VALID_TRAINING_ANCESTRYis 1000G superpopulations plusmulti; the Catalog's categories addNR(not reported),ASN,GMEandOTH, which no member covers. A score whose distribution names one of those has an ancestry this tier cannot spell, so an authored code the published set does not carry might be exactly that one — and a difference nobody can stand behind is not a finding.PGS_ANCESTRY_CATEGORIESis total over the categories the service serves, with the unmappable ones written out rather than left to a.getdefault, and a category the Catalog invents later is logged. MAEandMAOare bags, and the two directions are not symmetric. Multi-ancestry excluding European and multi-ancestry including European each stand for two or more superpopulations the Catalog did not break down, so both map tomultifor a positive answer and both block a negative: an authored code the rest of the distribution does not name may be inside the bag. Without that,MAOreported a false drift against an authoredEUR— a population it includes by definition._UNENUMERABLE_CATEGORIESis that half, kept apart from theNoneentries because the two are different facts: aNonecategory has no member at all, these have one and still cannot be exhausted.- The comparison is one-directional. An authored cell narrower than the published set is not flagged: a score the Catalog evaluated in eight ancestries at a percent each is not thereby validated in all eight, and claiming fewer is the conservative direction.
training_cohortis free text, so the disagreement bar is high. The cell agrees when any non-structural word of it appears anywhere in the record's own account of its training samples — cohort names, and the per-sample ancestry prose beside them, becauseFINandAshkenaziare the field's own examples and are ancestries rather than cohorts. Matching is substring rather than whole word, soFINis vouched for byFinnish. Words that describe what a cohort is (cohort,study,project) are dropped: they match almost every record, and letting one vouch would make the check unable to fail while looking like one that passes.match_rate_floorandresearch_tierare deliberately not checked, and the reason is written into the skip's owndetailso nobody adds them on the symmetry.match_rate_flooris described in its ownFieldas an author-set floor andresearch_tieris a two-member curator judgement; the Catalog publishes neither, so a check over them could not fail (@tautology-zero).- The overlay has nothing to say here, and that is structural.
overrides.csvcorrects derived rows andOVERRIDABLE_TABLESis built from the derived registry;pgs.csvis authored, so a disagreeing cell is edited in place rather than answered. There is no suppression path to wire and none is missing. - One accession on two rows is one claim only while the two rows agree.
PgsRowis keyed(pgs_id, trait_efo_id), so a pleiotropic score is two rows; the cell key carries the authored value, which collapses identical claims and keeps differing ones apart. Keying on the pair alone dropped whichever row came second, so whether a stale value was caught depended on row order. - A drift never fails
--strict. An accession the Catalog does not hold is a broken identifier and does, exactly like a retired HGNC symbol; a cell two authorities disagree about is not, and the finding's own message tells the author to leave it if the curation is deliberate. Failing a build on that would be a gate with no way to clear it — the shape@clinsig-never-escalateskeeps out. The green line is withheld all the same, so a clean exit never reads as agreement. - A Catalog outage reaches only these two records. The leg records
unreachableon the report rather than raising:check_identifiersputs questions to four registries and the command writes one record per check, so an exception here would stampunreachableagainsttrait_currencyandgene_symbol_currencyfor registries that answered. The ontology legs run first for the mirror reason — an OLS4 failure aborts before this leg spends requests and writes a licence table the run is about to discard. Called withvariants=rather than a spec directory the record isunsupported, because no row model a caller can hold carries apgs_idat all.
Currency is read, not built. /rest/info publishes latest_release — date, score count,
publication and trait counts — beside the REST API's own version, so pgs_catalog joins
currency.PROBE_SOURCES as the second shipped probe with no builder and no snapshot behind it.
pgs.dataset_label spells the release for the probe and for the SourceRow.dataset the check
stamps, which is clinvar_dataset_label's rule: two spellings of one release make the comparison
quietly never match, and never-matches renders as unreadable rather than as a bug. A release record
that cannot be read leaves pgs_release at None and the accession verdicts standing — it is a fact
about the Catalog, not about any accession. The floor row's dataset is paired with
withdraw_stale_dataset, exactly as the two drafting providers pair theirs: merge_sources_csv never
clobbers, so without it a row written under an older release would hold that label forever and
verify_datasets would report it behind on every run with the pass that owns it unable to refresh it.
It only ever blanks — one column cannot name two releases.
The terms are per score, and this is a correctness requirement rather than thoroughness. license
is a field on the score record. Measured over the first 250 of the corpus: most carry the Catalog's
generic sentence deferring to "any licensing restrictions set by the authors", a handful are "freely
available to the academic community for research use" — academic-use-only, the class this workspace
already classifies as barring redistribution outright — and a couple are CC0. So PGS_TERMS is
written as a floor in GWAS_CATALOG_TERMS' shape (EBI's terms-of-use URL, every gating axis
None), and the pass writes one further SourceRow per score, source=pgs_catalog:PGS000013, with
that score's own licence string in notice and hashed into license_sha256; license itself carries
a short name (CC0-1.0, Academic research use only, or null for the generic sentence) because
the compiler compares that column against module_spec.yaml's declared licence by equality, and a
250-character sentence there would put a paragraph into manifest.sources.licenses and into a warning
the author cannot act on. All rows sit at the annotation layer, which is where
taints_commercial_use reads them, and most-restrictive-wins does the rest. Measured end to end: a
module naming one academic-use-only score fails the compile licence gate by name until its author
records a non-commercial declaration, where a single generic row would only have warned. A single constant covering the source would be a false claim for a measured minority in the
permissive direction, which is the direction that matters — the same shape as @per-article-terms and
as the SpliceAI note inside GNOMAD_TERMS. A licence string the classifier has not read is unknown on
all three axes and logged; unknown is not permission.
No --use flag and no check_declared_use call, for pubmind_draft's reason: every gating axis
on the floor is None, so the gate would refuse every run unconditionally, and reading a public
endpoint to check an identifier is a read rather than an acquisition anyone has gated
(@acquisition-gate-is-not-a-read-gate).
ACMG secondary findings (acmg.py + acmg_build.py) — just-dna-enricher check-acmg¶
VariantRow.acmg_sf has been materialized into weights.parquet since 0.4 and checked against
nothing. The compiler cannot hold a gene list (that is the un-injected reference RM21 taught), and no
pass here had one, so the column was assertable and unfalsifiable. This closes it.
Read this part first: the scraped list is a release behind, and the check now says so. ACMG
published SF v3.3 in June 2025 — 84 genes over 100 gene-condition rows, adding ABCD1, CYP27A1
and PLN — and NCBI still serves its adaptation of v3.2 (81/94). The five guards below all pass on
that page, because it is neither truncated nor re-laid-out; it is simply old. So the check reported
acmg_sf=true but ABCD1 is not on ACMG SF v3.2 about a row that is right, which is precisely the
short list failure the guards exist to prevent, arriving where no guard could see it.
Two things changed, and the --sf-list half is the one to use:
# once, from ACMG's supplementary workbook (assets/acmg_sf_v3.3.xlsx, or your own download)
just-dna-enricher acmg build assets/acmg_sf_v3.3.xlsx # → data/repro/acmg_sf
just-dna-enricher check-acmg spec/ --sf-list data/repro/acmg_sf --offline
- The list can be injected.
acmg buildturns ACMG's workbook intoacmg_sf.csv+release.json(declaredsf_version,source_sha256, DOI, counts), andload_acmg_snapshotreads it with the standard library — socheck-acmgis the first check here that works--offline. The workbook is the better artifact in every way that matters: version-pinned behind a DOI rather than hand-maintained, content-hashable, and carrying four columns the page does not (Inheritance,Phenotype Category, the release that first listed the gene, and ACMG's scope-of-reporting text). Only MedGen concept ids go the other way — they are on NCBI's page and not in ACMG's sheet, and no verdict reads them. - The scrape path carries a staleness tripwire.
KNOWN_LATEST_SF_VERSIONis one version string, not a gene list. When the list actually read is older, every disagreement — both directions, since ACMG can remove entries as well as add them — is demoted tounverifiable: reported as a warning, never astrictrefusal. A mismatch against a superseded list is a question, and answering it anyway is worse than saying nothing. The stale-constant risk is asymmetric on purpose: when v3.4 ships the constant under-warns, i.e. degrades to the previous release's behaviour, whereas a hand-kept gene list would make confident wrong claims about specific genes.
The scrape stays, because it is the only zero-setup path. Probed 2026-08-03: ClinGen's FTP publishes
gene-curation, region-curation, dosage and recurrent-CNV lists and no secondary-findings list;
ClinVar's FTP tree carries no ACMG flag (gene_condition_source_id, 13,478 rows, zero mentions). NCBI's
adaptation of ACMG's Table 1 at /clinvar/docs/acmg/ is still the only fetchable form of the list, so
it remains the fallback with the guards that branch was conditional on — now with its version checked.
The guards are load-bearing, and the naive parse really is wrong. Splitting the table on <tr>
returns 78 of the 81 genes, silently. The page is hand-maintained and shows it: two rows open with
a bare <td> after the previous </tr>, four leave a <td> unclosed with a stray trailing </td>,
and the gene cell links through three URL shapes (/gtr/genes/324, /gtr/genes/4089/,
/gene/3949). The three genes the naive split drops are TP53, COL3A1 and TPM1 — which is the
failure mode stated exactly: a short list makes correctly authored acmg_sf=true rows look wrong, and
it would have begun with the single most recognizable secondary-findings gene there is. So the parse
works in cells, not rows, and refuses rather than returning a short list:
| Guard | Catches |
|---|---|
the page declares ACMG SF vN.N |
a re-write, and it supplies the dataset label |
one table carries all four EXPECTED_HEADERS |
a re-layout, or a nav/footer table becoming "the list" |
<td> count divides exactly by four |
a column added or dropped |
| every four-cell group yields exactly one gene link | a gene cell that lost its link — a silent drop |
at least MIN_GENES distinct genes survive |
a truncated response, a JS shell, an error page |
MIN_GENES is a floor, not the real count. Hard-coding 81 would be the hand-transcribed gene list this
module exists to avoid, and it would go stale the day ACMG publishes v3.3 — which it since has, and
the version tripwire above is the answer that a count would not have been. The workbook parse reuses the
same floor and adds one guard of its own shape: ACMG's trailing disclaimer sits in the Gene column,
~1,200 characters of prose that a naive read counts as an 85th gene. It is skipped only when every other
cell in its row is empty; a symbol that cannot be read on a populated row refuses, because that is the
<tr> failure again.
The list is richer than a set of symbols, and cells hold more than one of things. Each row is a
gene–condition pair (94 pairs over 81 genes; TRDN is listed for two conditions), and a cell can carry
several MIMs and several MedGen concepts at once — SDHB names MIM 115310 and 171300 against MedGen
C1861848, C0031511, linking to a MedGen search rather than a concept. disease_mims and
medgen_ids are therefore tuples; taking the first of each would be a silent truncation of the same
family as the <tr> split.
The verdicts are the house tri-state, and a blank cell is never a defect. agree (either way
round) and blank are silent; not_listed (claimed true, gene absent) and denied (claimed false,
gene present) are findings that warn in best_effort and refuse in strict; unverifiable is either
of those two against a superseded list, which warns in both modes; unstated (blank, gene
listed) is a note, because blank means "not stated" and turning that into a defect is the
None-means-False collapse this codebase refuses everywhere else; unchecked covers a row naming no
gene and --offline without a --sf-list, which reports that nothing was asked rather than that
nothing was found.
Findings group by gene, because that is what they are about. Found by running the real thing: the
HFE reference example is 13 variants in one gene, and a per-row report printed the same 220-character
sentence 13 times. Same aggregation rule CPIC already taught (~600 identical lines for CYP2C19). The
per-row verdicts stay on the report for a caller that wants them; AcmgReport.by_gene is what a report
prints.
Gene-level, and the column says so — but ACMG's own list is not purely gene-level. acmg_sf is
documented as "True when the gene is on the ACMG secondary-findings list", so that is what is
compared. ACMG itself is finer-grained in at least one place: the HFE entry reads "Hereditary
hemochromatosis (c.845G>A; p.C282Y homozygotes only)". The parse keeps that text and the denied
message points at it, telling an author to leave the cell blank rather than false when a row is
about a variant in a listed gene that is not itself a reportable finding. Reading the column as
per-variant reportability would make the format decide disclosure policy, which it does not do.
This pass records no SourceRow, the deliberate exception to "a pass that consults a source writes
one". That rule is about a module carrying a source's data. Nothing lands in the module here:
acmg_sf was authored by a human before this ran, exactly as a gene symbol was, and this asks a
registry whether the authored value is still right. It is check-identifiers' shape (HGNC and OLS4 go
unrecorded too), not dosage's. The corollary is in the other direction: acmg_sf joins
hints.REDUNDANCY_BEARING, so no lookup or hint may fill it — a cell filled from the list this checks
against would make the check vacuous.
Repeat-allele bands against STRchive (strchive.py + strchive_build.py) — check-repeat-bands¶
repeat_alleles.csv was the one binning kind nothing in this tier had ever asked a question about,
and the corpus behind it is two hand-authored modules. STRchive (dashnowlab/STRchive, MIT) publishes
82 tandem-repeat disease loci with coordinates, motifs, the motif structure and three bands each —
benign_*, intermediate_*, pathogenic_*.
The item splits the source by column: the identity half is drafted, and the band half is checked and never written. That split is the whole finding, and it was decided by measuring both corpus modules rather than by caution.
just-dna-enricher strchive build --release v2.26.0 # → data/repro/strchive
just-dna-enricher check-repeat-bands spec/ --catalogue data/repro/strchive
What the measurement said¶
reference_examples/htt_repeat_expansion reproduces the catalogue exactly where the catalogue speaks:
STRchive gives benign 6–26 and intermediate 27–35, and the shipped table gives 6,26 and 27,35,
independently authored. reference_examples/fmr1_cgg_repeat is where the source is coarser —
STRchive runs one intermediate band from 45 to 200 and the module states 45–54 and 55–200, and
55 is the premutation threshold, the line between a carrier who is not at risk and the FXTAS/FXPOI
range. A drafting provider would have written the three bands as the answer and erased it.
Building the check turned up a third thing the item had not measured: HTT is finer than the catalogue too. The module divides STRchive's single pathogenic band at 40, which is the reduced-penetrance / full-penetrance line. So both corpus modules refine the catalogue, in two different places for two different reasons, which is a stronger argument for the split than the one that motivated it.
pathogenic_max is reported and never written¶
STRchive's HTT pathogenic_max is 250 and the module leaves its top band open. A catalogue's
pathogenic_max is the longest allele the literature records — an observation, not a clinical bound.
Written into measure_max, a 300-repeat allele would match no bin at all, silently and under
--strict too, which is the exact silence RM55 shipped a loud warning about. It gets its own finding
kind, ceiling_only_in_source, so that "the source states a ceiling the module does not" can never be
confused with "the two ceilings disagree", and a test asserts the number reaches no cell.
How the comparison is put¶
The authored rows are partitioned by binning._bin_groups — _KEY_FIELDS + (trait_efo_id,), with
unresolved sentinels dropped — because that is the partition the overlap rule is enforced over, and
grouping any other way would answer this question against a partition nothing else uses.
A group joins a catalogue locus on gene plus the motif, matched against both orientations the
source publishes: repeat_unit is one string with no column beside it to say which strand it is
written on, so a minus-strand locus is legitimately authored either way.
The comparison reads lower bounds, which is the one representation both tilings share: adjacent
bins are [6,26] [27,35] under quantised tiling and [6,26] [26,35] under continuous
(@dense-bin-boundary), so an upper bound means two things while the next bin's lower bound is the
division either way. A source division is a closed window [previous.max, next.min] rather than a
point, so the same division stated in either spelling matches; resolve_tiling decides only which
endpoint the message names, and whether a cut at a bin's own top edge divides it.
Eight finding kinds, one sentence each, asserted equal to the arms compare_bands can emit: a
boundary only the module has or only the source has, a floor that differs or is open on one side, and
the same three for the ceiling.
What it withholds, and what it never does¶
Three absences get three different sentences, because a reader chasing a skipped locus looks in three
different places: the catalogue carries no such gene; it carries the gene under no matching motif; or
it states no bands for it at all (twelve of the 82 loci state none). A (gene, motif) key several
loci claim — ARX has two and HOXA13 three in the published file — is reported as contested and
compared against nothing, because picking one would make the verdict depend on record order.
--strict never escalates a band difference, in either direction and on either corpus module.
Two expert bodies drawing a threshold in different places is a difference between authorities, and a
compile that refused would have this format pick the winner — the rule the ClinVar clin_sig and PGx
allele-function checks already follow. What strict still refuses is structural: a
repeat_alleles.csv that will not load raises in both modes.
This check records no SourceRow, on acmg's rule rather than as an oversight: sources.csv
exists so a module can account for what it carries, and nothing from STRchive lands in the module on
this path — the bands were authored by a human before it ran. The drafting half does carry the
source's data and does write one.
It is also deliberately absent from provenance.DRAFT_PROJECTIONS. That map is for a source that
drafts a column one of its own checks later compares; the split means the columns this check reads
were never copies of anything, so a projection over measure_min/measure_max would let the
tautology skip fire on cells no drafter ever wrote.
The snapshot¶
strchive build streams STRchive-loci.json and writes it beside a release.json carrying
source_url, source_sha256, locus_count, dataset, license, built_at and builder_version.
The catalogue is copied rather than re-serialized, so source_sha256 describes bytes a reader can
verify with sha256sum. There is no parquet and no [dev] extra: the shape the readers want is the
shape the source publishes.
Pin a release. --release v2.26.0 builds from the tag and records dataset=strchive_v2.26.0,
which is what the verification record then names. Building from the default branch is allowed and
records no label at all rather than inventing one out of a date — an unlabelled snapshot is honestly
unlabelled. There is no --offline (the off-switch is passing --catalogue instead of downloading)
and no --use (MIT grants the fetch, so the gate would answer the same on every run).
Drafting the identity half (strchive_draft.py) — draft-repeats¶
A drafted row is thinner than the item that ordered it expected, and that is this module's own
finding. The proposal listed chrom/start_hg38/stop_hg38, locus_structure, ref_copies and
the disease identifiers as the identity half. RepeatAlleleRow has a column for exactly one of them:
trait_efo_id. Repeat coordinates are RM65 and the motif structure is RM66, both deferred, so what a
row can actually carry is the gene, the motif as the catalogue spells it, the trait CURIE and the
fixed measure_kind, with conclusion left as <<REPLACE>>.
The rest is counted and reported, not dropped: how many loci state a fractional ref_copies (a
reference allele that is not a whole number of motif copies — RM55's case arriving in a source rather
than in a caller's VCF), and how many publish a locus_structure. Neither is rounded into a cell and
neither is silently discarded.
STRchive's own evidence grade is named, since RM276. Every locus carries a ClinGen-style grade
(Definitive … Refuted), and no authored column holds it, so a locus STRchive grades Refuted
(DMD) or Disputed (DIP2B, NIPA1, POLG) used to draft exactly like HTT. A drafted row
carrying one of those two grades is now named in a note, in both modes and never as a gate; a
Provisional locus (STRchive's not yet curated) is named apart and never called weak. The grade is
read off the catalogue file by a private reader; with no file to read, the note says the grades were
unread rather than implying none is doubted. Whether the grade should travel into the module is RM294.
What the thin row is still worth is the part an author cannot get from anywhere else: the list of
loci, the motif in the orientation the catalogue publishes, and the MONDO id. On HTT the provider
derives MONDO_0007739 from mondo and the shipped module's human author wrote the same value —
independent agreement, which is the strongest evidence a drafting provider gets. On FMR1 the
catalogue names three MONDO ids (fragile X syndrome, FXTAS, FXPOI) and the provider writes none,
which is what the shipped module does too.
Never drafted: the band columns. measure_min, measure_max, measure_tiling and the per-band
direction/clin_sig/phenotype. The withheld set is derived from the drafted one, so a column
added to the kind later is withheld by default rather than silently written empty, and a test asserts
the two partition the model's authored fields.
A (gene, repeat_unit) key several loci claim is drafted for none of them — ARX and HOXA13 are
the published cases — and contestation is decided over the whole admitted set before the gene filter
runs, so a filter cannot leave one claimant looking uncontested. The skip guard for a locus with no
usable identity is read out of authoring_requirements, not restated beside the model.
Rows match on (gene, repeat_unit) and go at the end. trait_efo_id is deliberately outside the
match key even though the bin-group key includes it: it is a cell the author may clear, and a re-draft
after they did would otherwise append the locus twice. A re-draft after the human has written the
bands and split one drafted row into four still reports already_present.
STRCHIVE_TERMS is MIT — share_alike=False, commercial_use=True, redistribution=True, all facts
rather than the None every other axis in this round had to say — so a module drafted from it stays
sellable and check_declared_use answers None on every declaration. The row is written at the
annotation layer with the snapshot's release in dataset, because this half does carry the source's
data into the module.
Pharmacogenomics and data-source licensing (pgx.py, licensing.py, pharmvar.py, cpic.py)¶
Pass 5 cross-checks a module's star-allele tables against the nomenclature authorities and records what was consulted and on what terms into the module's licence table. It is the first pass whose primary output is provenance rather than facts.
That file is licensing.csv since 0.6, and sources.csv is the deprecated spelling (RM51,
removed at 1.0). No pass names either by hand: licensing.sources_path — and sidecar_path for the
other machine-written tables — resolves it through just_dna_format.layout, which also accepts a
derived/ subdirectory (RM49). Two rules every pass inherits from that, and both matter:
write to the file you read (writing the current spelling onto a module carrying the older one, or
the root onto a module that is split, leaves two copies), and two copies is a refusal naming both
paths, raised as the calling pass's own error. Never a merge and never newest-wins — these tables
are hand-overridable, so two copies are two claims.
The bottom line first: a PGx module is non-commercial only¶
Every PGx upstream forbids sale, and PGx tables are the layer that taints, so one drafted row settles
it for the whole module. ClinPGx, CPIC and PharmVar each carry CC BY-SA 4.0 plus a contractual bar on
selling the data, which licensing.{CLINPGX,CPIC,PHARMVAR}_TERMS each record as
commercial_use=False. The PGx tables —
haplotypes.csv, allele_function.csv, diplotypes.csv, pharm_variants.csv — are the module's own
authored annotation, so their SourceRow sits at the annotation layer, and that is the one layer
sources.taints_commercial_use treats as tainting. The verdict is most-restrictive-wins, module-wide:
mixing in a permissive source cannot launder a restricted one, and the compile refuses in both modes
unless the licence table records declared_use=non_commercial for every tainting source. unstated is not a
loophole — it is the absence of a declaration, which is exactly what the gate is looking for.
reference_examples/cyp2c19_star_alleles/licensing.csv is the shape: one CPIC row,
commercial_use=false, declared_use=non_commercial.
reference_examples/pgx_slco1b1_simvastatin/licensing.csv is the same with a licence hash pinned from the
bytes it was read out of (license_sha256, dataset=clinpgx_2025-07-05).
And declaring is asserting, not proving — the gate's own closing sentence says so. Recording
non_commercial states how the module will be used; nothing in the compiler can check that, and it does
not pretend to.
Two things the flat "non-commercial" summary does not say, both of which matter:
- It is not "unrestricted if you give it away." Sale and distribution are different rights.
redistributionis a third recorded axis and it is deliberately not gated (RM27) — all three PGx sources recordredistribution=true, since CC BY-SA expressly permits sharing under share-alike plus attribution, so nothing in this workspace trips it today. But the axis exists precisely because academic-use-only sources (OMIM, dbNSFP) permit neither, and "non-commercial" is the reading that would hide the difference. - PharmVar is stricter than any column can express. Its recorded
noticereads research use only … not intended for direct diagnostic use or medical decision-making, and its API key is personal and non-transferable under its terms §2.commercial_use=Falseis the nearest a column gets; the rest stays prose innotice, on purpose rather than by omission — a fourth licensing axis means a newSourceRowcolumn, and that is a design round rather than a drive-by addition, so it is a 1.0 change and not a minor one. The restriction is therefore recorded and legible; it is not machine-enforced, and nobody should file the column as cheap.
The licensing picture, probed 2026-08-02¶
| Source | Endpoint | Auth | Licence | Sellable |
|---|---|---|---|---|
| ClinPGx (ex-PharmGKB) | api.clinpgx.org |
none | CC BY-SA 4.0 + no-sale clause | ❌ |
| CPIC | api.cpicpgx.org/v1 (PostgREST) |
none | same policy | ❌ |
| PharmVar | www.pharmvar.org/api-service |
Api-Key header, 2 rps |
CC BY-SA 4.0 + research-use-only | ❌ |
| Ensembl / dbSNP | already in the chain | none | unrestricted | ✅ |
| PubMind (read 2026-08-28, RM134) | openbioinformatics.org/annovar/download/hg38_pubmind_db.txt.gz |
none | none stated for the data | ❓ |
PubMind is the only ❓ and the row is dated separately, because it was read almost four weeks after the
others and it is a different kind of answer. The 🔒 sources publish terms that forbid sale; PubMind
publishes no data terms at all. LICENSE.md covers the software (academic, non-commercial, CHOP
tech-transfer for anything else), the paper is CC BY-NC-ND 4.0, and the ANNOVAR-redistributed table —
the only per-variant channel — states nothing. So every axis is None: unknown is not permissive, and
None is not False. See The caches for what that does and does not gate.
Three things follow, and each shaped the code:
api.pharmgkb.org is gone. Retired 2026-07-20; the successor is api.clinpgx.org with paths and
formats unchanged. ClinPGx is the umbrella that merged PharmGKB, CPIC and PharmCAT.
CPIC is not an escape hatch. cpicpgx.org/license/ 302-redirects to the ClinPGx data usage
policy, so CPIC carries ClinPGx's terms. Preferring PharmVar for the star-allele layer is a
data-authority choice — it is the naming authority for CYP star alleles — not a licensing one.
No PGx source is sellable — the consequence for a module is drawn out above. Each layers a contractual bar on sale on top of the CC grant, so a bare "CC BY-SA 4.0" line is not permission to sell; read the surrounding terms. Since the coordinate layer is already covered by Ensembl/dbSNP, ClinPGx and CPIC are deliberately never wired as resolution links — that keeps coordinates unrestricted and leaves nothing to declare there.
The complete roster — TERMS_BY_SOURCE, every source this tier has terms for¶
The table above is the PGx picture and is scoped to it. This one is the registry: every key in
licensing.TERMS_BY_SOURCE, which is what a pass reaches for when it writes a SourceRow. A source
absent from it has no recorded terms, which is a different state from having permissive ones —
@no-named-licence, and the reason every column below can read —.
| Source | Licence | Sellable | Share-alike | Redistributable | Note the column cannot hold |
|---|---|---|---|---|---|
clinvar |
public domain | ✅ | — | ✅ | |
ensembl |
Apache-2.0 | ✅ | — | ✅ | |
gnomad |
CC0-1.0 | ✅ | — | ✅ | |
clingen |
CC0-1.0 | ✅ | — | ✅ | |
gencc |
CC0-1.0 | ✅ | — | ✅ | |
civic |
CC0-1.0 | ✅ | — | ✅ | the only annotation snapshot that may be published |
strchive |
MIT | ✅ | — | ✅ | |
mitomap |
CC BY 3.0 | ✅ | — | ✅ | a floor: the host grants it "unless otherwise noted", so a per-record note outranks it |
alphagenome_avi |
AlphaGenome Services Additional Terms | ✅ | — | ✅ | Permissive Use (RM195); the sign-in bars classes of holder, which no column says |
clinpgx |
CC BY-SA 4.0 | ❌ | ✅ | ✅ | + a contractual bar on sale, on top of the CC grant |
cpic |
CC BY-SA 4.0 | ❌ | ✅ | ✅ | cpicpgx.org/license/ 302s to ClinPGx's policy |
pharmvar |
CC BY-SA 4.0 | ❌ | ✅ | ✅ | research-use-only, and the key is personal and non-transferable |
alphagenome_atlas |
AlphaGenome Output Terms of Use | ❌ | — | ✅ | a second name for one service, because one (source, layer) key cannot carry two licence classes |
pubmind |
— | ❓ | ❓ | ❓ | publishes no data terms at all; the software licence and the paper's are not the table's |
mane |
— | ❓ | ❓ | ❓ | NCBI states a policy rather than a licence — nothing to refuse and nothing to permit |
pgs_catalog |
— | ❓ | ❓ | ❓ | a floor: terms are per score, so PGS_LICENSE_CLASSES is consulted per record |
gwas_catalog |
— | ❓ | ❓ | ✅ | redistribution is stated; the other two are not |
clingen_allele_registry |
— | ❓ | ❓ | ❓ |
❓ is None and never False. A source whose terms could not be established has not been shown
to permit anything, and "we could not read the terms" is not a finding that they forbid anything
either — so it is skipped in every --use column with a message saying which of the two it is. That
is the house algebra landing on a licence: unknown never becomes permission, and it never becomes a
refusal either.
Four of the eighteen are a floor rather than a total, and that is a property of the source, not of
the constant: mitomap grants site-wide "unless otherwise noted", pgs_catalog is per score,
gwas_catalog states one axis of three, and alphagenome_avi's permissive Output terms sit behind
an eligibility clause about the holder that no column models
(@a-hosts-terms-are-not-its-contents-terms).
declared_use — a third axis, not a mode¶
--use is unstated (default) | non-commercial | commercial, threaded to enrich_pgx(declared_use=).
It is orthogonal to mode on purpose (Principle 5): mode says how hard to fail on a finding,
declared_use says who is using the data and why. A three-state string rather than a bool pair,
because a bool cannot express the default and defaulting either way would have the tool assert a
purpose on the user's behalf.
| Source terms | --use unstated |
--use non-commercial |
--use commercial |
|---|---|---|---|
| forbids sale | skip + warning | fetch, record the declaration | refuse, fetch nothing |
unknown (None) |
skip + warning | skip + warning | skip + warning |
| permits | fetch | fetch | fetch |
Unknown never becomes permission. A source whose terms could not be established has not been shown to permit anything, and "we could not read the terms" is not a finding that they forbid anything either — so it is skipped in every column, with a message that says which of the two it is.
The refusal lives here, at acquisition, because under a data-usage policy that is when the terms are accepted, and because refusing here means nothing is fetched rather than merely nothing written.
A module's own recorded declaration counts when the flag states none (S105, RM252). Every gate that
has a module directory — pgx's two legs, the CPIC/ClinPGx/ClinVar/AlphaGenome drafters and checks —
goes through licensing.effective_declared_use(spec_dir, terms, declared_use) before
check_declared_use: the flag when it states one, otherwise the row an earlier run recorded for that
source at that layer in the licence table, otherwise unstated. So pgx on a module drafted under
--use non-commercial no longer says "cpic forbids sale and no use was declared" about the file it
has just read; it runs the CPIC leg, reports recorded_use={"cpic": "non_commercial"} (the summary
line says where the declaration came from), and PharmVar — which has no row — still asks. Three
things it is not: not a default (unstated on disk is not a declaration, and with no row the tool
asserts nothing); not per module (a declaration for CPIC says nothing about PharmVar); and not a change
to the flag, which outranks the file in both directions — --use commercial against a recorded
non_commercial still refuses. The cache lanes (cache pull, the builders) have no module and gate on
the flag alone, and test_declared_use_recorded.py walks every check_declared_use call site to keep
the two sets exact.
Where the terms come from¶
The per-source SourceTerms constants (CPIC_TERMS, PHARMVAR_TERMS, …, collected in
licensing.TERMS_BY_SOURCE) hold only the residue that cannot be read from a payload. Where a source ships its
own licence — ClinPGx bundles a LICENSE.txt inside every archive — the pass reads it from the same
bytes it took the data from and records license_sha256, which makes the recorded terms provably
contemporaneous with the recorded data instead of a lookup that was true once. Both halves of a static
table went stale inside one release (the retired hostname, the moved licence page); a hash turns the
next such change into a finding.
The compiler holds no source→licence map — that would give it a source convention (Principle 2) and an un-injected reference. It reads only what the enricher recorded.
A grant may be a floor rather than a total, and two constants are written that way. The PGS
Catalog's terms are per score. MITOMAP's page grants CC BY 3.0 for all of Mitomap.org "unless
otherwise noted" (MITOWIKI/HelpTerms r5, read from a browser on 2026-09-03 — the data surface serves
the dump to plain curl while the web surface 403s), so MITOMAP_TERMS states what the host grants
and never that every cell in the dump carries it: a per-record note outranks the site default
(@a-hosts-terms-are-not-its-contents-terms). Its commercial_use=True is stated rather than
inferred from the CC grant — the page names individuals, clinical labs and commercial services, with no
permission and no fee — which is why that axis is a True and not the None an unestablished
permission gets. One trap, since a search reproduces it: MITOMAP's NAR article is CC BY-NC, and that
is the paper's licence, not the database's.
On a host, or in a service — the gated sources need a cache (RM38)¶
Everything above assumes the shape this tier was written for: an author runs the enricher on their own machine. That case is fine as it stands, and the reason is worth naming, because it is exactly what stops being true elsewhere. The author accepts the source's terms themselves, spends their own rate budget, and holds their own PharmVar key.
A hosted enricher — the same functions behind an HTTP endpoint, enriching modules for whoever is calling — is a different act, for two independent reasons. Either one alone justifies the conclusion, and they have different consequences, so they are worth keeping apart:
- Terms and identity. The operator's acceptance stands in for every end user's, and the operator's
personal, non-transferable PharmVar key becomes the key third parties query on. There is no per-user
switch: the presence of
PHARMVAR_API_KEYin the server's environment is the switch. - One shared budget. Every published figure in the rate table is per IP, so a server does not get a budget per caller — it multiplies its callers onto one allowance. PharmVar's 2 rps and gnomAD's 10-per-60s are the whole deployment's, and an overspend limits the service rather than the caller who caused it.
The resolution half always handled this, which is why only the gated sources were at issue. Ensembl,
ClinVar and gnomAD constraint each have a snapshot, a locations resolver and an ensure_*, so a hosted
enrich() is cache-served and --offline is genuinely zero-egress — and all three are
commercial_use=True, so there is nothing to gate anyway. The gated set is precisely the three PGx
sources, and until 0.5.1 none of them had a working cache path:
| Gated source | Snapshot builder | locations resolver |
ensure_* |
Runtime pass |
|---|---|---|---|---|
| ClinPGx | clinpgx_build.py (shipped 0.5) |
✅ 0.5.1 | ✅ 0.5.1 | resolve → provision → skip with a reason |
| CPIC | ✅ cpic_build.py (0.5.1) |
✅ 0.5.1 | ✅ 0.5.1 | snapshot → live → skip with a reason |
| PharmVar | ✅ pharmvar_build.py (0.5.1) |
✅ 0.5.1 | none, by design | snapshot → live → skip with a reason |
Before that, --offline was a no-op for pgx (it warned and returned, because there was nothing to
fall back to) and was absent entirely from draft, so a hosted PGx path had two options — fetch, or
skip the check — and neither was the one it wanted. See The caches above for the operator's side.
The rule this tier follows: a hosted surface reaches a gated source through a snapshot the operator
built once, never live per request. --offline is the only switch and an explicit --snapshot /
--*-cache path is the inject-only escape hatch — the same shape the ClinVar and constraint snapshots
already use, and deliberately not a second flag.
One asymmetry in publishing such a snapshot. The recorded terms permit redistribution for all three,
so ClinPGx and CPIC snapshots follow the full ClinVar pattern — build, publish, ensure_*. PharmVar
cannot: bulk data pulled under a personal, non-transferable key is not covered by any axis the terms
record, and an unestablished permission is never a permission here — the same None ≠ False rule that
governs share_alike and commercial_use, applied to a clause no column models. A PharmVar snapshot
therefore stays operator-built and inject-only: a resolve_pharmvar_reference and a builder, and
deliberately no ensure_pharmvar_snapshot and no pharmvar publish.
And a coordinate bug the snapshot turned from latent into written. PharmVar publishes each defining
variant against both assemblies and lists GRCh37 first; _merge_variants was first-wins over any
NC_ row, so 451 of the 739 rsID-keyed defining variants would have carried a GRCh37 position
(DPYD rs868235016 as chr1:97547910 rather than its GRCh38 place). The accession version cannot separate
them — chr10 is .10/.11 and so is chr22 — but referenceCollections does, exactly. Nothing consumed
PharmVarAllele.variants before, which is why it never bit; a snapshot stores them. See
pharmvar.PHARMVAR_GENOME_BUILD.
CPIC does publish a chromosome, on gene.chr. An earlier probe read sequence_location alone — which
genuinely has none — and concluded CPIC publishes none at all, so the drafting provider skipped every
defining variant CPIC gives no rsID for: 18 in CYP2C9, 14 in TPMT, 4 in NUDT15. Joining gene.chr onto the
symbol the location row already carries is a lookup in CPIC's own tables, not the inference that probe
rightly refused, and draft --gene CYP2C9 now writes 17 coordinate-only haplotype rows it used to drop.
Every pass records what it consulted — and a link is not a source (RM33)¶
licensing.record_source_terms(names, layer, path) is the one place that turns "this pass consulted
these sources" into sources.csv rows. Three passes were missing it entirely, which is why
VALID_SOURCE_LAYERS had members nothing ever wrote:
| Pass | Layer | Sources recorded |
|---|---|---|
enrich (resolution) |
resolution |
whichever of ensembl / clinvar / gnomad answered |
frequencies |
frequency |
gnomad |
gene_metrics |
gene_metrics |
gnomad (and clingen writes its own dosage row) |
None of these layers can taint a module — only annotation does, because a coordinate or an AC/AN is a
fact the source reports rather than expression it owns. So what these rows carry is attribution,
which gnomAD, Ensembl and ClinVar each request and none enforces, and that is as much what the table is
for as the prohibitions are. declared_use is unstated: none of them forbids sale, so these passes
never have to ask.
resolution.csv records both a link and an authority. source names which link answered
(ensembl-rest, cache, …) and authority names the licensed source it speaks for (ensembl), which
is what sources.csv joins on. RESOLUTION_AUTHORITY_BY_LINK is that map, and it lives here rather than
in the compiler for the reason stated one paragraph above. A link with no entry — authored, reversed,
manual — keeps an empty authority, because the module's own bytes are not a licensed source.
gene_metrics.csv had the same overloading and was fixed the other way: it records gnomad, and which
release answered stays in dataset, where the two-constraint-routes distinction already lived.
A gene panel is drafted, never decided¶
clinvar_draft.draft_gene_panel is the provider RM4 waited for, in the shape the charter allows: it
drafts rows a human then owns, with no compile-time reference materialization. It was blocked on a
real problem rather than on effort. VariantRow.genotype is required and ClinVar publishes
alleles, not genotypes — whether carrying a pathogenic allele once is informative (a carrier, an
affected proband, neither) follows from the condition's inheritance mode, which ClinVar does not
state. Writing A/G because the alt is G would be a clinical claim the source never made;
reference_examples/pathogenic_clinvar/ is a human having made that call by hand, row by row.
So the provider writes a partial row: everything ClinVar publishes, with genotype carrying
vocab.TEMPLATE_PLACEHOLDER, which no mode compiles. The panel is authored, in place, in gene order,
and loudly incomplete until someone decides. A re-draft after those decisions adds nothing, because a
partial row matches on the identity columns rather than on the natural key — that key runs straight
through the column still holding the stub.
The alleles the author writes that genotype from stay in the report, and since 0.7 the report can
be asked again (RM71). A drafted row is rsID-only, so rs118203998 arrives with empty ref/alts
and the pair is stated only in the warning stream — one line per open row, uncapped, because each is a
task. Every place the alleles could legally be written is the wrong place: identity is filled whole
or not at all, alts is redundancy-bearing (filling it makes the compiler compare ClinVar with
ClinVar), a comment column is full cost on the most expensive table and rots at the next release, and
an unread sidecar becomes a permanent resident of every drafted module. So the repair is scope, not a
column. The worklist covers every row in the file whose genotype is still the placeholder, not
only the rows this run added, so a second draft-panel reprints it instead of falling silent — and
--dry-run obtains it without appending anything, meaning exactly what it means on draft. The
earlier narrowing survives for a better reason than scoping gave it: a row the model refused was never
written, and a row the author has settled no longer carries the stub. The state line moves with it,
since it reads as "these rows also". Where the run holds no record for an open row — usually another
gene, or a tighter --clin-sig/--min-review-stars, though the pass cannot tell which — the alleles
are genuinely unknown, so they are withheld and the rows are named by gene in one aggregated line
rather than guessed at or dropped from the count. The line states only that nothing this run selected
covers them; naming a cause it did not establish would be a guess dressed as a finding.
A scaffolded template row is not drafting work. scaffold writes a variants.csv row carrying
the placeholder in rsid as well as in genotype, and file scope would otherwise sweep it in — as an
inflated count, as a row supposedly belonging to some other gene, and as <<REPLACE>> in a list of
row labels. A row whose identity is itself the placeholder is skipped here: it has no identity to
match, so no run can ever hold its alleles, and the compiler already refuses it by name.
Two rules worth keeping in view. Identity is filled whole or not at all: the rsID, else the
complete coordinate, never a subset, because a lone alts on a position-only row makes
derive_variant_key mint a VRS ga4gh:VA.… id instead of chrom:start:ref — a partial coordinate
silently changes which variant the row is.
An rsID is only usable as an identity when it names one allele event here, and the predicate that
decides that is multi_allelic_rsids — which keyed the site on ref until S41. It grouped by
(rsid, chrom, start, ref) and fired on more than one alt inside that group, which reads as more
than one alt at one position and is not: an ordinary ClinVar dup/del mirror pair, A>AT beside
ATT>A at one position, is two groups of one alt each. So the rsID was never flagged, both records
took the same rsid-only identity, and append_partial_rows dropped the second as already_present —
a silent loss, since dedup is the normal case and nothing distinguishes it from a re-draft. The
event is now the whole (chrom, start, ref, alt), so a differing ref and a differing position both
flag. Measured on the 2026-06-27 snapshot over BRCA1/BRCA2/ATM/MLH1/MSH2: 942 rsIDs flagged before
and 1,589 after, the 647 new ones exactly the 647 collapsing identities, 725 records recovered, no
record made unkeyable. 187 of those collapses had dropped the better-reviewed record — a
consequence worth stating separately, because select_by_gene orders by ref before
review_stars DESC, so which of the pair survived was decided by allele spelling rather than by
evidence. Distinctness is over the allele event and not over records: a re-submission under a second
variation_id is one claim written twice, and coordinate identity could not separate those anyway.
A re-draft repairs the omission and cannot retract the collapse — so the drafter now names what it supersedes (S45). "Your module needs a re-draft" reads as a complete instruction and is not one.
The distinction that decides it: a drafting fix either skipped rows or wrote them under an
identity that has since moved, and only the second leaves anything behind. S44's widened genotype
gate is the first kind — measured on the same afternoon, a stale ClinPGx module of 18,691 rows
re-drafts to 18,895 with 0 missing and 0 stale, byte-for-byte a fresh draft, so it needs no caveat
at all. S41 is the second kind, and the two shipped in the same release, which invites a reader to
generalise the wrong remediation.
Drafting appends and never mutates, so re-running the fixed drafter over an existing spec adds the
coordinate-keyed rows beside the collapsed rsid-only row rather than replacing it: the module then
states both the right answer and the wrong one for the same locus, and the stale row is the one
carrying the mislabelled expansion. Reproduced on MLH1 at min_review_stars=2: a stale module of
996 rows re-drafts to 1,061 against a fresh draft's 1,030 — 0 identities missing (every dropped
record recovered) and 31 present that a fresh draft does not contain.
Those 31 are invisible from inside the module: _row_cells writes no rsid on a coordinate-identity
row, so the obvious predicate — an rsid-only row whose rsID also appears on a coordinate row — finds
0 of 31. What can see it is the pass itself, which holds ambiguous: an rsid-only row carrying an
rsID this run is deliberately writing by coordinate is a row no current draft would produce.
_superseded_rsid_rows reports them, aggregated through examples, after the append.
Reported, never removed, and that is the reporter's own preference for the reason they gave: a
drafted row is authored material by the time a re-draft runs, a human may have curated its genotype,
state and conclusion, and deleting curated work to repair a drafting defect is a trade only the
author can make. Drafting-appends-never-mutates is the rule one file over, and a provider that started
deleting rows would be the exception to it. The cleaner remediation is still a fresh directory
reconciled against the old one; the notice exists for the author who did the other thing, which is
what the instruction invited. And min_review_stars defaults to 2: a panel that
mixes a 0-star "no assertion criteria" submission with a 3-star expert-panel review without saying so
is worse than one that names its floor.
What it does not fill is as deliberate as what it does — no weight, direction or effect statistic
(ClinVar publishes none), no trait_efo_id (its condition is free text and MedGen, not EFO), no
acmg_sf, and no curator/method (the spec's defaults: block owns those).
A non-diploid contig gets its genotype written, because there is no judgement there to protect
(0.5.2). MT is haploid and chrY outside PAR1/PAR2 is hemizygous: exactly one genotype is expressible
per allele, so sole_expressible_genotype writes the ALT and the row arrives complete. The rule the
placeholder encodes is unchanged — it exists for a decision the source does not make — and this is the
case where that decision does not exist. Y is decided per locus through the same three-valued
vrs.in_pseudoautosomal_region the compiler's ploidy guard uses (XG and SPRY3 straddle a
boundary), and both True (diploid) and None (no PAR table for this build) keep the stub, which is
the house rule about unknowns. The run reports what it committed to in one aggregated line naming
the contigs: those rows read as homoplasmic/hemizygous, and a heteroplasmic level is a different
question with its own table kind. Without this every consumer rediscovered it the hard way — one
wrote A/G and A/A across 264 mitochondrial loci in a genome-wide panel and 260 in a cardiac one,
each asserting a second copy that is not there.
A citation ClinVar files under PubMed that is not a PMID is skipped and counted, never raised.
218 of the 3,952,341 PubMed rows in the 2026-06-27 file carry a nine-digit id (Variation 12606 cites
168335863; PubMed is at eight), and StudyRow.pmid rightly refuses them — but the refusal used to
surface as an unhandled ValidationError that aborted a 297-gene draft over one row in one gene. The
builder now drops them at the snapshot boundary using the format's own extract_pmids grammar rather
than a second opinion restated here, and the drafter survives one anyway, since every snapshot already
published carries them. The two shortfalls are reported apart: --max-citations is a choice this run
made, an unusable id is a defect in the source.
The snapshot is found, then provisioned — --snapshot used to be required (0.5). _resolve_snapshot
runs the ladder enrich() uses: an explicit path is taken as given (the inject-only escape hatch, and
what an air-gapped run passes), else the cache locations, else the published snapshot is downloaded unless
--offline. Until this, the published snapshot could not reach an author at all — they had to build 4.4M
records from a 200 MB VCF or already know the cache path — which mattered most for the citations,
since they are what makes a drafted panel compilable and they only started travelling with the snapshot in
the same release. Verified end to end from an empty cache: draft-panel --gene HFE provisions
data/ + citations/ + release.json and drafts 12 variant rows with 33 grounded study rows carrying
real PMIDs, then refuses to compile on the genotype placeholders — which is the designed state, not a
failure. No snapshot and --offline raises rather than drafting nothing: an empty draft would read as
"ClinVar has nothing for this gene".
Drafting from PubMind — a source with no gene column (0.7, RM134 § C)¶
draft-panel --source pubmind writes the same rows into the same table from the same gene argument,
so it is a flag on the existing command rather than a draft-pubmind beside it. The parts that
are hard to get right — the genotype worklist, the placeholder guard, the dedup-against-the-file pass,
the refusal summary — are the parts a twin command would have to carry a second copy of, and the
worklist's file-scoped seam above is the one this release had to repair. draft-clinpgx is a separate
command because it writes different tables; this does not, so pubmind_draft imports the ClinVar
provider's machinery instead of restating it.
PubMind names no gene, and this pass does not invent one. The snapshot is
(chrom, start, ref, alt) and nothing else locational. Turning --gene BRCA1 into positions needs a
gene→locus map, and this repo deliberately holds none — the compiler's own gene/locus check is
chromosome-granular for exactly that reason. So the map is ClinVar's own per-record gene attribution
matched at the exact position: every position ClinVar records for the gene, with no clinical or
review filter, because it is a locus universe rather than a selection and filtering it would narrow
what PubMind is even asked about, invisibly. A min/max span over those positions was refused: it
invents a boundary nobody defined and writes a gene cell that is a false claim wherever two genes
overlap. Both snapshots are therefore required, and each absence names its own switch —
$JUST_DNA_PUBMIND_CACHE or --pubmind-cache for one, the ClinVar ladder for the other. A missing
ClinVar snapshot raises this pass's PubMindDraftError rather than leaking ClinVar's, and the
message says what the second snapshot is wanted for: told "no ClinVar snapshot" by a PubMind command,
an author has otherwise been handed a puzzle.
The cost of that choice is stated rather than counted: a PubMind verdict at a position ClinVar has no record for cannot be reached by gene at all, and that class is not countable, since attributing it to a gene is precisely what there is no map for. It includes many of the codon-decomposed offsets.
Identity is the whole coordinate or nothing. The snapshot has no rsID column and most of the
source's rows carry no rs-number, so chrom/start/ref/alts go in together. The row still matches
on the same five identity columns a ClinVar-drafted row does, so a coordinate both sources speak about
is one row in the file rather than two.
A position two requested genes both claim leaves gene empty, counted and named. Both other
readings are wrong to write: BRCA1, BRCA2 in that cell is not a symbol check-identifiers can
resolve, and picking one is the gene model this pass went to ClinVar precisely to avoid inventing. The
row is still drafted — gene is optional and the coordinate is the identity — and the author is told
which coordinates and which genes, so filling it stays their call.
Five classes never become a row, and each is named at draft time. A contested key, a
length-changing row, a call outside --clin-sig, a confidence below --min-confidence, and a
confidence the source never stated — that last one its own class, because None is not 0 and reading
an unstated confidence as 0 invents a reading. Every candidate key is drafted or withheld under
exactly one reason, and candidates == drafted + Σ withheld is an equality over the walked reason set.
A class that withheld nothing reports no zero.
Contestation is decided over every PVID at the key, before either dial runs. PubMind's record id
keys on the text a model extracted rather than on a coordinate, so one position carries several
records and their calls can disagree. Choosing one needs an ordering nobody defined — mode() over an
unsorted group. A --clin-sig or --min-confidence applied first would remove the dissenting record
and pick that winner just as effectively, so the filters run second and the coordinate, its PVIDs and
its competing calls are all reported instead. Where the records agree, one row is written and the
record count, the PVIDs and the best stated confidence stay visible in the transcription.
No study row is drafted, and that is said out loud. The ANNOVAR-redistributed channel carries no
PMID and their API withholds per-record detail, so a PubMind draft grounds nothing — studies.csv is
mandatory, so the run names the gap rather than leaving the author to meet it at compile.
Unknown terms warn and never gate. check_declared_use is a gate on fetching and its unknown
branch skips, which is right for a pass that would go and get data whose terms nobody can state.
Nothing is fetched here: there is deliberately no ensure_pubmind_snapshot, the operator built the
snapshot with pubmind build, and refusing to read it would make that command's output a file nothing
may consume. The reason is reported in the source's own words instead, and the licence row records
None on every term — which does not taint, because taints_commercial_use requires an explicit
False. What the unknown answer governs is publishing such a module.
A module drafted from PubMind must not be able to confirm itself. pubmind is in
DRAFT_PROJECTIONS, projected onto clin_sig, so the drafter stamps a digest of the column the
concordance check later reads and the check can establish the copy rather than assume it
(@draft-digest). Its identity is the coordinate and not the provider's match_on: the source
states no rs-number, so an rs-number an author later adds is a change to the row's spelling and not to
the call.
What it does not fill is as deliberate as what it does. No clinvar, pathogenic or benign — all
three are ClinVar flags by their own field descriptions, and a position appearing in ClinVar's gene map
says nothing about whether this allele is in ClinVar. No phenotype, because the channel carries no
condition. state is folded from the source's own call exactly as it is on the ClinVar path, and left
as a stub for any call the fold does not cover.
Drafting from MITOMAP — the increment, never the photocopies (0.7, RM171)¶
just-dna-enricher mitomap build --out data/caches/mitomap # the pg_dump → parquet
just-dna-enricher mitomap miss --out data/caches/mitomap_miss # the join, from both parents
just-dna-enricher draft-panel spec/ --source mitomap-miss --use non-commercial
The fourth --source on draft-panel, and the first that is not asked for by gene. The other
three draft a panel: ClinVar's snapshot is 4.4 M records and CIViC's is a cancer corpus, so an
unfiltered draft from either is not a panel but the source. This one's snapshot is the increment,
so --gene filters where it exists and is not required; the other three still refuse an empty one.
What it writes. Identity as MITOMAP publishes it (chrom=MT, start, ref, alts), the gene
where locus names exactly one, the disease string verbatim as phenotype, and clin_sig from the
bracketed ClinGen mtDNA VCEP rating through the one shared normalizer. state is folded from that call
where STATE_BY_CLIN_SIG has an answer. Study rows come from MITOMAP's own reference.nlmid links and
are position-keyed, like ClinVar's, because a study is evidence about a locus.
What it refuses, and each for its own reason — three sentences rather than one skipped count, because they send an author somewhere different:
| bucket | why nothing is written |
|---|---|
| photocopy | the same event is in ClinVar under either spelling (indels left-aligned against a vendored rCRS on both sides since RM273), so that VCEP call already reaches this repository with ClinVar's own provenance. Drafting a second copy attributes it to the wrong publisher, and hands a ClinVar concordance check a copy of ClinVar to agree with (@tautology-zero) |
| unrated miss | absent from ClinVar, and MITOMAP published no class this tier may map — no bracket, a bare confirmation token, or [VUS*]. A real identity increment, counted, with no significance invented for it |
| unmintable | the published alleles do not spell a VCF pair: prose in an allele column, or no event. A : deletion is anchored on the vendored rCRS base at position-1 since RM293 and joined like any other row, unless MITOMAP's deleted bases disagree with rCRS |
genotype is a placeholder, and the reason is not the contig's. clinvar_draft.sole_expressible_genotype
fills the ALT on chrMT — a haploid contig leaves no zygosity open, so there is no decision for a
placeholder to protect (S6). That is right about ClinVar, whose record is a claim about an allele.
MITOMAP's row is a claim about a literature corpus: homo and hetero say whether the variant has
been reported in each state, and a share of the increment is reported only heteroplasmically. Writing
genotype=<ALT> there states the homoplasmic reading, which is exactly what
reference_examples/mt_heteroplasmy keeps in variants.csv and separates from its heteroplasmy.csv
bins. So the cell is stubbed, conclusion with it, and the draft prints one uncapped worklist line
per row carrying the alleles and the flags MITOMAP did publish. A MITOMAP-drafted module does not
compile until a human writes those cells; that is the cost of the adoption rather than a defect in it.
Four more things the draft says out loud. A drafted row keying on an indel is named. Since RM273
the lane compares events, so such a row is absent from ClinVar under any spelling; an increment built
before that compared spellings, and on the real lane five of the six rated misses were ClinVar's own
calls at another anchor, so the note tells an author holding an older build to rebuild it. A row whose allele name
states a variable number of copies while the allele columns state one definite pair is named too: the
source disagreeing with itself, kept rather than repaired, because rewriting it needs a rule for what
(n) means that MITOMAP has not given. A parent that has moved since the increment was built is
reported and the rows are still drafted — a stale increment is the increment against the older parent,
which is a fact worth stating and not an authoring error. And the SourceRow names mitomap, never
the derived lane: the computation is this repository's, the content and the attribution duty are the
source's, and dataset carries both parents because a derived artifact's identity is the pair it came
from.
Lookups answer, they never fill¶
lookup.py is the authoring counterpart to the passes above: same clients, same offline-capable
snapshots, and no writes at all — not a sidecar, not a cell. It answers what an author actually
asks. For an rsID: is it live, merged or absent (dbSNP is the oracle; Ensembl returns HTTP 400 on
some merged ids and would misclassify them), which coordinates it maps to, and — on demand —
whether that answer is ambiguous. For a coordinate: ref, alts, gnomAD populations with the
frequency computed as allele_count / allele_number (the API deliberately exposes no af),
ClinVar's own call, and — since 0.7 — PubMind's (RM134 § D). For a citation: whether PubMed has the record, plus the DOI and PMC id that
arrive free in the same response, with Crossref covering what PubMed does not index at all.
The PubMind leg is three-valued, and it may not fill the cell it reports on. clin_sig is what
the concordance check cross-examines, so a hint filling it from one of the authorities being compared
would make the check agree with the source it is checking (@hint-redundancy-bearing) — the same
defect @draft-digest handles one layer down, and here there is no digest to rescue it. So every
record comes back as an advisory with applied=False. No snapshot is nobody asked and says so,
naming $JUST_DNA_PUBMIND_CACHE: it is operator-built and there is nothing to download, so an empty
answer would otherwise read as "PubMind states nothing here". A snapshot holding no record at the
allele is the third state and is reported as an absence in their corpus — no paper survived their
triage, which is not a benign call and not a disagreement. Where their own records disagree, every one
is reported and none is picked. Unlike the ClinVar leg it also answers for a coordinate the caller
typed rather than only for one an rsID resolved to, because their channel is coordinate-keyed and most
of its rows carry no rs-number at all.
Every one of those comes back as an Alteration with applied=False and a refusal naming why the
value is the author's to type. That is not fastidiousness. resolution._verify compares an authored
coordinate against the table, sequences.verify_reference_alleles compares an authored ref against
the genome, literature._doi_conflicts compares an authored DOI against the registry — each has
force only because the human wrote the value independently of the oracle. Filling the cell from the
same source the checker consults makes the check vacuous, and for an rsid-only row _verify does not
run at all, so the row would move from honestly unverified to apparently verified. This package had
already made the argument for one field: Crossref is asked about the authored DOI because a
derived one "exists by construction".
Two operational notes. Clients are injected and reused (LookupClients) because each owns its
own PacingGate — gnomAD is one request per six seconds — and a fresh client per question discards
both the pacing state and the connection pool. And --offline yields unchecked, never absent:
a check that could not run is not a check that passed, and None is not False anywhere in the file.
An unfilled field is filled the same way on every leg (S92, RM206). LookupClients.ensure(name,
factory) builds a client under the bundle's lock on first use, stores it and returns it, and
close() walks CLIENT_FIELDS — derived from the dataclass, not listed — to close what was built.
Until then two legs assigned back onto the bundle and six built a per-request client they closed in a
finally, which is exactly the pacing state the paragraph above says to keep; a host filling seven
fields and leaving pmc_idconv unset had unpaced egress on one leg and no way to tell from the call
site. Ownership follows construction: hold a bundle for a session, and a lookup_* call given none
closes the one it built. A hosted CPIC draft is deliberately not a field here — pgx_draft.draft_gene
takes client=, so a host shares pacing by holding one CpicClient and passing it.
A cache miss falls through to live Ensembl (0.5), and until it did this surface was silently
weaker than the pass it advises on. hint variant --rsid rs1799945 answered "not found in Ensembl,
position remains unset" for HFE H63D — which live Ensembl serves at 6:26090951 — because the only
thing it had ever searched was a local snapshot that did not contain it. Two things were wrong and both
are fixed: the live link (V2 GraphQL → V1 REST) now runs on a miss, in enrich()'s own order so a
provisioned snapshot still costs no egress; and the snapshot's own warning stopped speaking for
Ensembl, saying "not in the injected Ensembl snapshot" — the thing it actually searched. Naming the
searched thing is the difference between "we did not look there" and "it is not there".
That warning kept a second clause it had no standing to write, and RM125 removed it. It read
"not in the injected Ensembl snapshot, position remains unset", and the trailing half is a claim
about a leg that has not run: online the live link follows and usually fills the coordinate, so
lookup_variant returned rs4988235 at 2:135851076 beside a finding saying the position was unset.
That is the ordinary three-valued failure — at emission the outcome is genuinely unknown, and the
house rule is to withhold rather than guess. Each link now reports only what it searched, and
lookup_variant, the one caller that sees both halves, states "{rsid}: position remains unset" once
at the end when nothing placed the variant, guarded on there being an rsID at all (a position-only
question fills rsid_candidates, never loci). The snapshot miss stays on the record either way,
because knowing a local cache is incomplete is what tells an author whether to warm it.
enrich() is untouched by this: it discards both links' warning lists, so lookup_variant was the
only reader either sentence ever had. The ClinVar twin needed the same repair and had not had the
first one — clinvar.lookup_loci is documented as signature-identical to resolver.lookup_loci,
"one implementation, no drift", yet it still said "not found in ClinVar, position remains unset",
speaking for the source exactly as the Ensembl half had before it was fixed. Both now say "not in the
injected <source> snapshot". A pair described as one implementation is still two strings, and only
one of them was reviewed.
A live locus is labelled live. The advisory rows were hard-coded to source="snapshot", which
became a lie the moment the live route landed: a network answer claimed to come from a pinned file, in
the one field an author reads to judge how reproducible the answer is. They now carry
ensembl-graphql/ensembl-rest, and a finding says re-running may differ as Ensembl advances. It is
still advisory — a live answer is not a licence to fill a redundancy-bearing cell.
Generation is not automatic¶
The PGx tables are authored _TABLE_KINDS, not fact sidecars: they carry AuthoredModel semantics,
the reserved-namespace guard and raw-byte input hashing. A network pass writing them would blur the
authored/derived line the 0.5 rework drew, and hand the human author a file they never wrote but are
accountable for. So the automatic pass only ever reads, and scaffolding is an explicit,
separate step.
PharmVar and CPIC gotchas¶
- The header is
Api-Key, noX-prefix, documented per-endpoint in the service's own OpenAPI document (docs/vendor/pharmvar_api_docs.json) rather than in asecurityDefinitionsblock. Every wrong spelling returns the same 401 as no key at all, so a wrong header is indistinguishable from a bad key — the error message says so rather than guessing. - The key is personal (PharmVar terms §2), so it comes from
PHARMVAR_API_KEYand is never written into a module, fixture, log or snapshot. - 2 rps, enforced by the shared
net.PacingGateon an injectable clock. The unfiltered/allelescollection is ~25 MB and silently ignores ageneSymbolparameter it does not define; use/genes/{symbol}. - CPIC
variantallelecarries two shapes the allele grammar cannot hold, and they are different findings. IUPAC ambiguity codes (Rat CYP2C19*2,Yat*4) are an uncertainty CPIC recorded — expandingRwould invent two defining variants where CPIC recorded one — and will never be expressible. Deletion/insertion and repeat notations (DELTCT,AAAGGGGCG(2),GGA(1), 23 of them in CYP2D6) are a grammar gap a release could widen to cover — and note that RM5 was not it: 0.6 widened the grammar to hold VCF's five symbolic structural alleles (<DEL:1500>), which is a different spelling from CPIC'sDELTCT. Both are skipped, never coerced;cpic.unusable_allele_reasonnames which, and they are reported as two aggregated lines with counts. Calling the second kind an ambiguity code — which the message did until a real CYP2D6 draft — is a false claim that points an author at the wrong fix. - A large star-allele gene needs
--allele(RM34).draft --gene CYP2D6unfiltered is 16,290 diplotype rows, 73%Indeterminate: faithful, and unreadable.--alleletakes the set the consumer's caller can actually emit (n alleles is n(n+1)/2 pairs, so six make 21 diplotypes) and applies it to all three tables at once, because a module that names an allele it never defines is what_cross_validate_haplotype_definitionsexists to warn about.*1is always kept — it is defined by carrying no variants, so it costs nothing, and without it*1/*2could not be drafted. An unknown allele refuses with the list CPIC publishes; the flag needs a single--gene, since*2in CYP2C9 and*2in CYP2C19 are different alleles. - CPIC activity scores are inequality strings (
"≥3.0","n/a"), not numbers, so they do not drop intoMeasureBinRow's numeric bounds; the raw string is carried and the parsing left to a human. - Coordinates are 1-based in both (verified against Ensembl for rs4244285 → chr10:94781859, which PharmVar, CPIC and our own resolution all agree on). Do not convert.
- CPIC recommendations are keyed by (gene phenotype, drug, clinical context) and the contexts
disagree — and since 0.5 (RM29b) that is no longer a refusal.
draft --drugused to stop and list the choices when CPIC scoped a pair to several settings, becauseDiplotypeRowhad nowhere to put the distinction: writing all of them collided on the duplicate-row key, and writing one asserted a clinical setting the author never chose.DiplotypeRow.clinical_contextis now part of that key, so every setting is drafted as its own row and the consumer picks — which indication a patient is being treated for is knowable at query time and not at authoring time. Drafted live,--gene CYP2C19 --drug clopidogrelyields 1,998 rows acrossCVI ACS PCI,CVI non-ACS non-PCIandNVI, and the disagreement is visible in them:*2/*2Poor Metabolizer isstrongin the first andmoderatein the other two, with different prescribing text forNVI.--populationsurvives as a filter for an author who wants one setting; an unknown value is still an error, since drafting nothing on a typo would look like "CPIC has no recommendations here". Three of CPIC's sixteen live context values carry trailing whitespace, so the column strips on load — unstripped,'CVI ACS PCI 'and'CVI ACS PCI'are two rows describing one setting.recommendation_strengthis still CPIC's andevidence_levelstill PharmGKB's; a provider fills only its own.
Pass 6 — ClinPGx summary annotations (clinpgx.py, offline capable)¶
pgx.py asks the nomenclature authorities about star alleles over the network; this pass asks
ClinPGx about summary annotations — which variant, which drug, at what evidence level — from a
local snapshot, exactly as the ClinVar cross-check does. clinpgx_build is the [dev] builder.
ClinPGx called these clinical annotations until 2025-07-29 and both names are still in use on its
own site; the archive, the members and the id column all say summary now.
The snapshot is read with duckdb, not polars, and that is deliberate: polars is a [dev]
dependency here (only the builders need it) while duckdb is core, so reading with polars would leave
this runtime pass unusable on a plain pip install just-dna-enricher. clinvar.py reads its snapshot
the same way for the same reason — the builder may be dev-only, the pass may not.
The snapshot pins its own licence. ClinPGx ships a LICENSE.txt inside summaryAnnotations.zip,
so the builder extracts it, records its sha256 in release.json, and the pass stamps that hash onto
the SourceRow. The recorded terms are provably the ones shipped with the recorded data — the
property a static source→licence map cannot offer.
The snapshot's grain is (annotation, genotype), joining summary_annotations.tsv to its
per-genotype child summary_ann_alleles.tsv. CREATED_<date>.txt is the release id, because
ClinPGx publishes no version number and does not refresh its archives in lockstep —
relationships.zip was a year newer than clinicalAnnotations.zip when this was written, which
RM175 later explained: that archive had stopped being rebuilt at all.
The archive is identified before it is read, and the retired spelling is refused (RM175). ClinPGx
renamed clinical annotations to summary annotations on 2025-07-29 and the archive followed —
clinical_annotations.tsv → summary_annotations.tsv, clinical_ann_alleles.tsv →
summary_ann_alleles.tsv, Clinical Annotation ID → Summary Annotation ID, with the evidence and
history siblings renamed the same way (this builder reads neither). clinicalAnnotations.zip was last
written on 2025-07-05, is on no downloads page, and the API still answers it 200 through a 303 to
that frozen object — so until this item every snapshot the lane built came out of a database fourteen
months old, and nothing in the download said so. require_current_archive therefore reads the
archive's member names first and answers in three arms: the current spelling builds, the retired one
is refused with the rename, its date, the retired filename and the URL to use instead, and an archive
that is neither says that instead. just-dna-enricher clinpgx build prints the refusal as
CLINPGX BUILD FAILED: … and exits 1.
Nothing about the table itself moved: the other fourteen columns are identical in name and order, and
Phenotype Category carries the same values with the same ; separator, so no vocabulary member, no
model field and no parquet column changed. The data moved, which is the point — 16,087 → 16,117
snapshot rows, 22 (annotation, genotype) keys gone and 52 new, 30 rows changing evidence_level and
every shared row rehosting its URL from pharmgkb.org to clinpgx.org — so a module drafted from
this lane can now see an evidence level move under it. That is the check working, not a regression.
The cross-check keys on the annotation, not the triple, and that is a bug fix rather than a
nicety. (rsid, drug, genotype) is not unique: rs4149056 + simvastatin is Metabolism/PK at 1A,
Efficacy at 3 and Toxicity at 1A. The first implementation indexed on the triple and reported all
three of the reference example's correctly-authored levels as stale. The lookup is now
annotation_id → (rsid, drug, genotype, category) → the bare triple, and when the bare triple
matches several annotations at different levels the row is reported as unchecked rather than
compared against an arbitrary one.
Severity follows the mode ladder, unlike the allele-function check beside it. An evidence level is ClinPGx's own metadata about its own annotation, so a difference means the module is stale — not that two expert panels disagree.
That holds only for a row compared with the annotation it cites (RM297, from S122). Nothing marks a
row as ClinPGx-derived, and annotation_id is the source's own accession whatever the source. So two
cases are reported in both modes and never refused under strict, in ClinPgxResult.withheld:
clinpgx_annotation_not_in_snapshot: the row cites an id the snapshot does not hold. The lookup stops there instead of falling through to a neighbour sharing the category. The warning names three readings (ClinPGx withdrew it, it is mistyped, it was never a ClinPGx accession) and what the snapshot holds at the row's triple.clinpgx_level_differs_from_uncited_annotation: the row cites no id, and the annotation its category or triple reached carries another level.
A cited id that is present but lacks the row's genotype keeps the fall-through and can still refuse.
Both codes are lane-local, not VALID_WARNING_CODES members. They are counted in the record's
findings and named in its detail, so an unchanged module publishes the same number it did before.
Knowing per row that a row cites ClinPGx is RM298's.
The declared-use gate still applies even though nothing is fetched: the terms were accepted when the snapshot was built, and using it is the same act.
del/del is skipped for a different reason since 0.6, and the old one became false. clinpgx_draft
used to report ClinPGx's structural genotypes as something the format could not spell. RM5 widened the
grammar — <DEL:1500> is authorable now — so the block moved: ClinPGx publishes no length, and a
lengthless symbolic allele is a rule the compiler drops. The pass therefore still declines to write
those rows, because a provider must not hand an author work the next command in the documented workflow
undoes, and the warning now names the length rather than the grammar. _CLINPGX_SYMBOLIC maps the
source's dialect (del → DEL) and lives here, at the boundary, not in the schema — the CC →
C/C rule beside it is the precedent, and a grammar that accepted every source's spelling would owe
every consumer the union of them.
But the rule beside it was narrower than the schema it writes into, and that cost two whole families
(S44). _authored_genotype took only CC — two unseparated bases — on the argument that the general
case needs the resolved ref/alt to disambiguate. True of an unseparated cell; false of the two shapes
ClinPGx actually publishes beside it, and validate_allele accepts any ^[ACGT]+$ allele:
- An already-separated call,
CTT/CTT. ClinPGx writes/wherever an allele runs past one base, so the source has already made the split and no decision arises. Declining it cost CFTR F508del among others — those annotations carry adel-spelled genotype and a pure-nucleotide one under the sameannotation_id, so dropping the annotation for the first discarded the second with it. A module shipped 176 CFTR rows and the drug elexacaftor / tezacaftor / ivacaftor while omitting the most common CF variant. - A single haploid allele,
A/CCCCCCC. The hemizygous/homoplasmic form the grammar already holds, and how ClinPGx spells an mtDNA call. Declining it cost every MT-RNR1 annotation — 24 annotations, 48 rows, 32 of them at evidence level 1A: aminoglycoside-induced hearing loss, a CPIC guideline. The reasoning isclinvar_draft.sole_expressible_genotype's, one source over: the placeholder protects a zygosity decision, and on a haploid contig there is none to protect.
Measured on the provisioned snapshot: 158 rows recovered, 36 of them at 1A, across MT-RNR1 (48),
HTR2C, ACE, TYMS, IFNL4, GSTM3 and CFTR. The unseparated multi-base cell the original rule guarded
against does not exist in the snapshot — every multi-base call arrives slashed — and it is still
declined. The general rule is now a test rather than a comment: every spelling this pass declines
must be one PharmVariantRow would also refuse, walked over the accepted set, so a provider can
never again be narrower than the format. The converse stays allowed, which is what keeps del/del
skipped.
And the terms are pinned now. SourceTerms.row has taken license_text= all along; this caller
passed only declared_use and dataset, so a share-alike source was recorded with a null
license_sha256 — the module named ClinPGx's licence without tying it to the text that governed the
bytes, which is the one thing that field exists for. clinpgx_build already extracts LICENSE.txt
beside the parquet, so the fix is to read it. Hashed from the file rather than copied from
release.json's stated hash: the file is what the module is claiming, so a truncated copy cannot pin
to a value it does not have. Absent stays None and warns — an older snapshot predates the extractor,
and inventing a hash would be worse than the null it replaces.
Exception contract — what a caller catches (RM97 + RM101)¶
Two layers, and a caller only ever touches the outer one. A client raises its own error type,
never httpx's (RM97). A pass raises its own type, never the client's (RM101). A consumer calls
a pass, so the pass's type is the one to write in an except; before RM101 five call sites let the
client's type through and the documented handler was silent for exactly the failure it was written for.
| you call | catch | and for "the source could not be reached" |
|---|---|---|
enrich_frequencies |
FrequencyEnrichmentError |
FrequencyUnavailable |
enrich_literature |
LiteratureEnrichmentError |
LiteratureUnavailable |
enrich_gene_metrics |
GeneMetricsEnrichmentError |
GeneMetricsUnavailable |
check_rsids / check_identifiers |
IdentifierCheckError |
IdentifierUnavailable |
enrich_dosage_sensitivity |
ClinGenError |
ClinGenUnavailable |
enrich_gene_validity |
GeneValidityError |
GeneValidityUnavailable |
verify_acmg_sf |
AcmgSfError |
AcmgListUnavailable (carries skip) |
enrich_gwas |
GwasError |
— client and pass share the type |
enrich_pgx |
PgxEnrichmentError |
— degrades per leg instead of raising |
enrich() |
— | — degrades and withholds; see below |
Three layers, really, and the third is where two clients were wrong (RM208). Inside a client the
retrying half is its own function and the translation sits outside it — eutils._request/_get,
cpic._request, gnomad._request, and since RM208 literature.CrossrefClient and
gwas.GwasCatalogClient too. Get that order backwards and one of two things happens, and the tier
had one of each:
- Translate inside the retry and the retry never runs.
CrossrefClient.existscaughthttpx.HTTPError— the superclass of the two types its own@retrymatched — so aConnectErrorbecame theNonewithhold before tenacity saw it. Measured: one upstream request whereattempt_floor(3)asked for three, and the knob a deployment raises moved nothing. - Re-raise for the decorator and translate nowhere and the last attempt leaks.
GwasCatalogClientre-raised the transport leg bare, correctly, andreraise=Truethen handed the rawhttpx.ConnectErrorto a caller told to expectGwasError.
test_retry_is_reachable.py walks every @retry-decorated function in the package and refuses any
whose body catches an ancestor of a type its decorator retries. It walks the package rather than the
client roster on purpose: both offenders sat in test_client_exception_contract.py's exempt set, and
a guard that iterates a roster inherits the roster's exemptions.
alphagenome check raises VariantImpactError / VariantImpactUnavailable (RM193), the pass
half of the pair above — a consumer calls the pass, so this is the type to write in an except.
Note that a transport failure inside it is neither: the pass records the variant as
unreachable and produces no finding, because a service that did not answer has said nothing about
the caller's data.
The Atlas client has its own ladder, and it is a client rather than a pass (RM192). It is listed apart because its third arm is not a failure at all:
| what happened | type | remedy |
|---|---|---|
| transport failed | AtlasUnavailable |
retry later: the client has already spent attempt_floor(5) attempts (RM280) |
REF disagrees with GRCh38 |
AtlasRefMismatch |
fix the caller's data — the server names the real base, which @va-omits-ref says only this tier can discover |
| an indel | AtlasNotScored |
none: the answer does not exist |
a quantile saturated at 1.0 |
AtlasNotScored |
read PHRED from the downloaded artifact instead |
AtlasNotScored deliberately does not derive from AtlasRefused, so an except AtlasRefused
cannot swallow it — a caller that recorded "no score" for a variant the service never claimed to
have scored would be @unreachable-not-absent in one line. AtlasRefMismatch is a subclass of
AtlasRefused, so the same handler-order rule below applies to it, and
test_handler_order_is_not_load_bearing_by_accident enumerates the ladder rather than leaving it to
a reviewer.
Every *Unavailable is a subclass of the type beside it, so except <Pass>Error keeps catching
everything it did (P3, additive within a major) and the narrower catch is new capability rather than a
migration. The client's exception is chained, so __cause__ still carries it — but it is no longer the
only way to tell the two apart, which is what the split is for.
Being a subclass makes handler order load-bearing, and that is the one thing the split costs. Two
separate except arms with the parent first leave the narrow arm dead — Python takes the first
matching clause — so a caller who wrote except <Pass>Error above except <Pass>Unavailable reports
an outage as an ordinary failure and never sets whatever an outage sets. It fails silently: nothing
raises, and a test asserting a status code sees a clean run. One tuple is unaffected, since both names
route to the same block. just-dna-registry lost three of four handlers this way on 0.6.2 (S38), which
is what enricher/tests/test_shadowed_handlers.py now walks the workspace for.
What the distinction means. The subclass says the source was asked and never answered —
nothing was established either way, so a caller records unchecked rather than a negative
(@unreachable-not-absent). The plain parent means the question was put and the answer is a
problem: a strict run with a variant that genuinely has no frequency, or a local sidecar that will
not parse. ClinGenError and GeneValidityError covered both histories with one type until RM101,
and the only way to separate them was reading __cause__ — chained for the fetch, bare for the table.
Do not separate them by message. Neither string is in the pinned warning-text catalogue, so a reword would silently flip a verdict from "unchecked" to "your table is broken". That is the argument just-dna-registry made in S37 against their own first option, and it is why this is a type question.
Two passes deliberately do not raise at all. enrich() catches GnomadError on its last-resort
gnomAD link ("a last-resort link must not sink the whole enrichment") and answers an unreachable rsID
with None from the Ensembl leg; enrich_pgx degrades per leg so one dead source does not take the
other's answer down. Both are the tri-state withhold, which is the correct shape where a pass has
already produced work worth keeping — the opposite end of the same rule, not an exception to it.
If you are adding a pass or a client, enricher/tests/test_client_exception_contract.py and
test_pass_exception_contract.py both discover by walking the package rather than by a list, and
will fail naming your addition until it is covered or explicitly exempted. That is deliberate: RM97's
guard walked a hand-written tuple of eight module names, identifiers was not one of them, and
OntologyClient leaked raw httpx for a whole release as a result.
The complete roster — every error type this tier defines (RM216)¶
The table at the top of this § is the one a consumer usually needs; this one is the registry.
Fifty-one of the tier's error classes were named nowhere in this document, in a section titled what a
caller catches — so a consumer meeting ClinPgxUnavailable or GatedSnapshotError had nothing to
look it up in. Walked by test_enricher_doc_registries.py against the package, so a new error type
joins this table by existing (@registry-completeness).
Read the third column as a ladder, not a set. A subclass makes a caller's except order
load-bearing (@client-exception-contract), so catching the base type first silently swallows every
narrowing beside it. AtlasRefMismatch is the one entry two levels deep, under AtlasRefused.
The *Unavailable convention is the tri-state, spelled as a type: it means the source could not
be asked, which is unknown and never a finding against the module. A base type with no narrowing
either has one failure mode or answers with a withholding value instead of raising — Grch37Client
and EnsemblResolver are the latter, and are exempt from the contract suite for exactly that reason.
Runtime passes — the type a consumer calling a pass writes in its except
| module | base type | narrowed by |
|---|---|---|
acmg |
AcmgSfError |
AcmgListUnavailable |
alphagenome_check |
VariantImpactError |
— |
assertions |
ClinicalAssertionError |
— |
civic_citations |
CivicCitationsError |
— |
clingen |
ClinGenError |
ClinGenUnavailable |
clinpgx |
ClinPgxEnrichmentError |
— |
currency |
ReleaseProbeError |
ReleaseUnavailable |
drug_labels |
DrugLabelError |
DrugLabelUnavailable |
enrich |
EnrichmentError |
— |
expression |
ExpressionError |
ExpressionUnavailable |
frequencies |
FrequencyEnrichmentError |
FrequencyUnavailable |
gene_metrics |
GeneMetricsEnrichmentError |
GeneMetricsUnavailable |
gene_validity |
GeneValidityError |
GeneValidityUnavailable |
gwas |
GwasError |
GwasNotFound |
identifiers |
IdentifierCheckError |
IdentifierUnavailable |
literature |
LiteratureEnrichmentError |
LiteratureUnavailable |
litvar |
LitvarError |
LitvarUnavailable |
mitomap |
MitomapError |
MitomapUnavailable |
pgx |
PgxEnrichmentError |
— |
strchive |
StrchiveError |
StrchiveUnavailable |
Clients — the transport layer, which a consumer calling a pass never meets
| module | base type | narrowed by |
|---|---|---|
atlas_client |
AtlasError |
AtlasNotScored, AtlasRefused, AtlasRefMismatch (under the one before it), AtlasUnavailable |
civic_api |
CivicApiError |
CivicApiUnavailable |
clingen_allele |
ClingenAlleleError |
— |
cpic |
CpicError |
— |
ensembl |
EnsemblError |
— |
eutils |
EutilsError |
EutilsRateLimitedError |
gene_spans |
GeneSpanError |
— |
gnomad |
GnomadError |
RateLimitedError |
pgs |
PgsCatalogError |
PgsCatalogUnavailable |
pharmvar |
PharmVarError |
— |
Builders and the publisher ([dev]) — what cache rebuild and a build command raise
| module | base type | narrowed by |
|---|---|---|
alphagenome_avi_build |
AlphaGenomeBuildError |
— |
atlas_protos |
ProtoFetchError |
— |
civic_build |
CivicBuildError |
CivicUnavailable |
civic_vcf |
CivicVcfError |
— |
clinpgx_build |
ClinPgxArchiveError |
ClinPgxUnavailable |
clinvar_build |
ClinVarBuildError |
ClinVarUnavailable |
constraint_build |
ConstraintBuildError |
ConstraintUnavailable |
cpic_build |
CpicBuildError |
— |
mane_build |
ManeBuildError |
ManeUnavailable |
pubmind_build |
PubMindBuildError |
PubMindUnavailable |
upload |
OrphanedSidecarError |
— |
upload |
PublishCollisionError |
— |
Snapshot readers — an absent or unreadable cache, deliberately a FileNotFoundError
| module | base type | narrowed by |
|---|---|---|
clinvar |
ClinVarReferenceError |
— |
download |
ConstraintReferenceError |
— |
download |
GatedSnapshotError |
— |
download |
OpenSnapshotError |
— |
download |
SnapshotNotPublished |
— |
pubmind |
PubMindReferenceError |
— |
resolver |
EnsemblReferenceError |
— |
Drafting providers, and the licence gate
| module | base type | narrowed by |
|---|---|---|
civic_draft |
CivicDraftError |
— |
clinvar_draft |
ClinVarDraftError |
— |
licensing |
LicenseRefusal |
— |
mitomap_draft |
MitomapDraftError |
— |
pubmind_draft |
PubMindDraftError |
— |
strchive_draft |
StrchiveDraftError |
— |
CLI¶
just-dna-enricher enrich spec/ --strict # write spec/resolution.csv, fail if unresolved
just-dna-enricher enrich spec/ --offline # cache-only (Ensembl + ClinVar), zero egress
just-dna-enricher enrich spec/ --no-clinvar # Ensembl links only
just-dna-enricher enrich spec/ --no-verify-ref # skip the reference-allele check
just-dna-enricher enrich spec/ --no-verify-clinsig # skip the ClinVar clin_sig cross-check
just-dna-enricher enrich spec/ --no-verify-rsids # skip the dbSNP merge/withdrawal check
just-dna-enricher enrich spec/ --no-verify-datasets # skip the recorded-release currency check
just-dna-enricher enrich spec/ --pubmind-cache pm/ # add PubMind as a second authority to the
# clin_sig check; without it that leg is unchecked
just-dna-enricher enrich spec/ --keep-par-twin # record both contigs of a pseudoautosomal locus
just-dna-enricher enrich spec/ --rederive # re-ask every recorded subject; report what moved
just-dna-enricher enrich spec/ --keep-staging # keep the staged answers after a successful commit
just-dna-enricher literature spec/ # pass 4: write spec/literature.csv (online only);
# reads studies.csv AND any binning row's pmid
just-dna-enricher literature spec/ --no-fulltext # existence + identifiers, skip the quote match
just-dna-enricher check-identifiers spec/ # trait CURIEs (OLS4) + gene symbols (HGNC) + gene/chromosome agreement
just-dna-enricher litvar coverage spec/ # RM167: which papers LitVar holds per locus, and at which tier
just-dna-enricher litvar coverage spec/ --offline # every locus recorded as unchecked, nothing asked
just-dna-enricher litvar gene HFE # every node under a gene, split by tier (rsID / CAID / gene / mention)
just-dna-enricher dosage spec/ --offline # no-op with a warning (ClinGen has no snapshot)
just-dna-enricher gene-validity spec/ # RM24: ClinGen expert-panel gene-disease assertions
just-dna-enricher gene-validity spec/ --source gencc # …or GenCC's aggregate of nineteen submitters
just-dna-enricher assertions spec/ # RM25: ClinVar's call + review tier per allele
just-dna-enricher assertions spec/ --offline # snapshot only; no snapshot → no-op with a warning
# Caches — provision once, then every gated pass runs with zero egress. See "The caches".
just-dna-enricher cache status # what is present, where, which release
just-dna-enricher cache pull # the ungated published lanes, from HuggingFace
just-dna-enricher cache pull --use non-commercial # …and ClinPGx, CPIC and the drug labels, which forbid sale
just-dna-enricher cache pull --only clinvar
just-dna-enricher cpic build --use non-commercial # whole CPIC → data/repro/cpic ([dev])
just-dna-enricher cpic publish cpic/ --repo org/cpic # redistribution is granted
just-dna-enricher clinpgx publish cp/ --repo org/clinpgx # LICENSE.txt travels with it
just-dna-enricher pharmvar build --use non-commercial # YOUR key; never published
just-dna-enricher pubmind build --download # literature verdicts; never published
just-dna-enricher pubmind publish # exits 1 and says why (by design)
just-dna-enricher mane build --download # the transcript numbering frame
just-dna-enricher pgx spec/ --offline --use non-commercial # both legs off snapshots
just-dna-enricher pgx spec/ --cpic-cache cpic/ --pharmvar-cache pv/
just-dna-enricher draft spec/ --gene CYP2C9 --offline --use non-commercial
just-dna-enricher acmg build assets/acmg_sf_v3.3.xlsx # once: the SF v3.3 snapshot
just-dna-enricher check-acmg spec/ --sf-list acmg/ # acmg_sf vs the ACMG SF gene list (offline-capable)
just-dna-enricher strchive build --release v2.26.0 # the repeat-locus catalogue
just-dna-enricher check-repeat-bands spec/ --catalogue strchive/ # repeat_alleles.csv bands vs STRchive (reports only)
# Authoring — templating and drafting (the compiler owns the offline half; see COMPILER.md)
just-dna-enricher template repeat_alleles.csv # header + required/one-of/never-empty defaults
just-dna-enricher draft spec/ --gene CYP2C19 # CPIC → haplotypes/allele_function/diplotypes
just-dna-enricher draft-repeats spec/ --catalogue strchive/ # STRchive → repeat_alleles.csv identity rows
just-dna-enricher draft spec/ --gene CYP2D6 --allele '*1' --allele '*4' --allele '*10' # 21 not 16,290
just-dna-enricher draft spec/ --gene CYP2C19 --drug clopidogrel # + every clinical context, as rows
just-dna-enricher draft spec/ --gene CYP2C19 --drug clopidogrel --population NVI # one context only
just-dna-enricher draft-clinpgx spec/ --snapshot cp/ --drug simvastatin --use non-commercial
just-dna-enricher draft-panel spec/ --gene MTHFR --gene BRCA1 # ClinVar gene panel (snapshot auto)
just-dna-enricher draft-panel spec/ --gene MTHFR --snapshot cv/ --offline # a snapshot you built
just-dna-enricher draft-panel spec/ --gene MTHFR --no-download # use a cached snapshot; fetch none
just-dna-enricher draft-panel spec/ --gene MTHFR --dry-run # the genotype worklist, appending nothing
just-dna-enricher draft-panel spec/ --gene BRCA1 --source pubmind --pubmind-cache pm/ # literature verdicts
just-dna-enricher draft-panel spec/ --gene BRCA1 --source pubmind --min-confidence 2 # a deeper floor
just-dna-enricher mitomap build # MITOMAP's pg_dump → both curated mtDNA tables (data/repro/mitomap)
just-dna-enricher mitomap miss # the join against ClinVar chrMT; both parents required
just-dna-enricher draft-panel spec/ --source mitomap-miss # the increment only; --gene filters, never required
just-dna-enricher clinvar citations --out data/repro/clinvar --download # add PMIDs so a panel can compile (--out names the EXISTING snapshot)
just-dna-enricher clinvar publish data/repro/clinvar # data/ + citations/ + release.json
# Authoring — lookups. These WRITE NOTHING: every answer comes back advisory, with a reason.
just-dna-enricher hint variant --rsid rs1801133 # validity, loci, ref/alts
just-dna-enricher hint variant --rsid rs334 --ambiguity # warn when the answer is not unique
just-dna-enricher hint variant --rsid rs1801133 --frequencies # + gnomAD populations (paced ~6s)
just-dna-enricher hint variant --rsid rs1801133 --offline --json
just-dna-enricher hint variant --chrom 17 --start 43093220 --ref C --alts T --pubmind-cache pm/
just-dna-enricher hint citation --pmid 9545397 # which paper it is + the DOI/PMC id it carries (--json)
just-dna-enricher hint citation --pmcid PMC3110566 # the PubMed id for a PMC id — reported, never written
just-dna-enricher hint trait EFO_0004340 # current | obsolete | absent
just-dna-enricher hint gene MTHFR # approved | retired | unknown
just-dna-enricher enrich-and-compile spec/ out/ # enrich, then compile from resolution.csv (offline)
just-dna-enricher upload out/coronary --dry-run # plan a module HF upload ([dev]); names both paths
just-dna-enricher upload out/coronary # push to data/coronary/ and data/coronary/v<version>/
just-dna-enricher clinvar build --vcf clinvar.vcf.gz # VCF → snapshot parquet ([dev]), data/repro/clinvar
just-dna-enricher clinvar build --download # fetch the NCBI VCF first, then build
just-dna-enricher clinvar publish data/repro/clinvar --dry-run # plan the reference-snapshot upload
just-dna-enricher clinvar publish data/repro/clinvar # create-or-update datasets/just-dna-seq/clinvar
enrich takes the mode and cache flags (--strict/--best-effort, --offline, --ensembl-cache,
--clinvar-cache), the per-link toggles (--clinvar/--no-clinvar, --gnomad/--no-gnomad,
--vrs/--no-vrs), the four verify toggles above (--verify-ref, --verify-clinsig, --verify-rsids,
--verify-datasets), --pubmind-cache, --keep-par-twin, and the two transaction flags
--rederive / --keep-staging. enrich-and-compile takes the mode and cache flags, the ClinVar and
gnomAD link toggles, and the --frequencies / --gene-metrics pass toggles — not --vrs, not the
verify toggles, not the transaction flags — it is the offline convenience wrapper, so those are
enrich's alone. That distinction is the rule worth writing down; the enumeration is not, because
a flag list restated in prose rots the moment a flag is added, and --help is the authority.
literature takes
--strict/--best-effort, --offline, --fulltext/--no-fulltext; check-identifiers takes
--strict, --traits/--no-traits, --genes/--no-genes, --pgs/--no-pgs (RM163) and, like check-acmg, writes no authored
cell and records that the question was put: there is no sidecar column for a module-level identifier
and filling one from the registry being asked about it would make the comparison vacuous
(hints.REDUNDANCY_BEARING), so the report is the whole output apart from the verification.json
attestation — five records for check-identifiers (the two PGS ones since RM163), one for check-acmg, unconditional, and an
attestation rather than a value; upload takes --repo, --name, --message,
--dry-run; clinvar build takes --vcf/--download/--out; clinvar publish takes --repo,
--message, --dry-run. enrich-and-compile runs enrich then compile_module(..., ensembl_cache=None,
strict=…), so compilation consumes the just-written resolution.csv (path 1) with no reference and
no network.
Command → Python API¶
Every command is a thin shell over a library function; nothing here is CLI-only logic. Use the API directly to compose passes, inject clients, or run in-process.
| Command | Python API |
|---|---|
enrich |
enrich.enrich |
frequencies |
frequencies.enrich_frequencies |
gene-metrics |
gene_metrics.enrich_gene_metrics |
dosage |
clingen.enrich_dosage_sensitivity |
litvar coverage |
litvar.check_literature_coverage + litvar.verification_records |
litvar gene |
litvar.LitvarClient.gene_nodes |
literature |
literature.enrich_literature |
pgx |
pgx.enrich_pgx |
check-identifiers |
identifiers.check_identifiers |
check-acmg |
acmg.verify_acmg_sf (+ AcmgReport.by_gene for the grouped view) |
check-repeat-bands |
strchive.check_repeat_bands |
cache status / cache pull / prepare / rebuild |
locations.resolve_*_reference / download.ensure_*_snapshot / caches.prepare_caches / caches.rebuild_lane (+ publish_reference_snapshot under --publish) |
cpic build / publish |
cpic_build.build_snapshot / upload.publish_reference_snapshot |
pharmvar build |
pharmvar_build.build_snapshot (no publish — see The caches) |
pubmind build / publish |
pubmind_build.download_pubmind_table + build_snapshot / refuses, cli.PUBMIND_PUBLISH_REFUSAL |
mane build |
mane_build.discover_current_release + download_mane_file + build_snapshot (no publish — see The caches) |
acmg build |
acmg_build.build_acmg_snapshot → acmg.load_acmg_snapshot |
draft |
pgx_draft.draft_gene |
draft-clinpgx |
clinpgx_draft.draft_pharm_variants |
draft-panel |
clinvar_draft.draft_gene_panel / pubmind_draft.draft_gene_panel_from_pubmind / civic_draft.draft_panel_from_civic / mitomap_draft.draft_panel_from_mitomap_miss, by --source |
mitomap build / miss / publish |
mitomap_build.download_mitomap_dump + build_snapshot / mitomap_miss_build.build_miss_snapshot (+ stale_parents) / upload.publish_reference_snapshot |
draft-repeats |
strchive_draft.draft_repeat_loci |
strchive build |
strchive_build.build_strchive_snapshot → strchive.load_strchive_catalogue |
clinpgx build / clinpgx check |
clinpgx_build.download_clinpgx_zip + build_snapshot / clinpgx.enrich_clinpgx |
clinpgx build-labels / check-labels |
drug_labels_build.download_drug_labels_zip + build_drug_label_snapshot / drug_labels.check_drug_labels |
civic build / citations / publish / reproduce |
civic_build.build_snapshot / civic_citations.draft_civic_citations (+ check_evidence_status_currency, folded into enrich) / upload.publish_reference_snapshot (CC0; RM176) / the CLI's own five checks |
clinvar build / citations / publish |
clinvar_build.download_clinvar_vcf + build_snapshot / download_var_citations + build_citations / upload.publish_reference_snapshot |
gnomad constraint build / publish |
constraint_build.download_constraint_tsv + build_snapshot / upload.publish_reference_snapshot |
vrs mint |
vrs.mint_resolution_rows |
hint variant / citation / trait / gene |
lookup.lookup_variant / lookup_citation / lookup_trait / lookup_gene |
upload |
upload.plan_upload / upload_module |
template |
just_dna_compiler.draft.blank_template (cross-package convenience) |
enrich-and-compile |
enrich.enrich → just_dna_compiler.compiler.compile_module |
Two library functions have no command, on purpose. clinical.verify_clin_sig and
sequences.verify_reference_alleles are checks that run inside enrich (--verify-clinsig,
--verify-ref), because their verdicts land on resolution.csv and on the enrichment report. Running
either standalone would compute a finding with nowhere to put it.
And two take rows rather than a spec_dir — both now take either (RM41, 0.5.1).
acmg.verify_acmg_sf and identifiers.check_identifiers are the only passes whose input is a list of
VariantRow, which left every caller turning variants.csv into models itself — and the only thing that
does that correctly is just_dna_compiler.load_csv_rows, which was private until 0.5.1. Re-implementing
it is a trap rather than a chore: an empty cell becomes None with the key kept (so a defaulted-but-
not-Optional field like MeasureBinRow.measure_kind fails on type rather than taking its default),
and genome_build is told to each row, so a loader that omits it mints GRCh38 identities for a
GRCh37 module. compiler.load_spec_variants(spec_dir) does the load, the injection and the
_restamp_for_build in one call, and both checks accept spec_dir= alongside the existing variants=.
Exactly one, never both — two answers in mind, and silently preferring one is a guess.
Since 0.5.4 the two forms are no longer equivalent for check_identifiers, and the report says which
you got: the gene ↔ locus comparison needs a chromosome per row, and an rsID-only row has one only
after resolution, so spec_dir= lets it read the injected resolution.csv beside the spec while
variants= alone limits it to authored coordinates. Where that leaves nothing to compare,
gene_loci_not_checked names the reason rather than reporting a clean zero.
Since 0.7 the roster is every authored table carrying the column, and only spec_dir= can reach
them (S86). check_identifiers built its roster from variants.csv alone while eleven authored
models declare trait_efo_id or gene — StudyRow has carried the trait column since 0.3 — so a
module whose traits live in studies.csv reported nothing checked and nothing flagged, and could ship
a retired CURIE with every gate green. The set is derived from DRAFTABLE rather than listed
(@registry-completeness): nine tables carry trait_efo_id, nine carry gene, and a table kind added
later joins by existing. The three derived models that also carry these columns —
GeneMetricsRow, GeneValidityRow, GwasEffectRow — are deliberately outside it: they are
machine-written, so a stale id in one is the source's currency and no author can act on it.
The unreadable 0 was the item, not the omission. traits checked: 0 said this module declares
no trait and its traits are in a table nobody read in the same breath, and a reader took the second
for the first. IdentifierReport now carries trait_tables_read/trait_tables_not_read and the gene
pair beside them, the CLI count names the tables it is out of, and a table that exists and will not
parse is reported as unread — an absent optional table is every module's normal shape and says
nothing, while an unreadable one means ids the module carries went unchecked. report.clean no longer
prints "all identifiers current" over an empty roster, which was the same vacuous pass one level up.
A caller passing variants= still gets the narrow roster and is told so in *_tables_not_read, since
rows in hand are all that form has.
Since 2026-09-13 an unreadable id-bearing table also fails the gate, and clean says so. The
report was honest and the verdict was not: the five stale lists are empty when a table carrying
identifiers will not parse for exactly the reason they are empty when everything agreed, so
--strict printed the unreadable table and then exited 0 beneath all identifiers current. A
terminal reader saw both halves; a library caller reading the dataclass — which is how this arrived,
as an MCP tool — saw only the verdict. clean is now a Verdict: falsy when it carries any code,
pass when empty, and if report.clean: is unchanged across the change. AcmgReport.clean is the
same type, so the two gates in this tier answer in one kind.
The set holds errors, not non-answers, which is the line that decides membership. offline and
tables_unreadable are errors and fail. A check the caller switched off, and a module with no row a
check applies to, are not: --strict --no-traits still exits 0, and check-acmg over a module
stating no acmg_sf cell now passes with its denominator printed rather than withholding, which
is the one behaviour RM234 set differently. unreachable is not a member and could not be — a
registry outage raises IdentifierUnavailable before a report exists, and that path already exits 1
with an unreachable attestation for all five checks.
A prose filter was the defect underneath it. unreadable_tables kept everything in
*_tables_not_read except the literal "not present", so the row-taking call form's reason — a
statement about how the check was invoked, not about a file — counted as eight tables that would not
parse. Nothing caught it because the CLI always passes spec_dir. Both benign reasons are now one
constant each (NOT_PRESENT, ROWS_PASSED_IN), written by the producer and read by the filter. That
is still prose matched at two sites; the structural repair is to carry the roster's read_errors
onto the report, which is a field and therefore not a patch.
And since RM156 a module carrying no variants.csv can reach that roster at all. The widening
left two gates in front of itself, both keyed on the one table it had stopped depending on:
check_identifiers(spec_dir=) loaded variants.csv unconditionally and raised
variants.csv is invalid: ... not found, and the command returned "no variants.csv — nothing to
check" before ever calling it. The table has never been mandatory, and four of the nine that carry
gene are the PGx kinds a module is built entirely out of — this repo ships four such reference
examples, and cyp2c19_star_alleles names CYP2C19 on every row it has. The rows are wanted for one
thing, placing a symbol against the chromosome its variant sits on, so an absent variants.csv is now
no rows and gene_loci_not_checked says why; a variants.csv that exists and will not parse still
raises, because that is a module whose rows exist and cannot be read. The command's guard is the
roster instead of a filename: nothing to check means no id-bearing table was read, and only then
is no attestation written, since with no question put there is nothing to record having asked.
template is the one command that duplicates the compiler's, and the other four offline authoring
commands (stub, requirements, describe, hint, scaffold) are deliberately not mirrored. The
offline authoring surface has an owner — just-dna-compiler — and the single mirror exists so a PGx
author working through this binary does not have to switch tools for a CSV header. See
COMPILER.md § CLI for the compiler's own table, including the three schema-tier
functions (keygen, sign, reference) that surface there because just-dna-format ships no CLI.
A quote that is the article's own title (S54)¶
quotes_found checks a provenance_quote against the article's fulltext, and a title appears in
its own fulltext — so a quote copied from the article's metadata makes the check pass while
establishing nothing about whether any claim is in any paper. Four published modules are in exactly
that state: 3,668 rows, one distinct quote per PMID, each verbatim the title. A passage located
for a specific claim varies row to row, because different rows cite the same paper for different
findings; one string per paper repeated across every citing row is structurally a property of the
article, not of the claim.
The pass reports it. LiteratureResult.titles_as_quotes lists the PMIDs whose every
provenance_quote is the title, and the CLI prints it in yellow — a warning, never an exit code,
because whether a title is an acceptable locator for a claim is the author's decision and what the
tool can honestly say is that quotes_found is not evidence here.
The discriminator is the metadata, not the shape of the string. A minimum length does not
separate a title from a passage — seventeen words is an ordinary title and an ordinary sentence — and
requiring a provenance_regex instead is no help, since a regex is as copyable as a quote. The
comparison is against the title already held for that article (it arrives in the same esummary
response that answers existence), so it costs no request and it answers for a paywalled article
too, which is where quotes_found stays null and a reader has nothing else to go on.
Two deliberate narrownesses. Normalisation is case, whitespace and a trailing period and nothing more — a quote that merely contains the title is a real quote of a paper that names itself, and an aggressive normalizer would start rejecting located passages. And the report fires only when every quote for a citation is the title: a module quoting the title on one row and a passage on another has an author doing the work.
It answers for a pinned row too, and that is the whole difference between working and not.
literature.csv is merge-not-clobber, so a row an earlier pass wrote is not refetched while it still
counts the module's quotes (since RM277 a pin at another count is re-fetched and re-checked online;
offline keeps it) — and on the four modules that motivated this, every row is pinned, so a check living only in the fetch loop would
have fired on none of the 3,668 quotes. The summary is therefore fetched for any cited PMID carrying
a quote, pinned or not; esummary batches, so it costs no extra round trip in the common case. The
pinned row itself stays authoritative and untouched — the merge rule is not the thing that was wrong.
One correction to the report, found against the recorded fulltext for PMC5753237. A title appears
in its own fulltext nearly always rather than always: esummary gives it with a trailing period, the
JATS body carries it without, and quote_matches does not strip one — so that exact pair misses. The
substance is unaffected, since the miss is punctuation rather than evidence, and a module whose two
spellings agree gets the green check the report describes. Both states are pinned.
What makes two rows of a sidecar the same row (S51)¶
Every derived sidecar is merge-not-clobber, so each pass holds an existing map and must decide when
an incoming row is one it already has. That decision is published: each row model declares
_KEY_FIELDS, just_dna_format.base.merge_key(row) is what every pass keys its map on, and
just_dna_compiler.hints.key_fields(csv_name) reports the same columns to a consumer holding only a
filename. The pass and the surface are one statement rather than two, which is the half that keeps them
from drifting — before 0.6.5 the key existed only as a dict-key expression inside each pass's body, so
a tool re-deriving a sidecar had to guess it, and the guess was coarse on exactly the two tables where
one subject legitimately carries several rows.
| file | key | rule | note |
|---|---|---|---|
resolution.csv |
variant_key |
subject | one rsID resolves to several loci (locus_index orders them); the pass replaces the group whole, so several rows sharing this key is the normal shape and not a duplicate |
frequencies.csv |
(variant_key, population) |
equality | one variant carries a row per ancestry group |
gene_metrics.csv |
(gene, dataset) |
equality | two writers (gnomAD constraint, ClinGen dosage); one gene legitimately carries a row per authority |
gene_validity.csv |
assertion_id, falling back to (gene, disease_id, moi, submitter, dataset) |
equality | the source's own id where it published one. Never (gene, disease) alone, which collapses 59 real ClinGen curations, and never without submitter, which collapses the disagreement GenCC exists to publish |
clinical_assertions.csv |
(variant_key, variation_id) |
equality | a null variation_id is a value, carrying the not_found row that states the archive was consulted and holds no record |
gwas_effects.csv |
association_id |
equality | per record, not per variant — a variant gains associations as papers publish |
literature.csv |
pmid |
equality | one article, one row |
sources.csv / licensing.csv |
(source, layer) |
equality | one source legitimately appears at two layers |
What a re-run does since 0.7, and what it no longer costs (RM124)¶
Merge-not-clobber's behaviour is unchanged and deliberately so: a re-run gap-fills — it fills subjects with no row and leaves recorded rows alone. Re-asking every subject on every pass was rejected, because it would put the full resolution time on every run to buy drift detection nobody asked to run continuously.
What changed is what leaving a recorded row alone risks, and the answer is now nothing. A
curator's correction to a covered derived table lives in overrides.csv
beside the spec, not inside the sidecar; the compiler applies it on every build. So every derived file
this tier writes is a pure build product, hand-editing one is not expected, and rm plus a re-run —
the crude form of a full re-derivation — costs nothing. Each writer's docstring says so at the point a
reader outside this repo actually meets it.
Two consequences for a pass author. A row this tier writes carries no authored content, so a pass may reason about it as the source's answer and nothing else. And a difference between a fresh row and a recorded one now means the source revised, full stop — the ambiguity that made "re-derive the machine rows and keep the overrides" unimplementable is gone, because the tier never has to tell a curator's edit from what the source said last time.
Read rule as well as columns. resolution.csv is the only subject table today, and treating
its key as a uniqueness constraint would report a legal one-to-many file as a duplicate. And read
fallback — gene_validity.csv is the only table with one, and a consumer ignoring it is right
about the other seven and wrong about the one where it matters. The two levels are tagged ("id" /
"grain"), so a grain tuple can never collide with an id that happens to equal it.
A check's findings, and where each one's survive (S70)¶
verification.json records counts, and until S70 clinical_significance was the only one of
the checks enrich() puts whose findings then survived nowhere at all:
| check | where its findings survive |
|---|---|
reference_allele |
detail=, via summarize_ref_mismatches — grouped by diagnosis, three keys named per group plus "and N more" |
rsid_coordinate_agreement |
detail= — up to DETAIL_LIMIT disagreements, plus what was not compared and why |
rsid_currency |
per row in resolution.csv — rsid_status and rsid_current are columns of the written file |
genome_build_agreement |
count only, but its subjects are reference_allele's mismatches, so the candidate rows are reachable |
clinical_significance |
detail= since S70, grouped on opposed. Before that: the logger, and nothing else |
The conflicts always existed at runtime — compare_clin_sig returns them and every one is logged —
but stderr survives the process and nothing else does, so a record reading findings: 20,
detail: null left an author able neither to defend the twenty nor correct them, and re-running to see
the log again costs the full ClinVar comparison. _clin_sig_detail groups on ClinSigConflict.opposed
because that is the distinction that decides what to do (an opposed pair is worth acting on, a
difference is worth knowing) and names both values, since the standing instruction on a mismatch is
to check both sides and a bare key cannot be checked at all.
The compiler now says so too. _findings_warning reports a non-zero findings at validate and
compile; nothing read VerificationRecord.findings before, so the counts reached
manifest.verification.checks[] and an author running the validator saw a green result. A record
reporting zero says nothing — a check that could not fail must not report one.
The other half shipped in 0.7 as RM130:
a detail makes the rows nameable to a human and not joinable, and the concordance record above is
what makes them joinable. RM124's question about whether one record serves both an overlay and
outranks was answered as a dated succession, so the record is the overlay's input side — a
conflict is a question and an overrides.csv row is the answer.
How a sidecar is written, and why merge-not-clobber makes that load-bearing (S66)¶
Every sidecar writer goes through just_dna_format.layout.atomic_writer / atomic_write_text —
a temp file in the same directory, fsync, then os.replace. Same directory because os.replace is
only atomic within a filesystem; fsync before the rename so a machine losing power cannot expose an
empty file where a complete one used to be.
This sits under the merge table above rather than beside the writers because the merge is what makes
truncation dangerous. A writer that truncates in place leaves, on a kill, a file that is
syntactically valid and simply short — and a short resolution.csv is not detected as damaged by
anything, it is read back, keyed on subject, and believed. Three branches in enrich() write no
row for a subject nothing could answer (@unreachable-not-absent, and each is correct), so fewer
rows is a state the table reaches legitimately: a truncated table and a module whose author resolved
less are the same bytes. The reported incident is the sharp form — a run killed client-side kept going,
reached the write, and replaced a restored 330-row table with 162 rows, after which the module
validated, closed and compiled green.
enricher/tests/test_atomic_sidecar_writes.py walks the writers rather than testing one: three of the
nine were reported and the other six had the identical shape, so a fix scoped to the report would have
left the next writer inheriting whichever neighbour it was copied from.
What this did not fix, and what RM128 then did. Atomic writing left enrich()
persisting nothing until its tail, so a run killed before the write had written nothing at all —
atomically. The lost-work half, the missing lock over a read-modify-write window that spans the entire
run, and the absent progress callback were each a decision rather than a missing line, and all three
shipped in 0.7: see The run is a transaction above.
resolution.csv is provisional (0.5)¶
The table's shape (ResolutionRow — columns, keying, the status vocabulary, how one-to-many expansion
is encoded) arrived in 0.5 and is still not frozen at 0.7: SCHEMAS.md carries it as provisional,
and the additive-within-a-major / digest obligations have deliberately not been engaged for it. It is
free to be refactored wholesale while that status stands, and has taken several passes already. See
SCHEMAS.md § the resolution table. The compiler's
consumption contract (digest parity between the resolution.csv path and the DuckDB path, offline
round-trip) holds regardless of the table's internal shape.
Downstream adoption (cross-repo)¶
The enricher is the single source of truth for variant resolution. The intended migrations, tracked in ROADMAP.md, live in their own repos:
- ensembl-mcp keeps only its FastMCP wrapper and imports the query core from here.
- just-dna-lite / just-dna-pipelines drops its resolver shim + HF download/upload copies and imports
the enricher (
just_dna_enricher.uploadis the canonical publisher API; pipelines still carries a local copy while it is pinned to format/compiler<0.4and cannot import the 0.5 enricher yet).
Open questions the code does not answer¶
Written down because each was reached for and not found, and a reader who reaches for the same thing
deserves to know it is absent rather than hidden. None is a defect with a filed RMn — a defect is a
place where the code and a stated rule disagree, and these are places where nothing states a rule.
- The Ensembl snapshot's builder and its
release.json.download.ensure_snapshotprovisionsdatasets/just-dna-seq/ensembl_variations/data/homo_sapiens-*.parquetandresolver.pyreads a fixed column set from it (id,chrom,start,ref,alt), but nothing in this package builds that snapshot, unlike the ClinVar and gnomAD ones.enrich._snapshot_releasereads adatasetkey no builder here writes, sorsid_coordinate_agreement.releaseisNonein practice unless something outside this repo produces one. Worth knowing before treating that field as provenance. - What the ClinVar snapshot's
cache statuslabel should be.clinvar_build._write_release_jsonwrites nodatasetkey andcache statusprintsrelease.get("dataset") or "", so a correctly built ClinVar snapshot shows a blank label. Whether that is intended — ClinVar has its own richerclinvar_dataset_label— or an oversight is not decidable from the code. enrich_gwas(mode=…). Accepted, defaulted, never read; the CLI's--stricthelp promises a "severity ladder" and the pass docstring describes none. Whether the intent was astrictrefusal onmissing(as the sibling passes have) or no ladder at all is undetermined, and no test covers either. The dead parameter itself is filed as part of RM100; what it should do is this question.identifiers.IdentifierCheckErroris declared with a docstring and never raised anywhere, insrcor in the tests.CpicClient.knows_drug's return type. Annotatedbool | None; the live client can only returnboolor raise, and onlyCpicSnapshotClientreturnsNone.pgx_drafthandles all three, so the widening looks deliberate — the live client's docstring just does not say so.- Whether
check-identifiersis meant to have an--offline.identifiers.verification_recordsargues it must not; its siblingcheck-acmghas one. Self-consistent, undecided as policy.
Testing¶
uv run pytest enricher/tests. All network-free: the cache side uses a tiny synthetic
ensembl_variations parquet; the live-Ensembl side uses an httpx.MockTransport; upload is tested
against a fake HfApi. Coverage includes offline enrich → compile matching the DuckDB digest, --offline
making zero network calls, the V2 503 → V1 REST fallback, tenacity retrying a transient error, strict
failure, one-to-many expansion, and the upload plan/token paths. Integration tests that need a real cache
are @integration (skipped without JUST_DNA_ENSEMBL_CACHE).
Two files carry an unusual shape on purpose. test_query_shapes.py guards the hash-join probe
(0.5.2): its load-bearing assertion is on the query plan, so it needs no clock at all, and its one
timing check runs the OR-chained shape and the joined shape in the same process against the same
synthetic snapshot — a relative bound, because an absolute one just measures the runner. And
test_locations.py runs every probe in a subprocess with a controlled cwd and environment: the
bug it pins is the first resolve in a process, and load_env mutates os.environ for the rest of the
session, so an in-process test would neither reproduce the defect nor stay isolated from the rest of the
suite. It also demonstrates the old arrangement failing rather than asserting that it used to.
The gnomAD tests are driven by real recorded payloads committed under assets/
(gnomad_v4.1_variant_payload.json, gnomad_gene_constraint_payload.json,
gnomad_v4.1_constraint_slice.tsv), replayed through httpx.MockTransport. Recording rather than
fabricating matters here: the quirks under test — the "Multiple variants found" error sitting beside
valid data, the duplicated XX/XY entries, the two mane_select=true rows per gene — are ones a
hand-written fixture would have quietly omitted, and the tests would then have passed against the naive
implementations they exist to catch. Coverage: batching and the pacing gate (on an injected clock, so
the suite never really sleeps six seconds), partial-error resilience, AC/AN against the payload's own
exome.af, faf95 landing on one group only, the MANE pick fed in both orders, byte-identical
snapshot rebuild, the chain order (a both-links variant keeps source="cache"), --offline making zero
gnomAD calls, an unqueryable ClinVar cache degrading instead of crashing, VRS minting and the
snapshot-vs-API dataset labelling. Live queries and indel normalization are @integration and opt-in
via JUST_DNA_NETWORK_TESTS=1.
The ClinVar tests (test_clinvar.py) build against a small committed real slice
(assets/clinvar_GRCh38_slice.vcf.gz), computing expected values from that VCF at runtime: build
correctness (columns, clin_sig ⊆ VALID_CLIN_SIG, clin_sig_raw preserved), the cross-source
coordinate agreement / off-by-one guard, byte-identical rebuild, the ClinVar-after-Ensembl chain order
(no compiled digest moves), one-to-many expansion, the allele-aware back-fill + ambiguity marking, the
compile → reverse → compile fixpoint, and the mocked reference-snapshot publish. A full build against
the real local VCF is @integration (skipped when absent).
The clinical cross-check (test_clinical.py) builds a snapshot from that same slice and leans on a
coincidence in it that is too useful to be luck: the slice contains both of rs334's opposed records
(T>A pathogenic, T>G likely_benign), so allele-exactness is provable on real data rather than on a
constructed pair. Coverage: an opposed call reported with its star count, the same call on the other
allele correctly staying silent, uncertain_significance not counting as disagreement, the locus-wide
fallback in both directions, strict not escalating, and a foreign parquet degrading instead of raising.
The literature and identifier tests replay recorded payloads from the live services
(pubmed_esummary_payload.json, europepmc_search_payload.json,
europepmc_fulltext_PMC5753237.xml, dbsnp_esummary_payload.json, ols4_terms_payload.json,
hgnc_fetch_payload.json). Each recording carries a quirk a hand-written fixture would have smoothed
away, and the tests exist to pin exactly those: a nonexistent PMID arriving as a normal-looking record
with an error key; a real PubMed record that is simply not in PMC (exists-yes, retrievable-no);
Europe PMC silently omitting ids it does not know; and — the one most worth guarding — rs11273140
(withdrawn) and rs2000000000 (never assigned) returning byte-identical responses, asserted on the
recordings themselves so that a future dbSNP release which does separate them fails the test rather
than silently invalidating the design. The quote matcher is exercised against the real JATS fulltext
with a phrase read out of that same document, and the regex bound is demonstrated on a genuinely
catastrophic pattern (it must return not checked, never not found).
Each of these files also carries an opt-in live probe (JUST_DNA_NETWORK_TESTS=1) that re-asks the real
services the same questions, so a recording that has drifted away from reality fails loudly instead of
letting the unit tests pass against a fiction.
GWAS effect sizes (gwas.py, online only) — RM90¶
just-dna-enricher gwas <spec-dir> fills gwas_effects.csv with the NHGRI-EBI GWAS Catalog's
published effect sizes for every rsID the module names. One row per published association, not
per variant — rs1800562 alone carries 186 — keyed on the Catalog's own association_id and merged
never clobbered.
The subjects come from enrich.collect_subjects since RM158, not from variants.csv alone. Five
authored models carry rsid, so a module whose rsIDs live in haplotypes.csv or pharm_variants.csv
got no associations and no line saying none had been asked for. The collector is the same one
resolution uses and answers the same question — which rows in this spec ask about a variant — so this
pass inherits its precedence (variants.csv first, first occurrence winning, so a PGx row cannot take
an identity a SNP row minted) rather than restating the loop. studies.csv carries rsid and is
deliberately not a subject: a study row references the variant it grounds, which the module already
carries as a row of its own.
It does not fill weight, and that is the point of the pass rather than a limitation of it. A
consumer asked for exactly that (S36); MODULE_LIFECYCLE § Stage 3 names the cell, and every check
here reports rather than repairs. The effect sits beside the authored column and a consumer picks one
wholesale.
Everything below came from probing the live API, not its documentation.
- The association payload is thin.
pmid,study_accession,ancestry,traitandtrait_efo_idall sit behind_links.studyand_links.efoTraits, so a complete row costs two extra requests and the pass is1 + 2Nper variant._LinkCachememoizes by resolved URL — and the measurement refuted the prediction that motivated it: onhfe_hemochromatosisit saved nothing, because rs1800562's 189 associations each name their own study. 382 requests, zero cache hits.--no-study-factsexists because of that number and drops the cost to one request per variant, keeping the effects and losing the linked metadata. The loss is permanent for the rows it writes, and this is not the per-run trade the sentence above reads like (S50). The merge is keyed onassociation_idalone, and the skip happens before the row is built — so a row written with--no-study-factskeepspmid,study_accession,ancestry,traitandtrait_efo_idnull, and a later run with study facts skips it rather than filling it in. Deletegwas_effects.csvto re-derive them. Every other delete-to-regenerate case in this tier is about a stale value; this one is about a value that was never fetched, so the file looks complete — every column present, most cells populated — and cannot be repaired incrementally. Back-filling instead was considered and refused: it would make the pass rewrite existing rows, which is what merge-not-clobber exists to prevent, and a nullpmidis not distinguishable from a study record that has none — the casefollow's 404 arm deliberately produces. - A 404 is the empty answer, not an outage. The Catalog holds only variants with a published
association, so it 404s on a rare clinical one.
GwasNotFoundkeeps that typed:associations_forreports the empty answer (recorded as anot_foundrow),followwithholds one association's study facts and keeps the effect. The first version of this pass read a 404 as a transport failure and died on the first variant of the first real module it met. pvalue: 0.0is an underflow the Catalog really publishes.p_value_numisgt=0because such a value is not a probability; the pass withholds the number, keeps the verbatimp_valuestring, and keeps the row — an early version let theValidationErrordiscard a real published effect over one derived column (six on rs1800562; 189 rows became 195).riskAlleleNameisrs4149056-?when a study never established which allele carries the effect. Null, never a guess, and the row is kept and counted — 42 of 195 on that module.
Licensing. GWAS_CATALOG_TERMS is the first entry here with no named licence: EBI states its
terms in prose (read 2026-08-17). redistribution=True is directly supported; commercial_use stays
None because the page permits "use" generally but conditions it on the original data owners' terms,
which for an aggregator of thousands of publications are not established. Unknown is neither permission
nor refusal — taints_commercial_use requires an explicit False, so a null warns rather than gating.
Do not tidy it to True; a test pins it.
LitVar2 / PubTator3 — literature coverage, and the tier that answered (litvar.py, online) — RM167¶
NCBI's LitVar2 indexes the literature by variant, and it does so at three tiers that sit beside
each other as separate nodes. A node id is litvar@<clingen_id>#<rsid>#<gene_id> with an unfilled
slot collapsing to a bare #, so litvar@rs1800562## is the position node and
litvar@CA113795#rs1800562## is the allele node beside it, named by a ClinGen canonical allele
id. just-dna-enricher litvar coverage spec/ asks, per module locus, which of those answered.
The finding is the tier, because the tier a locus is answerable at is a property of the locus and not of the source. Measured 2026-09-01:
| locus | position node | allele node(s) | on the position node and no allele node |
|---|---|---|---|
| BRAF rs113488022 | 32,095 | 31,276 + 99 + 41, three CAIDs | 801 (2.5 %) |
| HFE rs1800562 | 3,053 | 2,693, one CAID | 360 (12 %) |
| APOE rs429358 | 3,945 | 328, one CAID | 3,617 (92 %) |
BRAF's three CAIDs are three distinct ALTs at one codon — V600E, V600G, V600A — and they differ by three orders of magnitude, so allele resolution is doing real work. APOE is the case that decides the shape: 92 % of the literature at that locus is not allele-resolved, so a pass that reported the allele node's count as the answer would understate it twelvefold, and one that quietly substituted the position count would answer an allele-level question with a position-level fact. Allele-resolved, position-only and absent are three outcomes, and for once the source supplies the three states rather than the schema imposing them.
How a module locus reaches an allele node. The position node's own record carries clingen_ids;
each CAID goes through RM153's clingen_allele.ClingenAlleleClient, and its GRCh38 allele is compared
against the ones the module names — a real second authority rather than a rename. A resolution.csv
that already carries a caid short-circuits that, because the module has then stated its allele
identity outright. A one-sided indel — the shape the registry states with an empty referenceAllele
or an empty allele — is anchored through clingen_allele.anchor_indel using the module's own
ref base at its own start, and withheld anywhere the module does not state one: a guessed
anchor puts a wrong ref on a right position, which is a false match rather than a missing one. That
last leg needed a repair one file over to work at all — _parse computed the one-sided allele and
then dropped it on every record that also carried an rs number, which is most of them, so the registry
looked as though it held nothing comparable. It is carried now, and the corpus's own
hboc_palb2 rows are what the tier verdict turns on.
The asked tier is a property of the module's rows. A module that names an allele at a locus asks an
allele-level question; one that names an rsID and nothing else asks a position-level one, and the
position node answers it exactly. Without that split a purely positional module would have every locus
counted as a shortfall. The allele columns come off AuthoredModel.ALLELE_COLUMNS, and the rsID roster
off DRAFTABLE, so a table kind added later joins by existing.
Four outcomes, and the fourth is the house algebra's. allele / position / absent /
unchecked, each with its own reason and its own sentence (coverage_reason, one arm per way of
getting there, walked by a test):
| tier | reason | what it means |
|---|---|---|
| allele | allele_node_matched |
an allele node for the allele this module names |
| position | row_names_no_allele |
the module asked at position level; this is the answer, not a shortfall |
| position | no_allele_node_at_locus |
the index holds no allele node here at all |
| position | allele_nodes_name_other_alleles |
it holds some, and none of them is this module's |
| absent | no_node_for_rsid |
the index holds nothing for this rsID at any tier — an answered absence |
| unchecked | offline |
nobody asked |
| unchecked | index_unreachable |
LitVar could not be reached |
| unchecked | registry_unreachable |
the CAIDs at this locus were never resolved, so none of them is this module's is not established |
| unchecked | allele_not_comparable |
the registry answered and holds no allele these columns can compare |
The last two are the collapse @answered-is-not-absent names, one tier out: an empty match is an
established negative only where every CAID was actually compared.
The residue is counted, never discarded (@dont-discard-computed). position_only_pmids is the
papers on the position node that no allele node at that locus claims — a set difference over the
union of every allele node, not a subtraction, because the allele nodes overlap each other. It is 92 %
of APOE and 2.5 % of BRAF, and both numbers are about the locus rather than about the module's allele.
It writes no row, and that is the whole artifact answer. A PMID list per variant is not a table
kind, literature.csv is keyed by article, and 32,095 PMIDs for one BRAF locus is a row-writer arguing
against itself. What lands is verification.json: one literature_coverage record whose subjects is
the loci the index answered about, whose findings is the loci where an allele-level question came
back position-level, and whose detail names the tier breakdown — because a coverage answer that does
not say which tier answered is the defect this lane is about.
No SourceRow either, and for the converse of @write-the-sourcerow: sources.csv travels to the
registry meaning this module uses this source, and nothing here reaches a module's tables.
identifiers.py is the precedent, civic_draft.py — which does put registry-derived values into a
module and does write its clingen_allele_registry row — is the contrast. The source is named on the
VerificationRecord instead.
What the corpus actually gets, measured 2026-09-01¶
Run over the 11 reference modules carrying a resolution.csv — 389 loci, one rsID shared by two
modules — with one shared client so a locus two modules name is one request:
| loci | |
|---|---|
| answered at allele tier | 165 |
| answered at position tier only | 92 |
| absent from the index at any tier | 122 |
| could not be asked | 10 |
| with at least one CAID node at the locus | 180 (46.3 %) |
Every locus in this corpus asks an allele-level question — no module names an rsID and no allele — so
all 92 position-tier answers are the finding rather than the question. Of the 180 loci that carry a
CAID node, 165 resolve to the module's own allele; the other 15 split three ways, and each keeps its
own reason: 10 allele_not_comparable (the registry holds a one-sided indel this module's row cannot
anchor), 3 allele_nodes_name_other_alleles, and 2 where the position node lists a CAID the index has
no node for.
14,168 papers across the corpus sit on a position node that no allele node claims, and APOE is most of the top of that list: rs429358 contributes 3,617 of 3,945 and rs7412 3,083 of 3,597. The PGx loci run about half — rs4149056 is 575 of 1,342, rs1057910 606 of 1,093. A pass that reported the allele node's count as the locus's answer would have been quietly wrong by those margins on nine of the eleven modules.
The 122 absences concentrate where you would expect a literature index to hold nothing — 105 of them
are in pathogenic_clinvar, a panel of rare ClinVar variants. What the measurement establishes is that
the index holds no node for those rsIDs; whether that is because no paper names them or because
nothing indexed the paper that does is a question this surface does not answer, and it is recorded as
absent rather than as either reading.
The bound: it answers which papers discuss an identified allele, not which allele a name meant¶
Do not reach for this to recover an identity. Those two read as the same question and are not, and
the measurement is on the two hardest records in this repository. CIVIC_LEGACY_INSERTIONS
works CIViC 1955 (VHL P71fs (c.211insT)) and 2131 (VHL Q73fs (c.214insGCCC)) down to four candidate
alleles with registered CAIDs. Asked of LitVar on 2026-09-01:
| asked | answer |
|---|---|
| CA2586965638, CA2501268513, CA2573048346, CA2499307076 | no node for any of the four |
c.211insT, 211insT, c.214insGCCC |
no node |
VHL P71fs |
one node — litvar@#7428#p.P71fsX, 1 PMID: 19996202 |
That single hit is none of the four source papers; it is an unrelated paper that happens to write
"P71fs" (@existence-not-identity), and its id is the fifth shape — litvar@#<gene_id>#<protein_name>,
all three flag_* false, an unnormalized text mention rather than a variant. Nothing here treats
one as an identity, and litvar gene prints the mention count precisely so a reader can see how much
of a gene's tail is text: of 588 HFE nodes, 220 are rsID-only, 69 carry a CAID, exactly one is the
gene node (litvar@#3077#, 3,285 papers) and the remaining 298 are mentions.
The reason is structural, which is what makes it a bound rather than a gap. PubTator3's export for all four source papers returns title and abstract only — two passages, and zero variant annotations in every one. None is in the PMC open-access subset. The alleles live in Table 3 of a paywalled 1996–2007 paper, and text mining over abstracts cannot reach a table. §8 of that probe went around the same wall through UMD-VHL's curated protein column — a curator's tabulation, one step from the primary. LitVar indexes text.
Two API facts, and one of them is a defect¶
variant/search/gene/GENEreturns line-delimited Pythonrepr(), not JSON — single-quoted keys, one dict per line.httpx's.json()raises on it. The other endpoints return proper JSON.parse_repr_linesis a literal parser (ast.literal_eval), nevereval, and a test walks the module's AST to assert nothing in the lane calls.json()at all. That endpoint also carries noflag_*at all, so the tier there comes from which keys a record has.autocompleteis a prefix search.?query=rs429358returnsrs42935848as a second hit — a real node for a different variant. Taking[0]off that list answers confidently about the wrong thing, so both lookups filter on an exact rsID or CAID. And node ids are never constructed from the grammar: every id handed togetorpublicationscame back from a listing verbatim. The first pass at this source read the trailing##as a suffix on an rsID, asked one endpoint, and concluded there was no allele tier at all — a confident negative about a whole tier, from one misread character.
An absent node is an answer: autocomplete returns 200 [], and variant/get returns 400 with
a body opening Variant not found. The discriminator is the body and never the status, because a
malformed query is also a 400 — the same rule as Ensembl's 400 on an unresolvable rsID.
Licensing: NCBI publishes a policy, and a policy is not a licence. NCBI states it "places no
restrictions on the use or distribution" of molecular data and, in the same passage, that it "cannot
provide comment or unrestricted permission concerning the use, copying, or distribution" because
submitters may hold rights it cannot assess. ClinVar escapes that through its own maintenance_use
page, which is why CLINVAR_TERMS records public-domain; LitVar has no such page, so under
@no-named-licence its gating axes are unknown rather than permissive. Recording it as public domain
by analogy with ClinVar is exactly the move that rule forbids. This is NCBI's side only — nothing
was read about EMBL-EBI's terms for the surfaces EBI co-hosts, and nothing here asserts anything about
them. There is deliberately no LITVAR_TERMS constant, and the rule is worth stating because that
file has two kinds of entry rather than one: a TERMS_BY_SOURCE member earns its place either through
a pass that records it into sources.csv — clingen_allele_registry through civic_draft, pubmind
through its drafting provider — or through a snapshot that keeps the source's bytes on disk, where the
unknown redistribution axis becomes load-bearing the moment somebody proposes publishing them, which
is MANE_TERMS. This lane has neither: nothing it reads reaches a module's tables and nothing is
stored. So the finding lives here and in the module docstring, which is where a reader would look.
Regulator drug labels (drug_labels.py + drug_labels_build.py) — clinpgx check-labels — RM166¶
ClinPGx's drugLabels.zip annotates the pharmacogenomic content of medicine labels published by
five agencies: FDA, Health Canada (HCSC), EMA, Swissmedic and PMDA. The item asked for the FDA and
the file supplies four more at no extra cost, which changes the shape from module ↔ authority ↔
authority into a lane where the number of authorities is a parameter — exactly what RM134's
vocabulary split was built to survive. Nothing in the surface names an agency. Baking one into a
published key is the mistake RM134 caught in ClinSigConflict before it shipped, so the check is
regulator_label_agreement, the command is check-labels, and the agencies are data in a regulator
column.
just-dna-enricher clinpgx build-labels --use non-commercial # → data/repro/drug_labels
just-dna-enricher clinpgx check-labels spec/ --snapshot data/repro/drug_labels --use non-commercial
A second archive from a source already adopted, and a broader finding behind it¶
Enumerating api.clinpgx.org/v1/download/file/data/ properly — a 303 is a real file, a 404 is not —
finds at least twelve published archives: clinicalAnnotations, variantAnnotations,
clinicalVariants, drugLabels, relationships, variants, genes, drugs, chemicals,
phenotypes, occurrences, pathways-tsv. Before this item, clinpgx_build downloaded one of them.
So the honest restatement of RM166 is that the PGx lane reads a fraction of a source it has already
adopted and gated, and the FDA question was a narrow way into a broad finding. clinicalVariants.zip
is the next one that bears on a shipped table kind and is deliberately not built here.
drugLabels.zip was 59 KB on 2026-08-05 and holds LICENSE.txt, README.pdf, drugLabels.tsv
(1,433 rows) and drugLabels.byGene.tsv (238 rows). The second is a pivot of the first by gene symbol
and carries no fact the label table does not, so nothing reads it.
Its own release.json, never the annotation lane's. clinpgx_build's docstring records
relationships.zip a year newer than clinicalAnnotations.zip, so the archives do not refresh in
lockstep — and RM175 later found the reason that gap was so wide, which was that ClinPGx had retired
the annotation archive under that name. The label snapshot is dated from its own CREATED_<date>.txt and labelled
clinpgx_drug_labels_<date> — distinct from clinpgx_<date>, because two surfaces have two
denominators (@two-surfaces-two-denominators).
Every cell is stored verbatim. Seven of the fifteen columns are flag-shaped — a blank or one
constant string — and coercing them to booleans would have the builder decide that Biomarker Flag is
one, which it is not: it carries three values, On FDA Biomarker List and Formerly on FDA Biomarker
List beside the blank. A blank becomes None and the reader decides what a cell means.
The separator is ;, and it is not vocab.MULTI_SEP¶
MULTI_SEP splits on ,;|. In this file a comma inside a Genes / Chemicals /
Variants/Haplotypes cell is data: DPYD c.1129-5923C>G, c.1236G>A (HapB3) is one haplotype named
by two variants, and Ascorbic acid (vitamin C), combinations is one chemical. Splitting on it reports
604 variant tokens where the file states 601 and turns three real names into six that match nothing.
Read with ;, the cell holds 601 tokens over 189 distinct names, 415 of them rsID-shaped and 178
star-allele-shaped.
Two join tiers, because they are not the same claim¶
Genes is populated on 1,248 of 1,433 rows (87 %) and Variants/Haplotypes on 217 (15 %). So a
module's claim is put at two granularities and the tier is a property of the subject:
(gene, drug)— the gene tier. What do the agencies say about this gene and this medicine?(gene, allele, drug)— the allele tier, where the allele is a star allele fromdiplotypes.csvor an rsID frompharm_variants.csv.
A label naming CYP2C19*2 answers both, because those are two questions rather than one asked
twice. The alternative — scoring every authored allele against whatever the gene-level labels say —
was written first and measured: on reference_examples/cyp2c19_star_alleles it reported the EMA's
single disagreement about clopidogrel 34 times, once per star allele, none of which the label
mentions. Under the two-granularity shape the same run reports it once at the gene tier and seven
times at the allele tier, for the seven alleles the labels actually enumerate.
The star tokens are not haplotypes.csv's key verbatim, which the item's entry says they are. The
file writes CYP2C19*2; the module writes *2 in haplotype_name with CYP2C19 in its own column,
so the join composes them. It composes them two ways, because the file spells a gene-qualified
token two ways: the star alleles run together (TPMT*3A) and the DPYD haplotypes are spaced
(DPYD c.2846A>T). Trying only the concatenation told a DPYD module its allele was named by no label
while two regulators named it exactly, which is a false coverage claim rather than a miss. The bare
spelling is tried last, for an rsID and for an HLA allele already authored with its gene.
Composition is spelling, never an allele-definition mapping, and the limit is worth stating. A
DPYD module authoring *2A gets DPYD*2A / DPYD *2A / *2A and matches nothing, because the file
spells that allele DPYD c.1905+1G>A (*2A) — a star name and the variant that defines it are the same
allele under two naming systems, and knowing that is PharmVar's job rather than this join's. So the
allele tier answers where the two sides spell the same allele the same way; where they do not, the
gene-tier subject still answers and the allele is counted as named by no label.
An authored allele no label names is neither withheld nor a finding: the gene-tier subject for the same pair is what answers for it, so it is counted and reported as a coverage number.
Testing Level is five members and a third of the file states none¶
Testing Required 374, Actionable PGx 312, Informative PGx 162, No Clinical PGx 87, Testing
Recommended 26 — and 472 rows stating none. A blank is an absence, not a no: it is unknown, it
is counted, and it withholds. Kleene, so it cannot un-see a disagreement already witnessed and it never
establishes an agreement on its own. The vocabulary is derived from the payload and a test asserts the
equality against the fixture, because a sixth member upstream has to be a visible edit rather than a
silent .get(x, default).
Two verdicts, and only the ends of the axis are placed¶
classify_labels is a pure function over an authored action and N calls at one tier, and it answers
two orthogonal questions the way classify_concordance does:
concordance—concordant/discordant/single/unstated/none. Level equality, with no ordering at all, so a level ClinPGx adds later still classifies correctly here.position—opposed/unplaced/unchecked/absent/no_label.
The module carries no testing-level column, so most of the axis is deliberately unplaced. Mapping
Testing Required onto recommendation_strength=strong would be this format inventing an equivalence
between a regulator's testing requirement and CPIC's prescribing strength. No Clinical PGx needs no
mapping — it is the negative claim by its own name — so the one authored arm fires when a module ships
a prescribing recommendation for a pair every agency that spoke calls No Clinical PGx. The three
middle levels are stated and unplaced, which is reported as such rather than as an agreement. The
reverse direction withholds too: a module declining to recommend for one diplotype is a statement about
that genotype, not about whether the medicine's label carries pharmacogenomics (@refutation-withholds).
evidence_level is not read here at all. It is ClinPGx's own metadata about its own annotation and
clinpgx check owns it; treating "there is 1A evidence" as "this module recommends" would put two axes
in one field.
What the corpus says¶
cyp2c19_star_alleles— clopidogrel and CYP2C19 isActionable PGxat four agencies andInformative PGxat the EMA. Three of the five name the star alleles and two name only the gene, so the same disagreement is established at both tiers and reported apart.pgx_slco1b1_simvastatin— four labels reach SLCO1B1 + simvastatin and two of them state no level. The two that do are both Swissmedic's, at different levels: one covers simvastatin and one the fenofibrate/simvastatin combination, and both namers4149056. So the pair isdiscordant, one agency disagrees with itself, and neither of its rows is picked as its opinion (@multiplicity-is-a-finding). The two blanks are counted into the record and named in the sentence as stating no level, never folded into the count of labels that stated one.cyp2c9_warfarin_grch37— CYP4F2 + warfarin is a claim no agency labels, at either tier. Withheld at both, named in the record, and never reported as an absence of pharmacogenomics. It is also what keeps the coverage number honest: an allele only counts as "answered at the gene tier" when the gene tier actually answered, and CYP4F2's did not.
Severity, and what is not written¶
--strict never escalates, and a test asserts the two modes report the identical list. Five expert
regulators genuinely disagree with each other and with a curator, and a compile that refused would make
this format arbitrate between its own authorities — the rule the ClinVar clin_sig, PGx
allele-function and repeat-band checks already follow. What strict still refuses is structural: a
diplotypes.csv that will not load raises in both modes.
This check writes no SourceRow, and either of two reasons would settle it. Nothing from the labels
lands in the module — it writes no authored cell, and sources.csv accounts for what a module carries
(@write-the-sourcerow's converse; check-repeat-bands is the shipped precedent). And the row is not
free: merge_sources_csv keys on (source, layer), clinpgx/annotation is already owned by
clinpgx check and draft-clinpgx, and that row's dataset is load-bearing — the evidence-level
check's tautology guard compares it against the annotation snapshot's label. Stamping
clinpgx_drug_labels_<date> into that slot would silently disable a shipped check. Every other layer
sits outside the compiler's orphan-check exemption and would warn source_row_unused on every module.
The licence gate still applies, at clinpgx_draft's reading rather than a third one: ClinPGx's
terms are accepted when the data is taken, so --use commercial refuses the read as well as the fetch,
and the skip is not_permitted — what clears it is a declaration, not egress.
What closed with this item¶
The entry wanted two things from the FDA: this concordance check, and a PGx lane member whose terms may not gate. The second is refuted by both routes and closed in writing. The ClinPGx route is CC BY-SA + no-sale, the same gate the rest of the lane sits behind, so it diversifies nothing. FDA's own Table of Pharmacogenetic Associations, probed directly on 2026-09-01, is 126 associations in an HTML page with no CSV or XLS download and no copyright or public-domain statement on the page at all — a quarter of the FDA content ClinPGx already carries, in a shape that has to be scraped, on terms that are unestablished. "US government work is public domain" is a rule with exceptions and the page does not settle it. Licence diversification for this lane is still worth doing, and it wants its own entry with candidates chosen for their terms first, which is the opposite of how this one chose.
AlphaGenome — a local artifact, a live service, and two licence classes (alphagenome_*, atlas_*) — RM191–RM200¶
AlphaGenome is the only source in this tier reached on two surfaces that are not the same source.
One is a bulk artifact an operator already holds, re-encoded into a cache lane and read offline; the
other is a gRPC service queried per variant. They answer different questions, carry different terms,
and — since RM200 — write two different sources.csv rows, because one (source, layer) key
cannot carry two licence classes.
| Surface | Module | What it answers | Licence class |
|---|---|---|---|
AVI artifact → the alphagenome_avi lane |
alphagenome_avi_build |
how much does AlphaGenome think this variant matters, for 8.8 billion SNVs, offline | Permissive Use (RM195), commercial_use=True — so a module drafted from it stays sellable |
Atlas API (gdmscience.googleapis.com) |
atlas_client, atlas_protos |
the scores the artifact discarded — per gene, per tissue, with direction | non-commercial only; an undeclared run writes nothing |
The split is not a convenience. ALPHAGENOME_AVI_TERMS and ALPHAGENOME_ATLAS_TERMS are separate
SourceTerms constants under separate source names, and a module that used both would declare both
rows — which is the only way sources.taints_commercial_use can reach the right verdict. Reading the
Atlas under --use unstated is refused at acquisition, not at write time (@acquisition-gate-is-not-a-read-gate).
alphagenome build — the artifact, re-encoded (RM191, RM197, RM198)¶
Nothing here fetches, and that is not the usual inject-only rule — this is the network tier, so it
is allowed to. The AVI artifact is 88.5 GB behind a sign-in whose eligibility clause bars whole
classes of holder, so acquisition is the operator's own act under their own acceptance of the terms.
--input is required and has no default URL, which is that rule expressed as a flag.
Three properties of the built lane a consumer has to know, each of which is a measured decision rather than a format choice:
- The scores are integers and should be compared as integers.
raw_scoreisInt32at a scale of 10⁵ — exactly lossless, since the source prints at most five decimals. Recovering a float withraw_score_e5 / 1e5disagrees with the printed value on 53% of rows, because the division rounds a second time.score >= 0.1israw_score_e5 >= 10_000. PHREDis not stored, and the file that reconstructs it is not optional (RM198). It is an exact within-corpus rank, soavi_knots.parquetcarries the curve in 466 KB instead — as an interval per printed score, which makes threshold safety decidable in advance: a threshold is unsafe iff it lands inside a knot's span. Genome-wide exactly one does, at 3. A snapshot without that file holds scores nobody can rank, andalphagenome checkrefuses it. It is a sibling ofdata/, which is whylocations.SNAPSHOT_ROOT_FILENAMESexists. Both halves walk that tuple — the publisher did from the start and the puller did not until RM209, which is why a pulled lane held scores nobody could rank while two docstrings said the registry was walked.- The artifact is wide by position (RM197): one row per locus —
chrom, pos, ref, alt0, alt1, alt2— and noaltcolumn. Which base each column means is{A,C,G,T} − refascending, a function ofrefalone, so nothing has to travel beside the data.alphagenome_avi_build.to_long()recovers(chrom, pos, ref, alt, score)rows; apply it to a filtered frame, since over the whole corpus it is 2.9 billion loci becoming 8.8 billion rows, which is the shape the layout exists to avoid storing.
The lane is the first with a builder and rebuild=None: this tier cannot fetch the source, but the
re-encoded snapshot is publishable, so an operator who may not download the artifact can still
cache pull it. Having a builder and being rebuildable unattended are two different properties
(@a-cache-lane-has-three-stages-and-a-list-cannot-say-which-are-missing).
alphagenome check — the Atlas as a resolver, for the three questions the artifact cannot answer (RM193)¶
Attests as variant_impact_agreement, its own member rather than a second writer of
reference_allele — the Atlas answers that too, and two registries answering an overlapping question
get two names, or one source's outage writes a skip against the other's
(@one-registrys-outage-may-not-speak-for-another). Reports, never repairs. Mostly offline: without
--threshold there is no question the local snapshot cannot answer, so the live leg is the exception
rather than the path. Raises VariantImpactError / VariantImpactUnavailable — see Exception contract.
alphagenome expression — the axis the AVI artifact threw away (RM194 + RM200)¶
Fills expression_effects.csv: one row per (variant, gene) saying which way a variant moves that
gene's predicted expression, how many of AlphaGenome's 371 tissue tracks agree, and how far the variant
sits from the gene. A recording pass — it writes a derived sidecar, compares nothing authored, and
therefore has no verification-check member.
One scorer of twenty-two, and the other twenty-one are a measured refusal. RNA_SEQ is the only
one with a gene axis — its shape is (genes, tracks), the gene count varying with the window — so
AlphaGenome makes the gene attribution itself, which is what @gene-map-is-another-sources-attribution
requires. CHIP_TF looked like the useful one and was refused on measurement: over 751 factors the top
three carry 1–3% of the mass, the concentration does not track effect size, and every leading factor is
a singleton track, so the "top TF" is whichever one happens to be measured once. *_ACTIVE is an
activity level rather than a variant effect — it describes the locus, not the variant.
--gene is mandatory in both span forms, because the server-side gene filter is a requirement and
not an optimisation; an unfiltered interval query is refused before it is sent. The span comes from
--chrom/--start/--end when given, otherwise from the MANE lane widened by the ±512 kb horizon —
the Ensembl snapshot is eliminated by its own schema, which has no gene column at all.
distance_to_gene is the column that makes the table usable, and null is one of its values. Distal
scores run ~10× lower than scores at the gene, so a flat --min-score keeps only the proximal rows
while looking like it filtered on effect — the failure this pass exists to prevent. The MANE lane is
consulted even when the interval was supplied by hand, because those are two questions and an explicit
interval only answers the first; with no lane the column is null and the manifest publishes
without_distance, so a whole table of them is visible before a join. Never an interval edge
(@a-derived-lane-has-parents-and-an-absent-parent-is-not-an-empty-result).
--max-rows defaults to 50,000 and refuses rather than truncating; raising it is a deliberate act.
--offline is a no-op with a warning — this pass reads the Atlas, not a snapshot, so there is no
offline answer to give and pretending otherwise would report nobody-asked as nothing-found.
atlas generate — the bindings, and why neither the sources nor the generated code is committed (RM192, RM196)¶
uv add alphagenome costs 550 MB and 47 packages against a tier whose entire runtime list is
httpx/tenacity/huggingface-hub, and six of the twenty dependencies that wheel declares are never
imported on any scoring path. The .proto sources are Apache-2.0, so grpcio + protobuf reach every
Atlas RPC — 19 MB, the figure enricher/pyproject.toml measured in a clean venv — with score
payloads decoding through struct.unpack from the standard library. (22 MB appears in older text and
is the grpcio release current at the design round; RM221 swept it.)
That is the [atlas] extra; the alphagenome extra is deleted.
The repository carries neither the upstream sources nor the generated bindings. It carries the pin —
a commit id and a sha256 per file — which is what makes a fetch verifiable rather than merely
convenient, and what keeps a copy of somebody else's file from going stale silently. atlas generate
runs once per checkout; a released wheel carries them already. This is also why just-dna-enricher's
build backend is hatchling rather than uv_build (RM196): the bindings are generated at build time
and uv_build has no build hook. Backends are declared per package, so the other two tiers are
untouched.
Three defects only real execution could find, each one a case where every offline fixture agreed
with the code that built it: the Atlas wants chr22 on the wire where a module stores 22 (now fixed
by the client owning the spelling, as it already owns the 0-based/1-based conversion); a snapshot's
summary.parquet lives under data/ rather than at its root; and a lane directory existing is not
the same as a lane being built. A fourth is upstream's: the server returns a next_page_token on an
exactly-full final page and following it answers INVALID_ARGUMENT, so a faithful client crashes on
the one interval width that divides evenly — score_interval follows the token but also stops once the
requested interval is covered.
And one that cost a day: Interval.strand has no zero member. STRAND_UNSPECIFIED = 0 is the
proto3 default, so an omitted strand goes on the wire as a value the server rejects, reported as a
bare INVALID_ARGUMENT naming no field. Neither the x-goog-fieldmask header nor 32 bp chunking was
ever required — both were measured and both are the SDK's own choices.