just_dna_enricher.acmg¶
just_dna_enricher.acmg ¶
ACMG secondary findings — the list VariantRow.acmg_sf was, until now, checked against nothing.
acmg_sf has been materialized into weights.parquet since 0.4 and validated by nobody: the compiler
cannot hold a gene list (that is the un-injected-reference mistake RM21 taught), and no pass in this
tier had one to compare against. This module closes that, and the interesting part is why it took
until 0.5.1 — the answer is in parse_acmg_page, not in the check.
There is still no data file, and the scrape is real. Probed 2026-08-03: ClinGen's FTP publishes
gene-curation, region-curation, dosage and recurrent-CNV lists and no secondary-findings list;
ClinVar's own FTP tree carries no ACMG flag (gene_condition_source_id has 13,478 rows and zero
mentions of ACMG). The only machine-reachable form of SF v3.2 is NCBI's adaptation of ACMG's Table 1
at /clinvar/docs/acmg/, as HTML. So this is the "accept the guarded scrape" branch the roadmap left
open, taken with the guards that branch was made conditional on.
The guards are not ceremony — the naive parse is wrong on the real page. Splitting the table on
<tr> yields 78 of the 81 genes, silently. The page is hand-maintained HTML and it shows: two rows
open with a bare <td> after the previous </tr> with no <tr> of their own, four rows leave a
<td> unclosed and carry a stray trailing </td>, and the gene cell links through three different
URL shapes (/gtr/genes/324, /gtr/genes/4089/, /gene/3949). The three genes a <tr> split drops
are TP53, COL3A1 and TPM1 — which is the failure mode stated exactly: a short list makes a
correctly authored acmg_sf=true row look wrong, and it would have started with the single most
recognizable secondary-findings gene there is. The parse therefore works in cells, not rows
(<td> count must divide exactly by the header's column count), and refuses rather than returning a
short list.
This pass records no SourceRow, which is the deliberate exception to the standing rule that a
pass consulting a source must write one. That rule is about a module carrying a source's data:
sources.csv exists so a module can account for what is in it. Nothing here lands in the module —
acmg_sf was authored by a human before this ran, exactly as a gene symbol is, and this asks the
registry whether that authored value is still right. It is check_identifiers' shape (HGNC and OLS4
also go unrecorded), not clingen's.
Gene-level, and the column says so. ACMG's table is keyed on gene and condition — 94 gene-disease
pairs over 81 genes — and reporting is scoped to P/LP variants for the named condition. acmg_sf is
documented as "True when the gene is on the ACMG secondary-findings list", so that is what is
compared; the conditions are parsed and kept on SecondaryFinding so a caller can say which entry
matched, but they are not part of the verdict. Reading the column as per-variant reportability would
make the format decide disclosure policy, which it explicitly does not.
And the page turned out to be a version behind, which no guard above can see. ACMG published
SF v3.3 in June 2025 (84 genes over 100 gene-condition rows, adding ABCD1, CYP27A1 and PLN);
NCBI still serves its adaptation of v3.2 (81/94). The scrape's five guards all pass — the page is
neither truncated nor re-laid-out, it is simply old — so the check reported acmg_sf=true but ABCD1 is
not on ACMG SF v3.2 about a row that is right. That is the same short list failure the guards were
built for, and this is the fix, in two halves:
- the list can now be injected —
acmg_buildturns ACMG's supplementary workbook into a snapshot andload_acmg_snapshotreads it with the standard library, which also makes the check the first one here that works--offline; and - the scrape path carries a staleness tripwire,
KNOWN_LATEST_SF_VERSION. When the list actually read is older than a version this package knows was published, every disagreement is demoted tounverifiableand no longer refuses understrict. A mismatch against a superseded list is a question, not an answer, and answering it anyway is worse than saying nothing (the house tri-state rule: unknown withholds — it never reports and never negates).
One constant is deliberately hand-kept, and its failure mode is why that is acceptable: a version string is not the transcribed gene list this module exists to avoid. When ACMG ships v3.4 the constant says 3.3 and the tripwire under-warns — i.e. degrades to the behaviour of the release before this one — whereas a hand-kept gene list would make specific, confident, wrong claims about specific genes.
AcmgSfError ¶
Bases: RuntimeError
An ACMG SF fetch or parse failed in a way the caller must see rather than work around.
AcmgListUnavailable ¶
Bases: AcmgSfError
No usable list was obtained, so the check could not be put at all (RM72).
A subclass rather than a second exception, so every existing except AcmgSfError still catches
it. It exists because this module has two ways of raising and a caller that attests must tell
them apart: the list could not be read, or the module disagrees with a list that was read
perfectly well (the strict refusal). Recording the second as a skip would say the question was
never put, on the one run where it was put and answered badly — the answered-absence-versus-
unasked-question collapse this tier draws everywhere else.
skip carries the VALID_VERIFICATION_SKIPS member, decided where the failure happens rather
than sniffed from the message: unreachable when the source was asked and never answered, and
no_reference when something was there and no list could be read out of it.
Source code in enricher/src/just_dna_enricher/acmg.py
SecondaryFinding
dataclass
¶
SecondaryFinding(
gene: str,
gene_id: int,
gene_mim: str | None = None,
disease: str | None = None,
disease_mims: tuple[str, ...] = (),
medgen_ids: tuple[str, ...] = (),
phenotype_category: str | None = None,
inheritance: str | None = None,
since_version: str | None = None,
variants_to_report: str | None = None,
)
One row of ACMG's Table 1: a gene, and the condition it is reportable for.
disease_mims and medgen_ids are tuples because the real page puts several of each in one
cell — SDHB is listed for "Hereditary paraganglioma-pheochromocytoma syndrome (MIM 115310,
MIM 171300)" against MedGen C1861848, C0031511, and the MedGen cell then links to a search
(/medgen/?term=…+OR+…) rather than to a concept. Taking the first of each would have been a
silent truncation of the same family as the <tr> split.
AcmgSfList
dataclass
¶
AcmgSfList(
version: str,
findings: list[SecondaryFinding],
retrieved_at: str,
source_url: str = DEFAULT_ACMG_URL,
)
The parsed list, with the version it declares about itself.
superseded_by
property
¶
The newer SF version this package knows about, or None when this list is the newest known.
Drives the demotion of every disagreement to unverifiable: a list that is a release behind
cannot answer whether an authored flag is right, only whether it matches an old answer.
entries_for ¶
Every condition this gene is listed for, in page order (P7: never set order).
AcmgVerdict
dataclass
¶
One variants.csv row's answer, over the tri-state the authored column actually has.
agree— the authored value matches the list, either way round. Silent.not_listed— authored True, gene is not listed. A finding.denied— authored False, gene is listed. A finding, and the message names the entry.unverifiable— a disagreement (either of the two above) against a list this package knows is superseded. A warning, never astrictrefusal: ACMG SF v3.3 added three genes NCBI's v3.2 page does not carry, sonot_listedagainst v3.2 is exactly as likely to be the list being old as the module being wrong. Reporting it as a defect made the check confidently wrong about correctly authored rows, which is worse than not checking.unstated— blank, gene is listed. A note, never a finding: blank means "not stated", and turning that into a defect is exactly theNone-means-Falsecollapse this codebase refuses.blank— blank, gene is not listed. Nothing was asserted and nothing contradicts it.unchecked— the row names no gene, or the list could not be reached. The question was never asked, which is not the same as being answered "no".
AcmgReport
dataclass
¶
AcmgReport(
version: str | None,
verdicts: list[AcmgVerdict] = list(),
warnings: list[str] = list(),
)
mismatches
property
¶
The verdicts that are defects — what strict refuses on.
unverifiable
property
¶
Disagreements against a superseded list — reported, never refused. See AcmgVerdict.
checked
property
¶
Rows the question could actually be asked about — never the row count.
not_consulted
property
¶
Why no comparison happened, or None because one did — the two arms of a skip (RM234).
verification_record reads this rather than re-testing the same two conditions, so the
in-memory verdict and the attestation cannot disagree about whether this run compared
anything. They did disagree: the record has always returned a skipped here while clean
went on answering True.
clean
property
¶
Whether every stated acmg_sf agrees with the list, and what is wrong when it does not.
A Verdict since 2026-09-13, bool | None for one day before that, and a bare bool until
RM234. The original bug is unchanged and worth keeping in view: clean was
not self.mismatches, and mismatches selects not_listed/denied, so a run that reached no
list gave every row the verdict unchecked, left mismatches empty, and answered True — a
comparison that never happened reporting as one where everything agreed (@tautology-zero).
RM234 fixed that by withholding, which is the house rule for an answer nobody can give. The
retrofit is because a gate is the one shape that cannot withhold: --strict has to choose
an exit code, and None made the caller do the arithmetic anyway. The reasons now travel with
the verdict instead (verdict.py has the whole argument).
nothing_to_check is a pass now, and that is the one behaviour change. RM234 gave None
to both arms of not_consulted. Only one of them is an error: offline means no list was
obtained, so nothing was compared and the run cannot certify; nothing_to_check means the list
was read and the module states no acmg_sf cell, which is a module with nothing to disagree
about rather than a failed check. checked sits beside it as the denominator, the same way the
identifier command prints what it read.
if report.clean: is correct across all three shapes of this property, which is why the change
is safe to make twice. The arms stay not_consulted's, so this and the attestation cannot
disagree about whether anything was compared.
by_gene
staticmethod
¶
[(gene, [row numbers], message)] — the shape a human should be shown.
Every verdict here is about a gene, so a per-row list repeats one sentence once per row: the HFE reference example is 13 variants in one gene and printed the same 220-character note 13 times. That is the aggregation rule this codebase already learned from CPIC (~600 identical lines for CYP2C19 buried every other finding). The per-row verdicts stay on the report for a caller that wants them; the grouping is what a report prints. First-occurrence order, never set order (P7).
Source code in enricher/src/just_dna_enricher/acmg.py
parse_sf_version ¶
ACMG SF v3.3 / ACMG Secondary Findings v3.3 in text → "3.3", else None.
parse_acmg_page ¶
The NCBI ACMG page → the list it publishes, or a refusal.
Five guards, each of which is a way the page has actually been observed to fail or could fail into a short list rather than into an error — the outcome that would make this check worse than no check at all:
- the version string must be present (it is also the
datasetlabel); - the table must carry the four headers
EXPECTED_HEADERSnames; - the
<td>count must divide exactly by four — the page's rows are not reliably delimited, so the cells are the only trustworthy unit and a remainder means the shape moved; - every four-cell group must yield exactly one gene link (zero would drop a gene silently, two would mean the cell boundaries slipped);
- at least
MIN_GENESdistinct genes must survive.
Source code in enricher/src/just_dna_enricher/acmg.py
fetch_acmg_page ¶
Download the SF page (~75 KB of HTML — one GET, no pacing needed).
Source code in enricher/src/just_dna_enricher/acmg.py
check_acmg_sf ¶
Compare each row's authored acmg_sf against the fetched list. Reports, never repairs.
Source code in enricher/src/just_dna_enricher/acmg.py
load_acmg_snapshot ¶
An acmg build snapshot directory → the list, read with the standard library only.
Deliberately csv and json rather than the workbook reader: this is the runtime half, and a
runtime pass may not depend on a [dev] package (the rule clinpgx.py learned by reading its own
snapshot with polars). openpyxl stays in acmg_build, where it is a builder dependency.
The release.json version is authoritative and must agree with nothing else — the CSV carries no
version column, precisely so the two cannot drift.
Source code in enricher/src/just_dna_enricher/acmg.py
verify_acmg_sf ¶
verify_acmg_sf(
variants: list[VariantRow] | None = None,
*,
spec_dir: Path | None = None,
mode: str = "best_effort",
offline: bool = False,
url: str = DEFAULT_ACMG_URL,
page_text: str | None = None,
snapshot_dir: Path | None = None,
) -> AcmgReport
Read the list — injected snapshot first, then the live page — and check variants against it.
Pass either variants or spec_dir (RM41). The row-taking form stays, because it is the right
thing for an in-process caller that already holds the rows; spec_dir= matches the shape every
other pass in this tier has, and spares a caller re-deriving the load — which is not
csv.DictReader plus Model(**row) (see compiler.load_spec_variants).
An injected snapshot_dir wins over the network unconditionally, which is the whole point: it is
both the newer list and the one that works --offline. With no snapshot, offline reports nothing
checked rather than nothing found — the same unchecked ≠ absent distinction every other pass in
this tier draws.
strict escalates a mismatch to a refusal; list membership is a published fact, not a clinical
judgement, so unlike the clin_sig cross-check there is no reason to hold this one at a warning.
It does not escalate an unverifiable — see _disagreement.
Source code in enricher/src/just_dna_enricher/acmg.py
619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 | |
verification_record ¶
This pass's one check, as a record verification.json can carry (RM72).
subjects is report.checked — the rows the question could actually be asked about, which
already excludes a row naming no gene and every row of a run that reached no list. Counting all
the verdicts would publish a comparison for rows nobody compared, which is the shape the reference
-allele pass fell into on an unbuilt assembly.
findings is mismatches alone, and the detail sentence is what keeps that honest. A
disagreement against a superseded list is demoted to unverifiable by _disagreement — every
one of them, in both directions — so a run against NCBI's v3.2 page can hold ten disagreements and
still have an empty mismatches. Recording findings=0 with no further word would read as a
clean bill, so the count of unsettled disagreements travels in detail and release names the
list that could not settle them. unstated notes are authoring aids and are named there too;
they are deliberately not findings, because blank means "not stated" and turning that into a
defect is the None-means-False collapse this codebase refuses.
The offline-with-no-list return is a skip, not ran(0, 0): verify_acmg_sf is the only path
that can produce a report with no version at all, and it produces one precisely when no list was
consulted.