just_dna_enricher.verification¶
just_dna_enricher.verification ¶
Recording what a pass CHECKED, so the manifest can say it (RM45).
The enricher is the only tier that can compare authored data against reality, and every such check
already reports itself — to a log line, and onto an EnrichmentResult field that dies when the
process exits. manifest.compilation inherited none of it, so a module whose clinical calls were
cross-checked against ClinVar and one where the check never ran shipped identical manifests.
This module is the load-merge-write that closes that, deliberately in the shape
licensing.record_source_terms already has and for the same reason: RM51 estimated five write sites
for the licence table and found nine, and a count of call sites is exactly the thing that goes
stale. One function, so a new pass wires itself in by calling it rather than by growing a private
copy of the merge.
One proof-of-work per call, which means one per command. The work is ~0.7s of hashing and it
binds the whole document, so a pass that recorded checks one at a time would pay it per check for no
extra guarantee. enrich() therefore collects its records and writes once at the end, literature
writes its own three, clinpgx its one, and the merge below is what keeps several commands' records
in one document instead of overwriting each other.
Which commands actually call this — seven of them, covering every member of
vocab.VALID_VERIFICATION_CHECKS except the ones whose comment there says RESERVED. enrich
(reference allele, wrong build, clinical significance, rsID currency, rsid↔coordinate, dataset
currency, published refutation, evidence-status currency), literature (citation existence, citation
identifier, provenance quote), check-identifiers (gene symbol currency, trait currency, gene↔locus
agreement, PGS accession currency, PGS metadata agreement), clinpgx check, clinpgx check-labels,
pgx, vrs mint, check-acmg, check-repeat-bands and litvar coverage (one each).
The counts in that sentence are gone on purpose, and they were wrong when they were there. It read
"enrich (six: …)" while enrich emitted seven — RM170's published_refutation had landed without
the sentence learning about it, which is the same drift its own next paragraph is about. A number in
prose is a registry nothing iterates (@registry-completeness), so the members are named instead and
the grep below is what settles the list.
Count the call sites before you edit that paragraph, and edit it whenever you add one. The
sentence it replaces said "enrich (four checks), literature (three) and clinpgx (one)" and was
wrong three ways at once — enrich emitted five, and pgx and vrs mint had been wired for a release
without the sentence learning about them. Its own predecessor was wrong the other way, naming
check-identifiers as a writer years before it was one, so the merge was machinery tested against a
document no two commands produced. Its successor then said "fifteen of the seventeen" and went stale
again the moment a sixteenth member landed, which is why the totals are gone rather than corrected: a
number in prose is a registry nothing iterates (@registry-completeness), and "every member except
the ones marked RESERVED" is a rule a grep settles. A count of call sites is exactly the thing that
goes stale, which is this module's own argument for being one function; the honest version of that
argument is to derive the list with a grep rather than from memory:
grep -rn 'record_verification(' enricher/src/ # the writers
grep -n '"[a-z_]*", *#' schema/src/just_dna_format/vocab.py # the members, each with its emitter
examples ¶
A few names and a count, so a per-row list cannot become the message (the CPIC lesson).
Here rather than beside one caller, for this module's own stated reason: the second pass that
wanted it (identifiers, RM72) would otherwise have kept a private copy of the aggregation rule,
and a rule with two copies has one that is about to be wrong.
Source code in enricher/src/just_dna_enricher/verification.py
producer_label ¶
just-dna-enricher <version>, or an honest 'unknown' — mirrors _compiler_version.
Source code in enricher/src/just_dna_enricher/verification.py
record_verification ¶
record_verification(
records: Iterable[VerificationRecord],
spec_dir: Path,
*,
error: type[Exception],
now: str | None = None,
) -> VerificationDoc | None
Fold this run's check records into the module's attestation and rewrite it.
Returns the document written, or None when there was nothing to record — a run that put no
check has nothing to say, and writing an empty attestation would create a file asserting that a
module was checked and nothing was found.
Writes to the file it reads (layout.sidecar_write_path), so running this on a module whose
sidecars live under derived/ does not leave a second copy at the root — the collision that
RM49/RM51 made an error, arrived at by following the documented workflow.
The binding is recomputed here, from the module's authored bytes as they are right now. That
is what makes the record perishable in the way it should be: an author who edits variants.csv
after this ran gets the block dropped at compile with a warning, because the checks were put
against rows that no longer exist.
Existing records for checks this run did not put are kept (verification.merge_records) — a run
that did not ask a question has said nothing about it, and dropping the earlier answer would turn
"not asked this time" into "never asked", which is the collapse this whole item exists to undo.
Since RM72 a skip this run writes does not displace an earlier real answer either, for the same
reason one step further out — but only while the earlier answer still describes the module's
current authored bytes, which is what existing_still_binds carries. Where they have moved, the
older record is about rows that no longer exist and this run's honest "could not ask" wins.
An existing closure (RM73) is carried across, but only while the binding holds. Enrichment writes derived sidecars, which are outside the authored set, so the ordinary case is that a run here leaves a closed module closed — and it must, or every enrichment would silently un-close one. Where the authored bytes have moved, the closure is dropped rather than re-bound: re-stamping it would have this pass assert on the author's behalf, which is the one thing the deliberate-act decision forbids.
Source code in enricher/src/just_dna_enricher/verification.py
97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 | |
ran ¶
ran(
check: str,
*,
subjects: int,
findings: int,
source: str | None = None,
release: str | None = None,
detail: str | None = None,
now: str | None = None,
) -> VerificationRecord
A record for a check that RAN, with the denominator it ran over.
Two constructors rather than one with an optional skipped, because the two shapes cannot both
be filled and the model refuses a record that tries — making the split visible at every call site
is cheaper than discovering it as a validation error.
Source code in enricher/src/just_dna_enricher/verification.py
skipped ¶
skipped(
check: str,
reason: str,
*,
detail: str | None = None,
source: str | None = None,
now: str | None = None,
) -> VerificationRecord
A record for a check that did NOT run, carrying the machine key and the sentence.
The sentence is not optional in practice and is optional in the signature: clinical.tautology_
reason writes a good one and it must survive into the record, while not_requested needs none.