just_dna_enricher.clinical¶
just_dna_enricher.clinical ¶
The clinical cross-check: a module's authored clin_sig against ClinVar's own.
The cheapest bite out of the compiler's largest blind spot — "is the annotation right?" It cannot answer that, and does not try: ClinVar is not truth either, and two curators reading the same evidence may legitimately land in different places. What it can do is make sure an author who calls a variant benign knows that ClinVar calls it pathogenic, and with how much review behind it.
Why this lives in the enricher even though it never touches the network. The tier boundary is not
"online vs offline", it is does the check need a reference. The compiler is inject-only by charter
(Principle 2) and holds no ClinVar; this check needs one, so it belongs here — beside
sequences.verify_reference_alleles, which needs the genome for the same structural reason. Reading
that boundary as "offline ⇒ compiler" would put it in the tier that cannot host it.
It reports; it never repairs, like every check in this tier, and:
strict does not escalate it. This is the one deliberate exception to severity-follows-the-mode,
and the reason is charter-level rather than pragmatic. Failing a compile on a ClinVar disagreement
would make the format arbitrate a clinical dispute — deciding that ClinVar is right and the curator is
wrong — which is precisely what the data-agnostic north star forbids: the format supplies annotation
tables and never a gene–disease inference. A curator who has read the primary literature and disagrees
with a one-star ClinVar submission is doing their job, not corrupting their module. So both modes warn,
and review_stars is carried into the message so a reader can weigh the disagreement themselves.
ClinSigConflict
dataclass
¶
ClinSigConflict(
variant_key: str,
genotype: str,
chrom: str,
start: int,
ref: str,
alt: str,
authored: str,
authority_clin_sig: str,
review_stars: int | None,
review_status: str | None,
condition: str | None,
opposed: bool = False,
authority: str = CLINVAR_AUTHORITY,
)
One authored clinical significance an annotation authority's records do not support.
The authority is a field, not a field name (RM130). This carried clinvar: str until 0.7,
which read fine while there was one archive and would have cost a rename — major-only work under
the additive rule — the moment a second one arrived. clinvar survives as a read-only alias so
an existing caller keeps working; new code reads authority_clin_sig and authority.
PlannedComparison
dataclass
¶
PlannedComparison(
variant: VariantRow,
authored: str,
targets: tuple[tuple[str, int, str, str], ...],
locus_wide: bool,
)
One authored clinical call and the alleles an authority will be asked about for it.
The unit of the reference-independent half of the comparison: which subjects have a claim to
check and which coordinates that claim lives at. Extracted from compare_clin_sig by RM134 § B
because a second authority has to be asked about exactly the same alleles — building the plan
twice, once per snapshot, is how the two legs would come to disagree about what was asked.
subject
property
¶
The (variant_key, genotype) this comparison is about — the concordance record's key.
ClinSigComparison
dataclass
¶
ClinSigComparison(
compared: int = 0,
no_record: int = 0,
conflicts: list[ClinSigConflict] = list(),
subjects: list[ConcordanceSubject] = list(),
)
What the cross-check actually put, and what it found (RM4, simplified by RM73).
Two numbers and the findings. It carried a provenance breakdown until 0.6 — copied / authored /
conflicting / no_record — which existed for one reason: the module-level release marker could not
see a cell edited after the draft, so strict paid for a per-row look-up to recover the split the
marker was blind to. provenance.draft_digest answers that question directly and offline, so the
breakdown had no remaining consumer and the mode ladder under it collapsed: this check now behaves
identically in both modes, which is what RM4 wanted and could not have.
Removing it changed no verdict. The copied bucket was an early exit taken before the camp
logic, and an exact match agrees with itself, so every row it caught already reached "no conflict"
by the path below it.
compared— comparisons the snapshot could answer. One per resolved locus, which is one per row in the ordinary case, and the denominatorverification.jsonpublishes beside the findings.no_record— resolved, but nothing at those coordinates to compare (or no ALT to ask about). Counted rather than folded intocompared, which would claim a comparison nobody made.conflicts— the disagreements, exactly as before.subjects— every(variant_key, genotype)the comparison put a question about, with the authority's answer attached (RM130). The input to the concordance record, kept here rather than rebuilt by a second pass over the snapshot: re-asking would cost the whole comparison again, and a second implementation of the record-selection rule is a second place for it to drift from the one that decided the conflicts.
AuthorityLegOutcome
dataclass
¶
AuthorityLegOutcome(
authority: str,
state: str,
dataset: str | None = None,
reason: str | None = None,
)
What happened when one annotation authority was approached, before any subject was asked.
Three states, and the third is the one this release exists to keep honest:
consulted— a snapshot was open and answered. The only state that produces evidence.unchecked— nobody asked. No snapshot was provisioned, or one was present and would not answer. Never an absence of records and never agreement (@unreachable-not-absent).tautological— asking would compare the authority's own values against themselves, because this module's rows were drafted out of this very snapshot and have not moved since. The authority is left unconsulted and states nothing: a call recorded here would agree with the module by construction, and the record would publish that agreement as though somebody had checked it.
consulted
property
¶
Whether this leg produced evidence — the only state that lets a record be written.
ConcordanceRecord
dataclass
¶
ConcordanceRecord(
parents: list[ClinSigConcordanceRow],
calls: list[ClinSigAuthorityCallRow],
legs: tuple[AuthorityLegOutcome, ...],
subjects: int,
multi_record_subjects: int,
internally_contested: int,
)
A run's concordance record: the two tables, plus everything the run measured making them.
Nothing here is computed and discarded (@dont-discard-computed). The counts below are what
a caller would otherwise recompute, and recomputing them means a second implementation of the
selection rule — the drift RM130 built this record to avoid.
is_tautological_leg ¶
is_tautological_leg(
sources: Sequence[SourceRow],
authority: str,
dataset: str | None,
spec_dir: Path | None,
) -> bool
Would consulting this authority compare its own values against themselves? (RM134 § B)
The conjunction RM4 established and RM73 completed, stated once and applied per authority: the module's licence row must name this release of this source — so the values really were copied out of what is about to be read — and the drafter's digest over the checked column must still match, so no cell has moved since. Either half missing runs the leg, which is the conservative direction: an unknown is never a permission to skip.
Per leg, never per module, and that is the whole reason it is a predicate rather than the
module-level tautology_reason generalized in place. A module drafted from ClinVar still gets a
real comparison out of a second authority that copied nothing, so skipping the whole check would
throw away a genuine finding to suppress a hollow one — the shape enrich_pgx already found the
hard way when its CPIC leg went unmarked for two releases.
Returns a plain bool rather than a tri-state on purpose: the two unknowns it can meet — no
recorded digest, no spec_dir — both mean nothing was established, and nothing established is
exactly the state in which a check runs. There is no third answer for a caller to act on.
Source code in enricher/src/just_dna_enricher/clinical.py
leg_tautology_note ¶
leg_tautology_note(
sources: Sequence[SourceRow],
authority: str,
dataset: str | None,
spec_dir: Path | None,
) -> str | None
Why this authority's leg cannot fail on this module, or None when it genuinely can.
The sentence an author reads for any authority but ClinVar, which keeps its own shipped wording
in tautology_reason. The checked column is derived from DRAFT_PROJECTIONS rather than
named here, so a drafting provider added later says which cell it copied without this function
learning about it — and a source that drafts nothing at all can never reach this branch, because
it records no digest for the predicate above to match.
Source code in enricher/src/just_dna_enricher/clinical.py
tautology_reason ¶
tautology_reason(
sources: Sequence[SourceRow],
reference: Path | None,
spec_dir: Path | None = None,
) -> str | None
Why this check cannot fail on this module, or None when it genuinely can (S4, re-keyed by RM4).
A module drafted by clinvar_draft copied its clin_sig out of the snapshot this check reads,
so the comparison is a value against itself: on a 7,818-row panel a consumer measured 27.1 s with
the check on and 2.6 s with it off, byte-identical output, and 0 conflicts either way — necessarily
0. Reporting "0 conflicts" there is mild misinformation, since it looks like evidence and is not,
and the cost is real: 90% of the resolve time on a panel, ~83 s per batch on a genome-wide one.
The check itself is one of the best things in this tier wherever a human typed the value, so the default does not change; what is needed is a way for a provider-drafted module to say "this came from you".
That marker is the licence row's dataset, not an authored panel: block (RM4). The claim
being established is provenance — these rows came from this snapshot — and the tool that copied
them is the authority on it, so the enricher records it itself: clinvar_draft stamps the release
into the dataset column of the clinvar/annotation row it already had to write, and this
recomputes the same label from the snapshot in hand (clinvar_dataset_label, shared by both
sides). The alternative was asking every author to maintain a declaration whose only reader is this
one skip, which is bureaucracy the enricher exists to remove. The panel: block is deprecated in
0.6 and keys nothing here any more.
The row is matched at the annotation layer specifically: enrich() writes a second
clinvar row at the resolution layer for the coordinates it looked up, and a coordinate is not
a copied clinical call.
Three-valued in the usual way: a module with no licence table, one whose ClinVar row records no
dataset, an unreadable release.json, or a different release all return None and the check
runs. Only an established match skips it — an unknown is never a permission to skip.
The release is half the question, and RM73 supplied the other half. A matching release says
the rows were copied out of this snapshot; it says nothing about whether they still are, which
is the hole this function shipped with and named in its own message. spec_dir lets it recompute
the drafter's digest over the clin_sig column and answer that too, so the skip is now a
conjunction: same release and no checked value moved since. Either half failing runs the
check, which is why the expensive per-row strict audit this used to defer to is gone — the
digest answers the same question without a lookup, in both modes.
spec_dir is optional so an existing caller keeps working, and omitting it is not treated as a
pass: with no directory there is nothing to recompute, nothing is established, and the check
runs.
The rule is shared with every other authority's leg and the prose is not (RM134 § B). The
conjunction lives in is_tautological_leg, which the PubMind leg calls with its own label, so
two authorities cannot come to disagree about what "drafted from this very snapshot" means. This
sentence stays here, verbatim, because a warning's text is an API and this one has shipped.
Source code in enricher/src/just_dna_enricher/clinical.py
comparison_plan ¶
comparison_plan(
variants: list[VariantRow],
resolution_rows: list[ResolutionRow],
*,
authored: Callable[
[VariantRow], str | None
] = lambda variant: variant.effective_clin_sig,
) -> tuple[
list[PlannedComparison],
list[tuple[str, int, str, str]],
int,
]
(plan, alleles to ask about, rows resolved with no ALT to ask about).
Needs no snapshot and consults nothing, which is the point: every authority's leg is asked about the same alleles, and a leg that could not run does not change what a leg that did run was asked.
authored selects the cell being cross-examined and so decides which rows are subjects at all —
effective_clin_sig for the ClinVar leg, direction for the CIViC refutation leg (RM170). It is
a parameter rather than a second copy of this function because the plan is the thing both legs
must share: two authorities asked about different allele sets cannot be compared with each other,
and a private copy is how the second caller stops finding the first
(@roster-is-as-wide-as-the-tables-it-reads).
The allele list is in first-occurrence order and deduplicated, so a batch look-up and the findings that come out of it are deterministic (Principle 7). The third element is a count rather than a silence: a resolved row naming no ALT is a question nobody could put, not a question answered with nothing.
Source code in enricher/src/just_dna_enricher/clinical.py
verify_clin_sig ¶
verify_clin_sig(
variants: list[VariantRow],
resolution_rows: list[ResolutionRow],
*,
reference: Path | None,
) -> list[ClinSigConflict]
Compare each authored clin_sig against the ClinVar snapshot. Returns the disagreements.
The conflicts half of compare_clin_sig, which is where the comparison itself lives — one
implementation, so the buckets and the findings can never disagree about a row. Empty when the
snapshot could not be read at all; a caller that needs to tell that apart from "nothing
disagreed" calls compare_clin_sig and branches on its None.
Source code in enricher/src/just_dna_enricher/clinical.py
compare_clin_sig ¶
compare_clin_sig(
variants: list[VariantRow],
resolution_rows: list[ResolutionRow],
*,
reference: Path | None,
) -> ClinSigComparison | None
Compare each authored clin_sig against the ClinVar snapshot, and classify every comparison.
None — never an audit of zeros — when the comparison could not be performed at all: no snapshot
provisioned, or one that is present but not queryable. A check that could not run is not a check
that passed, and an all-zero audit would read as the latter.
Matching is on (chrom, start, ref, alt), never on rsID, for the reason lookup_clin_sig
documents. When the alt the annotation refers to cannot be determined, the comparison falls back to
the whole locus and reports only if no ClinVar record there matches the authored call —
conservative by construction, because the alternative is guessing which allele the author meant.
Variants with no clin_sig are skipped: the module makes no clinical claim, so there is nothing to
disagree with. effective_clin_sig is used rather than the raw column so a legacy row that set
only the pathogenic/benign booleans is covered too.
Source code in enricher/src/just_dna_enricher/clinical.py
424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 | |
concordance_sentences ¶
The findings a run reports about its concordance record, warning-tier, in both modes.
Empty for a run that could not put the question (record is None) and for one that found nothing
contested. A check that cannot fail reports no zero (@tautology-zero), and neither does one
that did run and found nothing — the denominator is on the record for a caller that wants it.
Every sentence carries its denominator, because "9 contested" is unreadable without the number compared, and names the authorities that actually answered: the same nine subjects mean something different when one archive spoke and when three did.
Never escalated under strict (@clinsig-never-escalates), and with more force here than
for ClinVar alone. A disagreement with a literature miner's aggregate over the field is a
statement about that extraction's limits at least as often as about the module, the corpus join
measured 62 % agreement, and discordant is a fact about the field rather than a defect in the
module. Reporting it as one would have the format arbitrate a clinical dispute.
Source code in enricher/src/just_dna_enricher/clinical.py
concordance_notes ¶
What a run should say about its concordance record that is not a finding about the module.
Info-tier, and said out loud rather than left silent: a leg that quietly did not run reads as a
leg that found nothing, which is the failure this whole tri-state exists to prevent
(@unreachable-not-absent). Three kinds of note:
- an authority nobody could ask — unchecked is not absent, and it is never agreement;
- an authority this module's own rows were drafted from, which is left unconsulted because its call would agree with the module by construction; and
- the multiplicity, where one authority holds several records for one allele. Counted rather than collapsed, and the subset of those that straddle the pathogenic/benign line is counted separately because those are the ones no single call can represent.
Source code in enricher/src/just_dna_enricher/clinical.py
answered_call_shift ¶
answered_call_shift(
spec_dir: Path,
calls: Sequence[ClinSigAuthorityCallRow] | None,
) -> AnsweredCallReport
Read this module's answers and its previous record, and compare them to calls (RM151).
The IO half of the staleness check: it names the overlay table, resolves what an answer means for that table's key, and hands two sequences to the pure comparison. It must be called before the commit, because the commit rewrites the very file it reads as the baseline.
A group-scoped answer is expanded against the record, not against the overlay. member is
genotype here, so an empty one answers every genotype of that variant_key — and the only
place the genotypes are known is the record itself. Expanding here rather than inside the
comparison keeps the table's key grammar out of a function that has no table.
The answered set is intersected with what is still on the record, and that is deliberate. An
answer whose subject has left the record is not uncomparable — it is
overlay_answer_vindicated, which the compiler already reports as good news. Counting it here
as a question nobody could put would put a second, gloomier finding on the same overlay row.
With no record to compare against (calls is None — nobody could be consulted, or the check is
off) the overlay's own pairs are the count, so the notes can still say how many answers went
unexamined. enrich() has already said why in that case.
Source code in enricher/src/just_dna_enricher/clinical.py
fold_authority_records ¶
(clin_sig, clin_sig_raw, internally_contested) for one authority's records about a subject.
One authority can hold several records for one allele — ClinVar's several submissions, PubMind's
several PVIDs over one coordinate — and the detail table is keyed
(variant_key, genotype, authority), so a subject carries one row per authority. Something
has to stand for the set, and the two ways of getting that wrong are the interesting part.
The camp guard runs first, and it is the whole safety property. When the records straddle the
pathogenic/benign line there is no representative call, and folding by severity would silently
answer pathogenic — severity ranks it above benign, so the more consequential verdict would
win a vote nobody held. That is choosing a winner by an ordering nobody defined, which this item
rejected outright. The answer is conflicting, the vocabulary's own word for exactly this, and
it sits in the undecided camp so it opposes nothing and manufactures no disagreement.
Within one camp the fold is the shared normalizer's own rule. CLIN_SIG_SEVERITY is what
resolves a composite token — Benign/Likely benign becomes likely_benign — so resolving two
records that say Benign and Likely benign the same way is one rule applied twice rather than
a second rule invented here.
clin_sig_raw keeps every distinct wording behind the fold, sorted and pipe-joined, so the
answer stays auditable and a token this release does not model is still visible.
Source code in enricher/src/just_dna_enricher/clinical.py
clin_sig_concordance ¶
clin_sig_concordance(
variants: list[VariantRow],
resolution_rows: list[ResolutionRow],
*,
reference: Path | None,
pubmind_reference: Path | None = None,
sources: Sequence[SourceRow] | None = None,
spec_dir: Path | None = None,
checked_at: str | None = None,
clinvar_comparison: ClinSigComparison | None = None,
) -> ConcordanceRecord | None
The N-authority concordance record for a module, or None when nobody could be consulted.
The record RM130 exists for and RM134 § B widened: the contested subjects, named, in a form something can join to — against the counts the check has always published and the log lines nothing kept.
The three-way check subsumes the two-way rather than running beside it. With no PubMind
snapshot that authority's calls read unchecked, which is what the tri-state is for, and the
degenerate case is exactly the finding the ClinVar-only check reported: the same subjects are
contested, the same conflicts are logged in the same words, and no author meets one disagreement
twice. What changes is only what the record withholds — authority_concordance reads
unchecked rather than single, because one authority speaking while another was never asked is
not corroboration and must not be recorded as any.
None rather than empty tables when nobody could be consulted, and the distinction is the
whole tri-state. Two empty tables are a claim — nothing here is contested — and writing them
after a comparison that never happened publishes that claim on no evidence. That is this
release's most-repeated defect, a check agreeing with its own derivation, and this check is its
likeliest home: it compares two normalizations, and a module drafted out of a snapshot it is then
compared against agrees with itself by construction. So a record is written iff at least one
authority was actually consulted, where consulted excludes:
- an authority with no snapshot, or one that is present and will not answer; and
- an authority whose values this module was drafted from and has not edited since — the
tautology, decided per leg by
is_tautological_leg, so a module drafted from ClinVar still gets a real comparison out of PubMind rather than losing the whole check to suppress half of it.
Nothing resolves a split. At five authorities in a two-against-three disagreement, precedence and
majority pick different winners and choosing needs a weighting model this workspace has declined
to invent three times, so authored_position stays a relation to the set and a consumer with
its own model computes what it likes from the detail rows.
clinvar_comparison is the comparison a caller has already run this run, reused rather than
repeated: enrich() performs it for the two-way findings, and asking the snapshot twice costs a
consumer the whole look-up again (25 s on a 7,818-row panel) for an answer already in hand. It is
ignored where the leg is tautological, which is the one case a caller holds no comparison for.
Source code in enricher/src/just_dna_enricher/clinical.py
910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 | |