just_dna_enricher.concordance¶
just_dna_enricher.concordance ¶
Turning a clinical-significance comparison into a record an author can act on (0.7, RM130).
The ClinVar cross-check has always known which rows disagree and has always thrown them away: it reports twenty of 141,616 and the twenty reach a logger. This module is the other half — the two verdicts, the per-authority calls behind them, and the writer that puts both beside the spec.
Two functions, and the split is the design. classify_concordance is a pure function over an
authored call and N authority outcomes; concordance_tables turns a run's findings into rows. The
classifier knows nothing about ClinVar, PubMind or any other source, which is what lets it be
exercised at three and five authorities today rather than at whatever N the current producer happens
to reach — and the arity of the vocabularies is the property the whole two-field split exists to
hold.
Nothing here resolves a split. With five authorities in a two-against-three disagreement,
precedence and majority pick different winners and choosing between them needs a weighting model.
This workspace has declined to invent one three times, so authored_position is a relation to the
set: computable with no weights, true at any topology, and leaving a consumer free to apply its own
model to the detail rows.
What contests a subject, stated once because it is the contract with the shipped check. A row is written when a disagreement is established:
- the module's own call and an authority's call sit in opposite camps — a pathogenic-class call against a benign-class one; or
- two authorities that spoke disagree with each other.
At one authority the second clause cannot fire, so the emitted set is exactly the conflicts the shipped check already reports — which is what keeps the three-way check a superset of the two-way rather than a second opinion beside it, and keeps an author from meeting one disagreement twice.
An authority that could not be consulted never contests anything on its own. unknown AND false
is false under Kleene, so an unreachable archive does not un-see a disagreement already witnessed;
but on its own it produces no row, because a question nobody could put is not a finding.
AuthorityCall
dataclass
¶
AuthorityCall(
authority: str,
status: str,
clin_sig: str | None = None,
clin_sig_raw: str | None = None,
confidence: str | None = None,
confidence_unit: str | None = None,
dataset: str | None = None,
)
What one annotation authority had to say about one subject.
The classifier's input unit, and deliberately not tied to any source: it carries a normalized
clin_sig, the raw token behind it, and a confidence in the authority's own units with the
name of the instrument beside it. Nothing converts one authority's confidence into another's —
a gold-star count and an evidence-depth count are not the same quantity, and folding them into
one number would put three axes in one field.
ConcordanceVerdict
dataclass
¶
ConcordanceVerdict(
authority_concordance: str,
authored_position: str,
opposed: bool | None,
contested: bool,
)
The two orthogonal verdicts, plus what follows from them.
contested is on the verdict rather than recomputed by the caller because it depends on facts
only the classifier holds — whether any authority actually spoke, and whether the camps in play
are opinionated. A caller re-deriving it from the two vocabulary members alone would call a
subject contested where nobody has a record, which is the vacuous-truth trap in the ∀ over an
empty set.
ConcordanceSubject
dataclass
¶
ConcordanceSubject(
variant_key: str,
genotype: str,
authored_clin_sig: str | None,
calls: tuple[AuthorityCall, ...],
)
One (variant_key, genotype) the check put a question about, and every answer it got.
CallShift
dataclass
¶
One authority's call about one answered subject, as it was recorded and as it reads now.
Both sides are whole ClinSigAuthorityCallRows rather than the two classifications, because the
dataset pair is what separates the two readings a reader actually needs: a re-released
archive that revised its call, and an archive that revised in place within one release. The
check does not branch on that difference — it has no rule for a dataset neither side recorded —
it just puts both in the message and lets the reader see which happened.
subject
property
¶
The (variant_key, genotype) the answer was written about.
wording_held
property
¶
Whether the authority's own words are unchanged while our reading of them moved.
True only when both sides recorded a verbatim token and the two are equal. That is this release's normalizer having moved, not the archive — a different finding wearing the same shape, and reporting it as an archive revision would accuse a source of a change we made.
AnsweredCallReport
dataclass
¶
AnsweredCallReport(
shifts: tuple[CallShift, ...],
normalization: tuple[CallShift, ...],
withheld: tuple[tuple[str, str, str, str], ...],
subjects: int,
answered: int,
)
What a run can say about the answers this module already carries.
Nothing here is computed and discarded (@dont-discard-computed): answered and subjects
are the two denominators, and a caller that recomputed either would be re-implementing the
answered-subject rule — the drift the overlay's own design refuses.
moved_subjects
property
¶
Answered subjects carrying at least one moved call. The numerator.
camp_of ¶
The camp a classification falls in, or None when there is no classification to place.
None rather than undecided for an absent value, because the two are different statements: a
source that said nothing has not landed in the undecided camp, it has not spoken at all. A
classification this release does not model reads as undecided — it was stated, and it opposes
nothing.
Source code in enricher/src/just_dna_enricher/concordance.py
classify_concordance ¶
classify_concordance(
authored_clin_sig: str | None,
calls: Sequence[AuthorityCall],
) -> ConcordanceVerdict
The two verdicts for one subject, at any number of authorities.
authority_concordance — do the authorities agree with each other?
A disagreement already witnessed is not un-witnessed by an authority that could not be asked
(unknown AND false is false), so discordant wins over unchecked. Otherwise an unreachable
authority leaves the question open, and the answer is unchecked rather than an agreement nobody
established. With everyone reachable it is concordant for two or more agreeing, single for
exactly one opinion — one voice is not corroboration — and none when every archive was asked
and none has a record.
authored_position — where does the module's own call sit relative to them?
matches_some is establishable under an unreachable sibling, because both halves of it are
witnessed by authorities that did speak. matches_all and matches_none are ∀ claims and need
the set closed, so an unreachable authority sends them to unchecked. absent is the module
stating no clinical call at all, which is a different thing from disagreeing with everyone.
Both are at camp granularity, the same coarseness the two-way check uses: pathogenic and
likely_pathogenic are one position, not two.
Source code in enricher/src/just_dna_enricher/concordance.py
154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | |
concordance_tables ¶
concordance_tables(
subjects: Sequence[ConcordanceSubject],
*,
checked_at: str | None = None,
) -> tuple[
list[ClinSigConcordanceRow],
list[ClinSigAuthorityCallRow],
]
The two tables for a run: contested subjects, and the calls behind each one.
Only contested subjects reach the record, and the detail rows follow the parent rather than
being written for every subject asked. A record of every agreement would be a copy of the
module's own clin_sig column with a second opinion attached, and the count of subjects
compared is already published as the check's denominator — recording it a second time as rows
would be the same number in two places.
Order is the subject order the caller supplied, and within a subject the call order it supplied, both preserved rather than sorted: parquet bytes depend on row order, and a sort over values would let a corrected classification move a row.
Source code in enricher/src/just_dna_enricher/concordance.py
write_concordance_tables ¶
write_concordance_tables(
spec_dir: Path,
parents: Sequence[ClinSigConcordanceRow],
calls: Sequence[ClinSigAuthorityCallRow],
) -> tuple[Path, Path]
Replace both tables beside the spec, and return where they were written.
These are pure build products, and they are rewritten whole rather than merged. Every other
derived sidecar gap-fills because a recorded row might carry a curator's judgement; this one
cannot, because the judgement about a contested subject goes in overrides.csv and never into
this file. Merging would be actively wrong: a subject the archive stopped contesting has to
leave the record, since a conflict that stops being reported is exactly how an author learns
the archive caught up with them.
A run that found nothing contested therefore writes two empty tables rather than leaving the previous run's rows in place — an empty record is a claim (nothing is contested), and it is a different claim from no record at all, which is what a run that could not put the question leaves behind by writing nothing.
Resolved through layout so a module carrying its sidecars under derived/ gets them written
where it keeps the rest, rather than growing a second copy at the root.
Source code in enricher/src/just_dna_enricher/concordance.py
read_recorded_calls ¶
The authority calls this module already has on file, or None when it has none (RM151).
The baseline half of the staleness comparison, and it must be read before the commit rewrites
it. write_concordance_tables replaces this table whole, so the previous run's rows are the
record as it stood when an author read it and wrote their answer — available for exactly as long
as this run has not yet committed. enrich() computes everything before its commit already, for
the unrelated reason that a refused strict run must change nothing, and that ordering is what
makes this readable at all.
None rather than [] for a module that has never had a record written, because the two are
different claims and the whole check turns on the difference: no file is nobody asked at record
time, an empty file is a run that asked and found nothing contested.
A file that will not parse is None too, and says so. Guessing at a half-read baseline would
manufacture a movement out of our own inability to read, which is the one wrong answer available.
Resolved through layout so a module keeping its sidecars under derived/ is read where it
keeps them (@sidecar-name-and-place).
Source code in enricher/src/just_dna_enricher/concordance.py
shifted_authority_calls ¶
shifted_authority_calls(
baseline: Sequence[ClinSigAuthorityCallRow] | None,
fresh: Sequence[ClinSigAuthorityCallRow] | None,
answered: Sequence[tuple[str, str]],
) -> AnsweredCallReport
The answered subjects whose authority calls have moved since the answer was written (RM151).
An overrides.csv row against the concordance record is a judgement about a particular
disagreement — the archive said X, the author says Y, and the reason column explains why. If
the archive later says Z, that reason was written about a value that is no longer there. Nothing
in the record distinguishes a justification that still describes the disagreement on file from
one that describes a disagreement since replaced by a different one, and this is the check that
notices.
The comparison is recorded-call against fresh-call, per (variant_key, genotype, authority),
and the previous clin_sig_authority_calls.csv is the only baseline this format has. That
table is what each authority actually said, with clin_sig, the verbatim clin_sig_raw and the
dataset it came from — so the comparison is available here and nowhere else. Do not promise
this for a table that records no prior value (@probe-names-the-table): an overlay row against
frequencies.csv or resolution.csv has no recorded baseline at all, so a general "the value
moved" check would be answerable for one table and silently absent for the rest.
It observes, it does not adjudicate. The disagreement you answered is not the disagreement that exists now is a statement about the record. Your answer may be wrong is a verdict, and this format does not put a verdict under a check that cannot see the reasoning — the same restraint the concordance tables keep everywhere else.
A subject that left the record entirely is not this finding. It is not in fresh, so it
never enters the loop: the authorities stopped contesting it, which is RM117's
overlay_answer_vindicated and is reported by the compiler as good news. Two findings about one
overlay row would be one of them firing with the wrong words on it.
Order is fresh row order, preserved rather than sorted: the message is read by a human and a
sort over values would let a corrected classification move a line.
Source code in enricher/src/just_dna_enricher/concordance.py
496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 | |
answered_call_sentences ¶
The findings a run reports about its answered subjects, warning-tier, in both modes (RM151).
Never escalated under strict (@clinsig-never-escalates), and with more force than for the
concordance record itself: nothing here is even a disagreement, it is a note that the record an
author reasoned over has been rewritten underneath their reasoning. Refusing a build over it
would gate an artifact on an archive's release schedule.
Two sentences, kept apart because they are about different parties. The first is the archive revising; the second is this release's own normalizer reading unchanged wording differently, which is a fact about our code and is not the author's to act on.
Empty when nothing moved and empty when nothing could be compared — a check that cannot fail
reports no zero (@tautology-zero), and the gap belongs in the notes rather than dressed up as a
clean bill.
Source code in enricher/src/just_dna_enricher/concordance.py
answered_call_notes ¶
What a run should say about answers it could not put the question for. Info-tier.
Said out loud rather than left silent, because a comparison that quietly did not run reads as one
that found nothing — the failure the whole tri-state exists to prevent
(@unreachable-not-absent). Grouped by reason rather than by row, like every repeated finding
here.
Silent for a module with no answers at all, which is every module today: a check that cannot fire must not announce a zero.