just_dna_compiler.resolution¶
just_dna_compiler.resolution ¶
Pure, source-independent variant resolution from an injected resolution.csv (0.5).
The compiler's preferred resolution path. It consumes a table of already-resolved facts
(just_dna_format.resolution.ResolutionRow) keyed by the frozen variant_key, and reproduces the
DuckDB resolver's fill / expand / verify semantics without any duckdb import, SQL, or Ensembl
convention. All source knowledge (where facts come from) lives in the separate just-dna-enricher
tier; this module knows only "read the facts I was handed" — the strict inject-only end state
(CONSTITUTION Principle 2).
Digest parity with the DuckDB path is deliberate and load-bearing: given the same facts, this
produces byte-identical weights.parquet (hence artifact.digest) as resolver.resolve_variants.
The one place row order could drift — a one-to-many expansion — is pinned by sorting the expanded
rows on (locus_index, chrom, start, ref), matching the resolver's ORDER BY id, chrom, start, ref.
ResolutionOutcome
dataclass
¶
ResolutionOutcome(
variants: list[VariantRow],
plain_warnings: list[str] = list(),
ladders: list[LadderFinding] = list(),
errors: list[str] = list(),
expanded_keys: int | None = None,
expanded_rows: int | None = None,
)
What resolution produced, split by how badly each finding bites.
Three channels rather than the usual two, because resolution has three distinct severities and collapsing any pair of them loses a real distinction:
plain_warnings— reported in both modes, never fatal. Thewarningsproperty adds each ladder's own warning to them, which is what every caller reads.ladders— the round-trip contract, asLadderFindings (RM246): conditions under whichcompile → reverse → compilecannot reproduce the injected table, plusambiguous, which is reproducible but rests on a guessed label.best_effortcarries each one's warning;strict, whose contract is a reproducible artifact, refuses with its refusal — a different and longer sentence than the warning, which is why the pairing has to be carried as a pair rather than as two lists that happen to be built together.strict_errorsderives from them and is no longer a field: a field could be set without a matching warning, which is the drift the pairing prevents.errors— fatal in both modes. Onlywithdrawnlands here: every other finding leaves the annotation intact, while a retracted variant may leave it describing nothing.
expanded_keys / expanded_rows are the one-to-many expansion's two counts, carried out for
manifest.compilation (S33). Two numbers rather than one, and never a ratio, for RM44's reason:
one authored key can expand to any number of rows, and a consumer told only "3 rows are expansion
members" cannot tell three keys of one locus each from one key of three. Rows are counted in the
artifact's own unit — a weights.parquet row — so the number is checkable against the file.
They are None on the non-GRCh38 early return and nowhere else: that path resolves nothing, so
0 would say "looked, found no expansion" of a module nothing looked at. Same tri-state as every
other unknown here — withhold rather than negate.
warnings
property
¶
Everything best_effort reports: the findings that never escalate, then each ladder's own.
Both halves, because a ladder member IS a warning under best_effort — and under strict too,
where the refusal is published beside it rather than instead of it. Derived rather than stored
for the same reason strict_errors is: the pair is the unit, and a stored list could carry one
half of it.
strict_errors
property
¶
What strict refuses with — each ladder's refusal, or its warning where the two agree.
Derived rather than stored since RM246. As a field it could be populated without the matching
warning ever being emitted, which would give strict a sentence best_effort never says and
no way to notice; as a projection of the pairs, a refusal cannot exist without its warning.
PositionalFill
dataclass
¶
PositionalFill(
filled: int = 0,
unplaced_ambiguous: int = 0,
unplaced_absent: int = 0,
contradicted: list[str] = list(),
)
What the positional fill did, per table. Counts, never a line per row.
unplaced_ambiguous and unplaced_absent are separated for the reason every tri-state in this
codebase is: "the table names this key at several loci and the compiler will not pick one" and
"nothing has resolved this key" are different situations with different next moves, and one
number reporting both says neither.
resolve_from_table ¶
resolve_from_table(
variants: list[VariantRow],
resolution: dict[str, list[ResolutionRow]],
genome_build: str = "GRCh38",
) -> ResolutionOutcome
Fill/expand missing rsid or position from an injected resolution table (no network, no DuckDB).
Mirrors resolver.resolve_variants:
- fill (1:1): a variant_key with exactly one usable row fills the missing coord or rsid,
keeping the frozen key.
- expand (1:N): a variant_key (an rsid) with N usable rows expands to N coord-keyed rows,
ordered by (locus_index, chrom, start, ref) so the parquet byte order matches the DuckDB
path (digest parity). A locus whose alleles cannot host the authored genotype is not
expanded onto — see _hostable_loci.
- verify: a row carrying both an rsid and a coordinate is checked against the table; a
disagreement warns in best_effort and refuses in strict.
GRCh38-bound, like the DuckDB resolver (RM15): a non-GRCh38 module is skipped with a warning, and a
resolution row whose genome_build differs from the module's is ignored.
Returns a ResolutionOutcome — see its docstring for the three severity channels, and
COMPILER.md § Resolution for the full matrix.
Source code in compiler/src/just_dna_compiler/resolution.py
108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 | |
resolve_positional_rows ¶
resolve_positional_rows(
rows: list[object],
resolution: dict[str, list[ResolutionRow]],
genome_build: str = "GRCh38",
) -> PositionalFill
Fill the resolved coordinate into one 0.4-family positional table, in place (RM43).
pharm_variants.csv, haplotypes.csv and heteroplasmy.csv may identify a row by rsID alone —
their own models say so — and until 0.6 the compiler resolved variants.csv and materialized
every other table verbatim. A consumer matching a patient VCF by position therefore matched
nothing, silently, as an empty result rather than an error: the coordinates existed in the
resolution.csv sitting beside the spec and reached the artifact by no path at all.
Four rules, and the last two are what keep this from being the naive repair the roadmap rejected:
- Fill only what the author left empty. A cell the author wrote is never overwritten; that is
enrich's inject-only doctrine (report, never repair) and it is also what makes the fill idempotent. - Fill from exactly one locus, or from none. One usable locus fills. Several are filtered by
hosting_verdictagainst whatever allele the row states — agenotypeon a pharm row, the definingalleleon a haplotype junction, and nothing at all on a heteroplasmy band, which is a measurement over a locus rather than a claim about a genotype. If that leaves one, it fills; otherwise the row stays unplaced and is counted. There is deliberately no expansion: a one-to-many rsID expandsvariants.csvinto N coord-keyed rows, and doing the same here would multiply a pharm annotation's(variant_key, drug, genotype, …)key across loci the author never named. - A row whose own coordinate contradicts the table is left alone.
haplotypes.csvdrafted from CPIC carries astartwith nochrom, so the fill has to complete a half-coordinate — and completing it from a locus whosestartdisagrees would build a coordinate no source ever stated. Reported, never repaired, and never fatal: the same shape asresolve_from_table's_verify, minus the strict escalation, because the row is left exactly as authored. - The row is mutated, not copied, and its identity is frozen.
variant_key/authored_identare stamped at load from the authored subset, so filling cannot re-key the row andreverse_modulere-emits the shape the author wrote (_write_table_csv).
GRCh38-bound for the same reason resolve_from_table is (RM15): the caller skips a non-GRCh38
module rather than joining rows this tier cannot re-derive an identity for.
Source code in compiler/src/just_dna_compiler/resolution.py
withdrawn_refusals ¶
withdrawn_refusals(
variants: list[VariantRow],
resolution: dict[str, list[ResolutionRow]],
genome_build: str,
) -> list[str]
The one resolution refusal that is fatal in both modes, as finished sentences (RM207).
A merged or absent rsID leaves the annotation intact — the module is dated, or the label is
unserved. A withdrawn one is dbSNP repudiating the variant, so the annotation may be describing
something that does not exist; carrying it under best_effort would be publishing a claim its own
source has retracted. Never produced by the automated check (a retraction is indistinguishable
from a never-assigned id through the live API), so it fires only where a curator recorded it.
Sentences rather than names, which is where this differs from unresolved_subjects beside it.
That one returns subjects and lets each caller phrase them, because both phrasings are one short
clause. This message names two things — the subject and the retracted rsID — and the standing
rule is to share the predicate and copy the error; copying a sentence with two interpolations into
a second caller is how the two drift, so the sentence is shared too and there is one of it.
Empty for a non-GRCh38 module, where resolve_from_table skips resolving wholesale.
Source code in compiler/src/just_dna_compiler/resolution.py
ambiguous_refusals ¶
ambiguous_refusals(
variants: list[VariantRow],
resolution: dict[str, list[ResolutionRow]],
genome_build: str,
) -> list[str]
strict-only: the table marks this rsid ambiguous, so the pick is deterministic, not a fact.
Source code in compiler/src/just_dna_compiler/resolution.py
ambiguous_warnings ¶
ambiguous_warnings(
variants: list[VariantRow],
resolution: dict[str, list[ResolutionRow]],
genome_build: str,
) -> list[str]
The best_effort half of the same finding — the pick is carried, and said to be a pick.
Source code in compiler/src/just_dna_compiler/resolution.py
unresolved_subjects ¶
unresolved_subjects(
variants: list[VariantRow],
resolution: dict[str, list[ResolutionRow]],
genome_build: str,
) -> list[str]
Authored rows the injected table cannot place, named by rsID (or variant_key).
The predicate resolve_from_table will apply, factored out so the pre-flight can ask it before
resolution runs and reach the same answer the compile does (S76). A row is unplaceable when it
authors no coordinate of its own and the table holds no usable locus for its key — the same
_usable_loci filter the fill uses, so a not_found sentinel and a row recorded under another
build count as no locus here exactly as they do there.
Deliberately not a re-derivation of the compile's check: compile_module runs its refusal over
outcome.variants, the post-expansion list, and the two agree because a subject with no usable
locus is carried through the expansion unchanged. Sharing the predicate is what keeps that true —
a second implementation beside the first is the drift this returns instead of.
Empty for a non-GRCh38 module, where resolve_from_table skips wholesale rather than resolving
less: reporting every row unplaceable there would restate the skip warning once per row and blame
the table for a limit of the tier.
Source code in compiler/src/just_dna_compiler/resolution.py
hosting_verdict ¶
Can a locus spelling {ref} ∪ alts host genotype? Three-valued (RM31).
True it can, False it positively cannot, None this tier cannot tell — the house algebra, and
the third value is the whole point: an indel has several valid spellings, so a string comparison
reporting "does not fit" was asserting a verdict it had not reached.
The ladder, in order, and the order is load-bearing:
- No
ref/altsrecorded →True. Nothing is known about the locus's alleles and rejecting for lack of evidence is worse than accepting. Unchanged. *abstains — on both sides — and the rest is still judged (RM59).*says this sample's allele could not be observed here; it names no allele, so a locus cannot fail to offer it and cannot contradict it, and it belongs in the allele algebra on neither side. Without this the widening that made*/Gauthorable would be self-defeating: no source spells*in an ALT list, so an rsid-authored*/Gresolved against aG>Alocus fell through to step 6 and came back a confidentFalse— the locus dropped from the expansion, the row unresolved, andstrictrefusing over a finding no authored edit could clear. Dropping the member rather than the whole verdict is what keeps the observable half sharp:*/Tat aG>Alocus is stillFalse, because theTreally is a contradiction. Nothing observable on either side isNone.
The locus side is not optional, and the first cut got this wrong on the reasoning that a *
in ref/alts is reachable today, so excluding it might move an existing verdict. It does move
existing verdicts, and it is still right: parsimony_reduce cannot strip a shared flank past a
member that has none, so a * left in the locus stops the set reducing and collapses RM31 —
hosting_verdict('C/CAG', 'AGAG', 'AG') is True while …('C/CAG', 'AGAG', 'AG,*') was
False, refusing a correctly transcribed indel under --strict and telling the author to
"replace it with the alleles the locus actually has". ALT=AG,* is ordinary joint-caller output.
What that costs is measured, and it is not only acceptances. Swept over every star-free
genotype — i.e. everything authorable before RM59 — against loci with and without a *: for a
locus spelled in nucleotides nothing changes at all, and for a locus spelling * verdicts
move in both directions, None → False among them. Those are corrections. * was doing two
things to the arithmetic at once, blocking the flank strip and lending its own single character
to _indel_shaped's length set, so ref=AGAG alts=AG,* stayed unreduced, read as indel-shaped,
and withheld on calls that were decidable all along — a genotype naming an allele the locus does
not have escaped the finding. Same bug class as the RM5 guard's stated reason: characters counted
out of a token that spells no sequence. No reference example carries a *, so the corpus could
not have shown any of this.
3. The raw strings match → True. Checked before any normalization, so this step can only
ever gain acceptances: whatever passed the raw comparison before still passes it, byte for byte.
Step 2 sits above it and costs it nothing, provably rather than nearly — the called side is
stripped first, so * is never on the left of the subset test, and dropping it from the right
therefore cannot change the answer. The guarantee stops here: steps 5–8 are arithmetic over
the reduced sets, and stripping the locus does change those, in both directions (see step 2).
4. Either side names a symbolic allele → None. A <DEL:1500> has no sequence, so it has no
flank for parsimony_reduce and nothing to compare character by character — and a symbolic
allele exists because the exact sequence is not known, so even two stated lengths that differ
are summary information rather than grounds for a contradiction. Undecided, never "no match"
(RM5). Below the raw comparison, so <DEL:1500> against a locus spelling <DEL:1500> still
matches exactly.
5. The reduced allele sets match → True. alleles.parsimony_reduce strips the flank each
collection shares, leaving the event; ClinVar's C/CAG and Ensembl's AGAG>AG both reduce to
{'', 'AG'}, the SHOX 2 bp deletion that used to resolve to not_found.
6. The locus is a substitution or MNV → False. No flank, so no spelling freedom: an A/G
genotype at a C>T locus is a real contradiction, and must stay one (a strand flip is exactly
what that check catches).
7. The genotype names fewer than two distinct alleles → None. A homozygous C/C carries no
frame — one string has nothing to be relative to — so against an indel locus there is genuinely
nothing to compare. Reported as undecided, never as a contradiction. A */C reaches this too
once the * has abstained, and for the same reason: one string has nothing to be relative to.
8. An event length the locus does not offer → False. The confident negative:
left-alignment moves an indel, it never changes how many bases the event adds or removes, so a
1 bp insertion cannot be a 2 bp deletion however it is spelled.
9. Otherwise → None. Same lengths, different content: one variant rotated inside a repeat, or
two different variants, and only the reference sequence can say which. The enricher can settle
it (seqrepo); the compiler holds no reference by charter (P2), so it withholds.
Public because both resolvers must agree on it: this module's injected-table path and the
deprecated DuckDB path in just-dna-enricher. Digest parity between the two is a documented
guarantee, so a filter applied on one side only would silently break it.
Source code in compiler/src/just_dna_compiler/resolution.py
642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 | |
undecided_reason ¶
Why hosting_verdict withheld — the clause a caller appends when the answer was None.
None has five causes and the message asserted one of them. Both reporting sites — this
module's expansion warning and the enricher's twin — said "the two spellings describe events of
the same size but different content, either one indel re-anchored inside a repeat or two different
variants", which is step 9's cause and false for the other four: a symbolic allele was never
compared at all (RM5), an all-* call observed nothing and an all-* locus offers nothing (RM59),
and a homozygous call carries no frame. Stating a cause the tier did not establish is the same
defect as the ([], None) collapse S20 was filed for — two ways of returning nothing rendered as
one sentence — and here it sends the reader to check a reference sequence for a row where no
reference could help.
Mirrors hosting_verdict's withholding branches in its order, so the two answer the same
question the same way; a test walks every None-producing shape and asserts the pairing, which is
what keeps a sixth branch from quietly inheriting a fifth's explanation.
The sets are built exactly as the predicate builds them — including dropping an empty ref rather
than folding "" in, which would put a member in locus that is neither observable nor an allele
and quietly defeat the branch below it. Keeping the two constructions identical is the entire point
of the pair.
Source code in compiler/src/just_dna_compiler/resolution.py
contradiction_reason ¶
Why hosting_verdict was a confident False — undecided_reason's twin, same contract (S85).
False has two causes and the enricher's message asserted one of them for both. Step 8 is the
event-length arm — a 1 bp insertion cannot be a 2 bp deletion however it is spelled — and "the
event sizes differ, which re-anchoring cannot change" is its reason. Step 6 is the substitution/MNV
arm, where the sizes are identical and what makes the verdict confident is the absence of a flank:
same-length alleles have no spelling freedom, so a genotype naming alleles the locus does not offer
is a real contradiction. Saying "the event sizes differ" there is a false claim about two 1 bp
substitutions, and it sends the author hunting a second variant sharing the rsID.
The case that made this worth splitting is a strand flip: a paper's supplementary published against
hg19 spells the submitted strand, so an authored A/G meets GRCh38's C>T. The verdict is correct
— on the strand it is written on, that genotype really cannot be hosted — and the remedy is a
column the old sentence never mentioned. So the flip is named where it is established, and the
step-6 fallback stays honest about what it did not establish rather than inventing a cause.
Mirrors hosting_verdict's False branches in its order, for the reason its twin does: a test
walks every False-producing shape and asserts the pairing, so a third arm cannot silently inherit
a second's explanation.
Source code in compiler/src/just_dna_compiler/resolution.py
genotype_fits ¶
Whether a locus can host genotype, collapsing "cannot tell" into "keep it".
The boolean face of hosting_verdict, kept because both resolvers and three call sites read it, and
because the collapse it performs is the module's existing doctrine: only a positive contradiction
rejects. An undecidable spelling is therefore kept, exactly as a locus with no recorded alleles is.
A caller that needs to report the difference asks hosting_verdict instead.
Source code in compiler/src/just_dna_compiler/resolution.py
spelling_caveat ¶
The clause to append when a "cannot host" verdict is really about allele spelling.
Empty for a locus spelled in nucleotides, which is nearly all of them, so a caller can append it unconditionally.
hosting_verdict reaches False for a substitution locus by set difference, and it is right to: a
substitution has no shared flank, so no spelling freedom, so A/G at a C>T locus is a real
contradiction and must stay one — that is what makes the strand-flip check sharp. But it reaches the
same False when the locus is spelled T>Y, and there the mismatch is between the cell and the
nucleotide alphabet, not between the genotype and the variant. Reporting the generic message there
sends the author to re-examine a genotype that was correct all along.
Three reasons reach this caveat, kept apart because what the author does next differs: an ambiguity
code is an uncertainty that can never be expanded into definite alleles (expanding N to A,C,G,T
asserts four alleles nobody stated), a well-formed symbolic allele is held by the grammar since
RM5 and so is never a spelling defect at all, and everything else non-nucleotide is a grammar gap.
None is repaired here — this tier reports.
* is the fourth classification and is excluded here (RM59), because hosting_verdict strips
it from both sides before comparing, so it can no longer contribute to the False this sentence
explains. Left in, it inverted the caveat's whole point: genotype=C/G at ref=A alts="T,*" is a
real genotype error, and the author was told to read it as a spelling problem instead.