Skip to content

just_dna_compiler.resolution

just_dna_compiler.resolution

Pure, source-independent variant resolution from an injected resolution.csv (0.5).

The compiler's preferred resolution path. It consumes a table of already-resolved facts (just_dna_format.resolution.ResolutionRow) keyed by the frozen variant_key, and reproduces the DuckDB resolver's fill / expand / verify semantics without any duckdb import, SQL, or Ensembl convention. All source knowledge (where facts come from) lives in the separate just-dna-enricher tier; this module knows only "read the facts I was handed" — the strict inject-only end state (CONSTITUTION Principle 2).

Digest parity with the DuckDB path is deliberate and load-bearing: given the same facts, this produces byte-identical weights.parquet (hence artifact.digest) as resolver.resolve_variants. The one place row order could drift — a one-to-many expansion — is pinned by sorting the expanded rows on (locus_index, chrom, start, ref), matching the resolver's ORDER BY id, chrom, start, ref.

ResolutionOutcome dataclass

ResolutionOutcome(
    variants: list[VariantRow],
    plain_warnings: list[str] = list(),
    ladders: list[LadderFinding] = list(),
    errors: list[str] = list(),
    expanded_keys: int | None = None,
    expanded_rows: int | None = None,
)

What resolution produced, split by how badly each finding bites.

Three channels rather than the usual two, because resolution has three distinct severities and collapsing any pair of them loses a real distinction:

  • plain_warnings — reported in both modes, never fatal. The warnings property adds each ladder's own warning to them, which is what every caller reads.
  • ladders — the round-trip contract, as LadderFindings (RM246): conditions under which compile → reverse → compile cannot reproduce the injected table, plus ambiguous, which is reproducible but rests on a guessed label. best_effort carries each one's warning; strict, whose contract is a reproducible artifact, refuses with its refusal — a different and longer sentence than the warning, which is why the pairing has to be carried as a pair rather than as two lists that happen to be built together. strict_errors derives from them and is no longer a field: a field could be set without a matching warning, which is the drift the pairing prevents.
  • errors — fatal in both modes. Only withdrawn lands here: every other finding leaves the annotation intact, while a retracted variant may leave it describing nothing.

expanded_keys / expanded_rows are the one-to-many expansion's two counts, carried out for manifest.compilation (S33). Two numbers rather than one, and never a ratio, for RM44's reason: one authored key can expand to any number of rows, and a consumer told only "3 rows are expansion members" cannot tell three keys of one locus each from one key of three. Rows are counted in the artifact's own unit — a weights.parquet row — so the number is checkable against the file.

They are None on the non-GRCh38 early return and nowhere else: that path resolves nothing, so 0 would say "looked, found no expansion" of a module nothing looked at. Same tri-state as every other unknown here — withhold rather than negate.

warnings property

warnings: list[str]

Everything best_effort reports: the findings that never escalate, then each ladder's own.

Both halves, because a ladder member IS a warning under best_effort — and under strict too, where the refusal is published beside it rather than instead of it. Derived rather than stored for the same reason strict_errors is: the pair is the unit, and a stored list could carry one half of it.

strict_errors property

strict_errors: list[str]

What strict refuses with — each ladder's refusal, or its warning where the two agree.

Derived rather than stored since RM246. As a field it could be populated without the matching warning ever being emitted, which would give strict a sentence best_effort never says and no way to notice; as a projection of the pairs, a refusal cannot exist without its warning.

PositionalFill dataclass

PositionalFill(
    filled: int = 0,
    unplaced_ambiguous: int = 0,
    unplaced_absent: int = 0,
    contradicted: list[str] = list(),
)

What the positional fill did, per table. Counts, never a line per row.

unplaced_ambiguous and unplaced_absent are separated for the reason every tri-state in this codebase is: "the table names this key at several loci and the compiler will not pick one" and "nothing has resolved this key" are different situations with different next moves, and one number reporting both says neither.

resolve_from_table

resolve_from_table(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str = "GRCh38",
) -> ResolutionOutcome

Fill/expand missing rsid or position from an injected resolution table (no network, no DuckDB).

Mirrors resolver.resolve_variants: - fill (1:1): a variant_key with exactly one usable row fills the missing coord or rsid, keeping the frozen key. - expand (1:N): a variant_key (an rsid) with N usable rows expands to N coord-keyed rows, ordered by (locus_index, chrom, start, ref) so the parquet byte order matches the DuckDB path (digest parity). A locus whose alleles cannot host the authored genotype is not expanded onto — see _hostable_loci. - verify: a row carrying both an rsid and a coordinate is checked against the table; a disagreement warns in best_effort and refuses in strict.

GRCh38-bound, like the DuckDB resolver (RM15): a non-GRCh38 module is skipped with a warning, and a resolution row whose genome_build differs from the module's is ignored.

Returns a ResolutionOutcome — see its docstring for the three severity channels, and COMPILER.md § Resolution for the full matrix.

Source code in compiler/src/just_dna_compiler/resolution.py
def resolve_from_table(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str = "GRCh38",
) -> ResolutionOutcome:
    """Fill/expand missing rsid or position from an injected resolution table (no network, no DuckDB).

    Mirrors `resolver.resolve_variants`:
      - **fill (1:1):** a `variant_key` with exactly one usable row fills the missing coord or rsid,
        keeping the frozen key.
      - **expand (1:N):** a `variant_key` (an rsid) with N usable rows expands to N coord-keyed rows,
        ordered by `(locus_index, chrom, start, ref)` so the parquet byte order matches the DuckDB
        path (digest parity). A locus whose alleles cannot host the authored genotype is **not**
        expanded onto — see `_hostable_loci`.
      - **verify:** a row carrying both an rsid and a coordinate is checked against the table; a
        disagreement warns in `best_effort` and refuses in `strict`.

    GRCh38-bound, like the DuckDB resolver (RM15): a non-GRCh38 module is skipped with a warning, and a
    resolution row whose `genome_build` differs from the module's is ignored.

    Returns a `ResolutionOutcome` — see its docstring for the three severity channels, and
    COMPILER.md § Resolution for the full matrix.
    """
    if genome_build != "GRCh38":
        msg = skipped_cross_build(what="Resolution-table fill", genome_build=genome_build)
        logger.warning(msg)
        return ResolutionOutcome(
            variants=variants,
            plain_warnings=[CodedWarning("resolution_skipped_cross_build", msg)],
        )

    warnings: list[str] = []
    ladders: list[LadderFinding] = []
    errors: list[str] = []
    patched: list[VariantRow] = []
    # Coordinate-authored rows the table knows no rsID for. Collected rather than reported per row:
    # `reference_examples/pathogenic_clinvar/` produced 26 of these, which buried the nine expansion
    # warnings and the duplicate-citation finding in the same run. It is also the *expected* state for a
    # coordinate-authored module — the row already has its identity, and an rsID is a convenience label
    # — so it earns one counted line, not a line each.
    no_rsid: list[str] = []
    # rsID → the usable loci found for each authored row carrying it, one entry per authored row. So
    # `len(entry)` is how many rows the author wrote at that key and `sum(map(len, entry))` is how
    # many artifact rows they became. Collected rather than reported in place — see the expansion
    # branch below.
    expansions: dict[str, list[list[ResolutionRow]]] = {}
    for v in variants:
        rows = resolution.get(v.variant_key or "")
        loci = _usable_loci(rows, genome_build)

        if v.rsid is not None and v.chrom is None:
            # need position: fill from the table, or expand a one-to-many rsid
            if not loci:
                warnings.append(
                    CodedWarning(
                        "rsid_unresolved",
                        unresolved_rsid(
                            v.rsid,
                            searched="the resolution table",
                            consequence="position remains unset",
                        ),
                    )
                )
                patched.append(v)
            elif len(loci) == 1:
                patched.append(v.model_copy(update=_coord_update(loci[0])))
            else:
                usable, rejected, undecided = _hostable_loci(loci, v.genotype)
                for locus in undecided:
                    # Kept, and said out loud. This tier cannot re-anchor an indel (that needs the
                    # reference sequence, which P2 keeps out of the compiler), so the row is carried and
                    # the reader is told the comparison did not reach a verdict — never that the locus is
                    # a different variant, which is what the old message asserted. `undecided_reason`
                    # supplies *which* of the four ways it withheld, rather than naming one of them for
                    # all four.
                    warnings.append(
                        CodedWarning(
                            "locus_hosting_undecidable",
                            f"{v.rsid}: whether {locus.chrom}:{locus.start} {locus.ref}>{locus.alts} can "
                            f"host the authored genotype {v.genotype} could not be decided here — "
                            f"{undecided_reason(v.genotype, locus.ref, locus.alts)}. The locus is kept.",
                        )
                    )
                for locus in rejected:
                    # Dropping a locus makes the emitted table smaller than the injected one, so the
                    # round-trip cannot reproduce it — strict must refuse rather than silently prune.
                    caveat = spelling_caveat(locus.ref, locus.alts)
                    refusal = (
                        f"{v.rsid}: locus {locus.chrom}:{locus.start} {locus.ref}>{locus.alts} "
                        f"cannot host the authored genotype {v.genotype}. Dropping it makes the "
                        f"compile non-reproducible from the injected table; fix the genotype or the "
                        f"table, or compile without strict.{caveat}"
                    )
                    ladders.append(
                        LadderFinding(
                            CodedWarning(
                                "locus_cannot_host_genotype",
                                locus_cannot_host(
                                    v.rsid,
                                    locus=f"{locus.chrom}:{locus.start}",
                                    ref=locus.ref,
                                    alts=locus.alts,
                                    genotype=v.genotype,
                                    instead=(
                                        "rather than emitted as a row asserting an allele it does not have"
                                    ),
                                    caveat=caveat,
                                ),
                            ),
                            refusal=refusal,
                        )
                    )
                if not usable:
                    # Every candidate contradicts the genotype: the rsid and the genotype cannot both
                    # be right. Leave the row unresolved rather than pick one — `_cross_validate` and
                    # the strict gate then treat it as the unresolved variant it is.
                    warnings.append(
                        CodedWarning(
                            "rsid_no_hosting_locus",
                            no_hosting_locus(v.rsid, loci=len(loci), genotype=v.genotype),
                        )
                    )
                    patched.append(v)
                elif len(usable) == 1:
                    patched.append(v.model_copy(update=_coord_update(usable[0])))
                else:
                    # Accumulated per rsID, reported once after the loop. Emitting here put the same
                    # sentence in `manifest.compilation.warnings` once per *authored row* — a site with
                    # two authored genotypes published it twice — and each copy said "expanded to 2
                    # rows" while the artifact gained four. Same rule as the counted line below: a
                    # message that embeds a count is owed the whole denominator, and the whole
                    # denominator is not known until every authored row at this key has been judged
                    # (which loci are usable depends on the genotype, so two rows at one key can
                    # legitimately expand onto different loci).
                    expansions.setdefault(v.rsid, []).append(usable)
                    for index, locus in enumerate(_sorted_loci(usable)):
                        update = _coord_update(locus)
                        # `build=` is redundant *today* — the function returns early for any other
                        # build 70 lines up — and is passed anyway: correct-by-construction beats
                        # correct-by-a-distant-guard, and every instance of this bug so far was a
                        # guard that existed somewhere else.
                        update["variant_key"] = derive_variant_key(
                            None,
                            locus.chrom,
                            locus.start,
                            locus.ref,
                            locus.alts,
                            build=genome_build,
                        )
                        # The expansion marker (RM87). This is the only site that knows a row is a
                        # member of anything — after here it is an ordinary well-formed row with a
                        # real coordinate, and nothing on it said so. Both numbers were already in
                        # scope: the ordinal is the `enumerate`, the total is `len(usable)`.
                        #
                        # The ordinal counts within `usable`, i.e. after `_hostable_loci` dropped any
                        # locus that positively contradicts the genotype — not within the injected
                        # table's own `locus_index`, which `_sorted_loci` only *orders* by. The two
                        # coincide under `strict`, where a dropped locus is a refusal, and that is
                        # also exactly what `reverse_module` recomputes by encounter order, so the
                        # round trip is a fixed point either way.
                        update["locus_index"] = index
                        update["locus_count"] = len(usable)
                        patched.append(v.model_copy(update=update))

        elif v.rsid is None and v.chrom is not None:
            # need rsid: fill from the single usable row (keeps the frozen coord key). `alts` is
            # filled too when the author left it out — the table knows the allele and the row would
            # otherwise reach the artifact without it, so reverse could not re-emit the resolved fact
            # and `resolution_signature` moved across the round-trip.
            row = next((lo for lo in loci if lo.rsid is not None), None)
            update: dict[str, object] = {}
            if row is not None:
                update["rsid"] = row.rsid
            if v.alts is None:
                supplier = row or next((lo for lo in loci if lo.alts), None)
                if supplier is not None and supplier.alts:
                    update["alts"] = supplier.alts
            if update:
                patched.append(v.model_copy(update=update))
            else:
                # Name the *place*, not the key. `variant_key` for a resolved substitution is a
                # `ga4gh:VA.…` allele id, so the old message read "Position
                # ga4gh:VA.aseiElOGc6FKVcLTpib-L1y4s1dwiYE2" — calling a content-addressed identity a
                # position, and giving the author nothing to look up. The row holds the coordinate.
                no_rsid.append(_locus_label(v))
                patched.append(v)

        else:
            # both authored (verify) or nothing to do
            if v.rsid is not None and v.chrom is not None and loci:
                _verify(v, loci, ladders)
            patched.append(v)

    if no_rsid:
        warnings.append(
            CodedWarning(
                "rsid_without_resolution_label",
                without_resolution_label(
                    subject=(
                        f"{len(no_rsid)} coordinate-authored row(s) stay coordinate-keyed "
                        f"({_examples(no_rsid)})"
                    ),
                    searched="the resolution table",
                )
                + "; re-run the enricher if you want the labels back-filled.",
            )
        )

    expanded_rows = 0
    for rsid, per_row in expansions.items():
        expanded_rows += sum(len(u) for u in per_row)
        warnings.append(
            CodedWarning("rsid_expanded_to_multiple_loci", _expansion_warning(rsid, per_row, genome_build))
        )

    # **The table is keyed by the AUTHORED key, so these two are asked over `variants`, not `patched`.**
    # An expansion rewrites `variant_key` to the locus's `ga4gh:VA.…` id, and a lookup of that id in a
    # table keyed by the rsID the author wrote misses every time — so the refusal the comment below
    # calls fatal in both modes was silently skipped on exactly the rows that expanded (RM207). Both
    # arms are also asked by the pre-flight through the shared helpers, which is what keeps a green
    # `validate` from being followed by a refusal (`@validate-refuses-all`).
    errors.extend(withdrawn_refusals(variants, resolution, genome_build))
    # Paired rather than accumulated into two lists: `strict`'s sentence for an ambiguous label is a
    # different and longer one than `best_effort`'s, and the two were only ever kept in step by both
    # loops iterating `_ambiguous_loci` in the same order.
    ladders.extend(
        LadderFinding(CodedWarning("rsid_ambiguous", warning), refusal=refusal)
        for warning, refusal in zip(
            ambiguous_warnings(variants, resolution, genome_build),
            ambiguous_refusals(variants, resolution, genome_build),
            strict=True,
        )
    )

    return ResolutionOutcome(
        variants=patched,
        plain_warnings=warnings,
        ladders=ladders,
        errors=errors,
        expanded_keys=len(expansions),
        expanded_rows=expanded_rows,
    )

resolve_positional_rows

resolve_positional_rows(
    rows: list[object],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str = "GRCh38",
) -> PositionalFill

Fill the resolved coordinate into one 0.4-family positional table, in place (RM43).

pharm_variants.csv, haplotypes.csv and heteroplasmy.csv may identify a row by rsID alone — their own models say so — and until 0.6 the compiler resolved variants.csv and materialized every other table verbatim. A consumer matching a patient VCF by position therefore matched nothing, silently, as an empty result rather than an error: the coordinates existed in the resolution.csv sitting beside the spec and reached the artifact by no path at all.

Four rules, and the last two are what keep this from being the naive repair the roadmap rejected:

  • Fill only what the author left empty. A cell the author wrote is never overwritten; that is enrich's inject-only doctrine (report, never repair) and it is also what makes the fill idempotent.
  • Fill from exactly one locus, or from none. One usable locus fills. Several are filtered by hosting_verdict against whatever allele the row states — a genotype on a pharm row, the defining allele on a haplotype junction, and nothing at all on a heteroplasmy band, which is a measurement over a locus rather than a claim about a genotype. If that leaves one, it fills; otherwise the row stays unplaced and is counted. There is deliberately no expansion: a one-to-many rsID expands variants.csv into N coord-keyed rows, and doing the same here would multiply a pharm annotation's (variant_key, drug, genotype, …) key across loci the author never named.
  • A row whose own coordinate contradicts the table is left alone. haplotypes.csv drafted from CPIC carries a start with no chrom, so the fill has to complete a half-coordinate — and completing it from a locus whose start disagrees would build a coordinate no source ever stated. Reported, never repaired, and never fatal: the same shape as resolve_from_table's _verify, minus the strict escalation, because the row is left exactly as authored.
  • The row is mutated, not copied, and its identity is frozen. variant_key/authored_ident are stamped at load from the authored subset, so filling cannot re-key the row and reverse_module re-emits the shape the author wrote (_write_table_csv).

GRCh38-bound for the same reason resolve_from_table is (RM15): the caller skips a non-GRCh38 module rather than joining rows this tier cannot re-derive an identity for.

Source code in compiler/src/just_dna_compiler/resolution.py
def resolve_positional_rows(
    rows: list[object],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str = "GRCh38",
) -> PositionalFill:
    """Fill the resolved coordinate into one 0.4-family positional table, **in place** (RM43).

    `pharm_variants.csv`, `haplotypes.csv` and `heteroplasmy.csv` may identify a row by rsID alone —
    their own models say so — and until 0.6 the compiler resolved `variants.csv` and materialized
    every other table verbatim. A consumer matching a patient VCF by position therefore matched
    nothing, silently, as an empty result rather than an error: the coordinates existed in the
    `resolution.csv` sitting beside the spec and reached the artifact by no path at all.

    Four rules, and the last two are what keep this from being the naive repair the roadmap rejected:

    * **Fill only what the author left empty.** A cell the author wrote is never overwritten; that is
      `enrich`'s inject-only doctrine (report, never repair) and it is also what makes the fill
      idempotent.
    * **Fill from exactly one locus, or from none.** One usable locus fills. Several are filtered by
      `hosting_verdict` against whatever allele the row states — a `genotype` on a pharm row, the
      defining `allele` on a haplotype junction, and nothing at all on a heteroplasmy band, which is a
      measurement over a locus rather than a claim about a genotype. If that leaves one, it fills;
      otherwise the row stays unplaced and is counted. There is deliberately **no expansion**: a
      one-to-many rsID expands `variants.csv` into N coord-keyed rows, and doing the same here would
      multiply a pharm annotation's `(variant_key, drug, genotype, …)` key across loci the author
      never named.
    * **A row whose own coordinate contradicts the table is left alone.** `haplotypes.csv` drafted
      from CPIC carries a `start` with no `chrom`, so the fill has to complete a half-coordinate — and
      completing it from a locus whose `start` disagrees would build a coordinate no source ever
      stated. Reported, never repaired, and never fatal: the same shape as `resolve_from_table`'s
      `_verify`, minus the strict escalation, because the row is left exactly as authored.
    * **The row is mutated, not copied, and its identity is frozen.** `variant_key`/`authored_ident`
      are stamped at load from the authored subset, so filling cannot re-key the row and
      `reverse_module` re-emits the shape the author wrote (`_write_table_csv`).

    GRCh38-bound for the same reason `resolve_from_table` is (RM15): the caller skips a non-GRCh38
    module rather than joining rows this tier cannot re-derive an identity for.
    """
    report = PositionalFill()
    for row in rows:
        loci = _usable_loci(resolution.get(getattr(row, "variant_key", None) or ""), genome_build)
        fillable = [
            name
            for name in ("rsid", "chrom", "start", "ref", "alts")
            if name in type(row).model_fields and getattr(row, name) is None
        ]
        # Deliberately **not** short-circuited on `not fillable`. A fully-populated row has nothing to
        # fill and can still contradict the table it is keyed into — reachable on `heteroplasmy.csv`,
        # the one positional model whose whole identity set is authorable — and the promise this
        # function makes is that such a disagreement is *reported*, never repaired. Skipping the
        # lookup there would have let a mistyped `start` sit beside a disagreeing `resolution.csv` row
        # in silence, which is the shape `resolve_from_table._verify` exists to catch on the SNP core.
        if not loci:
            if row.chrom is None or row.start is None:
                report.unplaced_absent += 1
            continue
        statement = _stated_allele(row)
        candidates = loci if len(loci) == 1 or statement is None else _hostable_loci(loci, statement)[0]
        if len(candidates) != 1:
            if row.chrom is None or row.start is None:
                report.unplaced_ambiguous += 1
            continue
        locus = candidates[0]
        conflict = _authored_conflict(row, locus)
        if conflict is not None:
            report.contradicted.append(conflict)
            continue
        if not fillable:
            continue
        for name in fillable:
            value = getattr(locus, name, None)
            if value is not None:
                setattr(row, name, value)
        report.filled += 1
    return report

withdrawn_refusals

withdrawn_refusals(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str,
) -> list[str]

The one resolution refusal that is fatal in both modes, as finished sentences (RM207).

A merged or absent rsID leaves the annotation intact — the module is dated, or the label is unserved. A withdrawn one is dbSNP repudiating the variant, so the annotation may be describing something that does not exist; carrying it under best_effort would be publishing a claim its own source has retracted. Never produced by the automated check (a retraction is indistinguishable from a never-assigned id through the live API), so it fires only where a curator recorded it.

Sentences rather than names, which is where this differs from unresolved_subjects beside it. That one returns subjects and lets each caller phrase them, because both phrasings are one short clause. This message names two things — the subject and the retracted rsID — and the standing rule is to share the predicate and copy the error; copying a sentence with two interpolations into a second caller is how the two drift, so the sentence is shared too and there is one of it.

Empty for a non-GRCh38 module, where resolve_from_table skips resolving wholesale.

Source code in compiler/src/just_dna_compiler/resolution.py
def withdrawn_refusals(
    variants: list[VariantRow], resolution: dict[str, list[ResolutionRow]], genome_build: str
) -> list[str]:
    """The one resolution refusal that is fatal in **both** modes, as finished sentences (RM207).

    A merged or absent rsID leaves the annotation intact — the module is dated, or the label is
    unserved. A *withdrawn* one is dbSNP repudiating the variant, so the annotation may be describing
    something that does not exist; carrying it under `best_effort` would be publishing a claim its own
    source has retracted. Never produced by the automated check (a retraction is indistinguishable
    from a never-assigned id through the live API), so it fires only where a curator recorded it.

    **Sentences rather than names**, which is where this differs from `unresolved_subjects` beside it.
    That one returns subjects and lets each caller phrase them, because both phrasings are one short
    clause. This message names two things — the subject *and* the retracted rsID — and the standing
    rule is to share the predicate and copy the error; copying a sentence with two interpolations into
    a second caller is how the two drift, so the sentence is shared too and there is one of it.

    Empty for a non-GRCh38 module, where `resolve_from_table` skips resolving wholesale.
    """
    if genome_build != "GRCh38":
        return []
    out: list[str] = []
    for v in variants:
        locus = _first_locus(v, resolution, lambda lo: lo.rsid_status == "withdrawn")
        if locus is not None:
            out.append(
                f"{v.variant_key}: dbSNP has WITHDRAWN {locus.rsid} — the variant itself was "
                f"retracted, so the annotation resting on it may be describing nothing. Remove the "
                f"row or re-key it onto a coordinate; this refuses in best_effort too, unlike a "
                f"merged or absent rsid."
            )
    return out

ambiguous_refusals

ambiguous_refusals(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str,
) -> list[str]

strict-only: the table marks this rsid ambiguous, so the pick is deterministic, not a fact.

Source code in compiler/src/just_dna_compiler/resolution.py
def ambiguous_refusals(
    variants: list[VariantRow], resolution: dict[str, list[ResolutionRow]], genome_build: str
) -> list[str]:
    """`strict`-only: the table marks this rsid ambiguous, so the pick is deterministic, not a fact."""
    return [
        f"{v.variant_key}: the resolution table marks this rsid ambiguous"
        + (f" (candidates: {lo.rsid_alternates})" if lo.rsid_alternates else "")
        + ". The label is a deterministic pick among equals, not a fact; an "
        "all-or-nothing artifact should not rest on it. Resolve it by hand in "
        "resolution.csv, or compile without strict."
        for v, lo in _ambiguous_loci(variants, resolution, genome_build)
    ]

ambiguous_warnings

ambiguous_warnings(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str,
) -> list[str]

The best_effort half of the same finding — the pick is carried, and said to be a pick.

Source code in compiler/src/just_dna_compiler/resolution.py
def ambiguous_warnings(
    variants: list[VariantRow], resolution: dict[str, list[ResolutionRow]], genome_build: str
) -> list[str]:
    """The `best_effort` half of the same finding — the pick is carried, and said to be a pick."""
    return [
        ambiguous_pick(
            subject=v.variant_key,
            observed="rsid resolved as AMBIGUOUS",
            candidates=lo.rsid_alternates or None,
        )
        for v, lo in _ambiguous_loci(variants, resolution, genome_build)
    ]

unresolved_subjects

unresolved_subjects(
    variants: list[VariantRow],
    resolution: dict[str, list[ResolutionRow]],
    genome_build: str,
) -> list[str]

Authored rows the injected table cannot place, named by rsID (or variant_key).

The predicate resolve_from_table will apply, factored out so the pre-flight can ask it before resolution runs and reach the same answer the compile does (S76). A row is unplaceable when it authors no coordinate of its own and the table holds no usable locus for its key — the same _usable_loci filter the fill uses, so a not_found sentinel and a row recorded under another build count as no locus here exactly as they do there.

Deliberately not a re-derivation of the compile's check: compile_module runs its refusal over outcome.variants, the post-expansion list, and the two agree because a subject with no usable locus is carried through the expansion unchanged. Sharing the predicate is what keeps that true — a second implementation beside the first is the drift this returns instead of.

Empty for a non-GRCh38 module, where resolve_from_table skips wholesale rather than resolving less: reporting every row unplaceable there would restate the skip warning once per row and blame the table for a limit of the tier.

Source code in compiler/src/just_dna_compiler/resolution.py
def unresolved_subjects(
    variants: list[VariantRow], resolution: dict[str, list[ResolutionRow]], genome_build: str
) -> list[str]:
    """Authored rows the injected table cannot place, named by rsID (or `variant_key`).

    The predicate `resolve_from_table` will apply, factored out so the pre-flight can ask it **before**
    resolution runs and reach the same answer the compile does (S76). A row is unplaceable when it
    authors no coordinate of its own and the table holds no usable locus for its key — the same
    `_usable_loci` filter the fill uses, so a `not_found` sentinel and a row recorded under another
    build count as no locus here exactly as they do there.

    Deliberately **not** a re-derivation of the compile's check: `compile_module` runs its refusal over
    `outcome.variants`, the post-expansion list, and the two agree because a subject with no usable
    locus is carried through the expansion unchanged. Sharing the predicate is what keeps that true —
    a second implementation beside the first is the drift this returns instead of.

    Empty for a non-GRCh38 module, where `resolve_from_table` skips wholesale rather than resolving
    less: reporting every row unplaceable there would restate the skip warning once per row and blame
    the table for a limit of the tier.
    """
    if genome_build != "GRCh38":
        return []
    return sorted(
        v.rsid or v.variant_key or ""
        for v in variants
        if v.chrom is None and not _usable_loci(resolution.get(v.variant_key or ""), genome_build)
    )

hosting_verdict

hosting_verdict(
    genotype: str, ref: str | None, alts: str | None
) -> bool | None

Can a locus spelling {ref} ∪ alts host genotype? Three-valued (RM31).

True it can, False it positively cannot, None this tier cannot tell — the house algebra, and the third value is the whole point: an indel has several valid spellings, so a string comparison reporting "does not fit" was asserting a verdict it had not reached.

The ladder, in order, and the order is load-bearing:

  1. No ref/alts recorded → True. Nothing is known about the locus's alleles and rejecting for lack of evidence is worse than accepting. Unchanged.
  2. * abstains — on both sides — and the rest is still judged (RM59). * says this sample's allele could not be observed here; it names no allele, so a locus cannot fail to offer it and cannot contradict it, and it belongs in the allele algebra on neither side. Without this the widening that made */G authorable would be self-defeating: no source spells * in an ALT list, so an rsid-authored */G resolved against a G>A locus fell through to step 6 and came back a confident False — the locus dropped from the expansion, the row unresolved, and strict refusing over a finding no authored edit could clear. Dropping the member rather than the whole verdict is what keeps the observable half sharp: */T at a G>A locus is still False, because the T really is a contradiction. Nothing observable on either side is None.

The locus side is not optional, and the first cut got this wrong on the reasoning that a * in ref/alts is reachable today, so excluding it might move an existing verdict. It does move existing verdicts, and it is still right: parsimony_reduce cannot strip a shared flank past a member that has none, so a * left in the locus stops the set reducing and collapses RM31 — hosting_verdict('C/CAG', 'AGAG', 'AG') is True while …('C/CAG', 'AGAG', 'AG,*') was False, refusing a correctly transcribed indel under --strict and telling the author to "replace it with the alleles the locus actually has". ALT=AG,* is ordinary joint-caller output.

What that costs is measured, and it is not only acceptances. Swept over every star-free genotype — i.e. everything authorable before RM59 — against loci with and without a *: for a locus spelled in nucleotides nothing changes at all, and for a locus spelling * verdicts move in both directions, None → False among them. Those are corrections. * was doing two things to the arithmetic at once, blocking the flank strip and lending its own single character to _indel_shaped's length set, so ref=AGAG alts=AG,* stayed unreduced, read as indel-shaped, and withheld on calls that were decidable all along — a genotype naming an allele the locus does not have escaped the finding. Same bug class as the RM5 guard's stated reason: characters counted out of a token that spells no sequence. No reference example carries a *, so the corpus could not have shown any of this. 3. The raw strings match → True. Checked before any normalization, so this step can only ever gain acceptances: whatever passed the raw comparison before still passes it, byte for byte. Step 2 sits above it and costs it nothing, provably rather than nearly — the called side is stripped first, so * is never on the left of the subset test, and dropping it from the right therefore cannot change the answer. The guarantee stops here: steps 5–8 are arithmetic over the reduced sets, and stripping the locus does change those, in both directions (see step 2). 4. Either side names a symbolic allele → None. A <DEL:1500> has no sequence, so it has no flank for parsimony_reduce and nothing to compare character by character — and a symbolic allele exists because the exact sequence is not known, so even two stated lengths that differ are summary information rather than grounds for a contradiction. Undecided, never "no match" (RM5). Below the raw comparison, so <DEL:1500> against a locus spelling <DEL:1500> still matches exactly. 5. The reduced allele sets match → True. alleles.parsimony_reduce strips the flank each collection shares, leaving the event; ClinVar's C/CAG and Ensembl's AGAG>AG both reduce to {'', 'AG'}, the SHOX 2 bp deletion that used to resolve to not_found. 6. The locus is a substitution or MNV → False. No flank, so no spelling freedom: an A/G genotype at a C>T locus is a real contradiction, and must stay one (a strand flip is exactly what that check catches). 7. The genotype names fewer than two distinct alleles → None. A homozygous C/C carries no frame — one string has nothing to be relative to — so against an indel locus there is genuinely nothing to compare. Reported as undecided, never as a contradiction. A */C reaches this too once the * has abstained, and for the same reason: one string has nothing to be relative to. 8. An event length the locus does not offer → False. The confident negative: left-alignment moves an indel, it never changes how many bases the event adds or removes, so a 1 bp insertion cannot be a 2 bp deletion however it is spelled. 9. Otherwise → None. Same lengths, different content: one variant rotated inside a repeat, or two different variants, and only the reference sequence can say which. The enricher can settle it (seqrepo); the compiler holds no reference by charter (P2), so it withholds.

Public because both resolvers must agree on it: this module's injected-table path and the deprecated DuckDB path in just-dna-enricher. Digest parity between the two is a documented guarantee, so a filter applied on one side only would silently break it.

Source code in compiler/src/just_dna_compiler/resolution.py
def hosting_verdict(genotype: str, ref: str | None, alts: str | None) -> bool | None:
    """Can a locus spelling `{ref} ∪ alts` host `genotype`? **Three-valued** (RM31).

    `True` it can, `False` it positively cannot, `None` this tier cannot tell — the house algebra, and
    the third value is the whole point: an indel has several valid spellings, so a string comparison
    reporting "does not fit" was asserting a verdict it had not reached.

    The ladder, in order, and the order is load-bearing:

    1. **No `ref`/`alts` recorded → `True`.** Nothing is known about the locus's alleles and rejecting
       for lack of evidence is worse than accepting. Unchanged.
    2. **`*` abstains — on both sides — and the rest is still judged (RM59).** `*` says *this sample's
       allele could not be observed here*; it names no allele, so a locus cannot fail to offer it and
       cannot contradict it, and it belongs in the allele algebra on neither side. Without this the
       widening that made `*/G` authorable would be self-defeating: no source spells `*` in an ALT list,
       so an rsid-authored `*/G` resolved against a `G>A` locus fell through to step 6 and came back a
       confident **`False`** — the locus dropped from the expansion, the row unresolved, and `strict`
       refusing over a finding no authored edit could clear. Dropping the member rather than the whole
       verdict is what keeps the *observable* half sharp: `*/T` at a `G>A` locus is still `False`,
       because the `T` really is a contradiction. Nothing observable on either side is `None`.

       **The locus side is not optional, and the first cut got this wrong** on the reasoning that a `*`
       in `ref`/`alts` is reachable today, so excluding it might move an existing verdict. It does move
       existing verdicts, and it is still right: `parsimony_reduce` cannot strip a shared flank past a
       member that has none, so a `*` left in the locus stops the set reducing and collapses RM31 —
       `hosting_verdict('C/CAG', 'AGAG', 'AG')` is `True` while `…('C/CAG', 'AGAG', 'AG,*')` was
       `False`, refusing a correctly transcribed indel under `--strict` and telling the author to
       "replace it with the alleles the locus actually has". `ALT=AG,*` is ordinary joint-caller output.

       **What that costs is measured, and it is not only acceptances.** Swept over every star-free
       genotype — i.e. everything authorable before RM59 — against loci with and without a `*`: for a
       locus spelled in nucleotides **nothing changes at all**, and for a locus spelling `*` verdicts
       move in *both* directions, `None → False` among them. Those are corrections. `*` was doing two
       things to the arithmetic at once, blocking the flank strip and lending its own single character
       to `_indel_shaped`'s length set, so `ref=AGAG alts=AG,*` stayed unreduced, read as indel-shaped,
       and withheld on calls that were decidable all along — a genotype naming an allele the locus does
       not have escaped the finding. Same bug class as the RM5 guard's stated reason: characters counted
       out of a token that spells no sequence. No reference example carries a `*`, so the corpus could
       not have shown any of this.
    3. **The raw strings match → `True`.** Checked *before* any normalization, so *this step* can only
       ever gain acceptances: whatever passed the raw comparison before still passes it, byte for byte.
       Step 2 sits above it and costs it nothing, *provably* rather than nearly — the called side is
       stripped first, so `*` is never on the left of the subset test, and dropping it from the right
       therefore cannot change the answer. **The guarantee stops here**: steps 5–8 are arithmetic over
       the *reduced* sets, and stripping the locus does change those, in both directions (see step 2).
    4. **Either side names a symbolic allele → `None`.** A `<DEL:1500>` has no sequence, so it has no
       flank for `parsimony_reduce` and nothing to compare character by character — and a symbolic
       allele exists *because* the exact sequence is not known, so even two stated lengths that differ
       are summary information rather than grounds for a contradiction. Undecided, never "no match"
       (RM5). Below the raw comparison, so `<DEL:1500>` against a locus spelling `<DEL:1500>` still
       matches exactly.
    5. **The reduced allele sets match → `True`.** `alleles.parsimony_reduce` strips the flank each
       collection shares, leaving the event; ClinVar's `C/CAG` and Ensembl's `AGAG>AG` both reduce to
       `{'', 'AG'}`, the SHOX 2 bp deletion that used to resolve to `not_found`.
    6. **The locus is a substitution or MNV → `False`.** No flank, so no spelling freedom: an `A/G`
       genotype at a `C>T` locus is a real contradiction, and must stay one (a strand flip is exactly
       what that check catches).
    7. **The genotype names fewer than two distinct alleles → `None`.** A homozygous `C/C` carries no
       frame — one string has nothing to be relative to — so against an indel locus there is genuinely
       nothing to compare. Reported as undecided, never as a contradiction. A `*/C` reaches this too
       once the `*` has abstained, and for the same reason: one string has nothing to be relative to.
    8. **An event length the locus does not offer → `False`.** The confident negative:
       left-alignment moves an indel, it never changes how many bases the event adds or removes, so a
       1 bp insertion cannot be a 2 bp deletion however it is spelled.
    9. **Otherwise → `None`.** Same lengths, different content: one variant rotated inside a repeat, or
       two different variants, and only the reference sequence can say which. The enricher can settle
       it (seqrepo); the compiler holds no reference by charter (P2), so it withholds.

    Public because **both** resolvers must agree on it: this module's injected-table path and the
    deprecated DuckDB path in `just-dna-enricher`. Digest parity between the two is a documented
    guarantee, so a filter applied on one side only would silently break it.
    """
    if not ref or not alts:
        return True
    locus = {ref.strip().upper()} | {a.strip().upper() for a in alts.split(",") if a.strip()}
    called = {a.upper() for a in split_genotype(genotype)}

    # RM59, and kept visibly separate from the RM5 guard below it because they are separate axes: this
    # one drops a *member* that makes no claim, that one withholds the whole *verdict* because a claim
    # cannot be compared. `*` is not an allele — it reports that the sample's allele could not be
    # observed here — so it is neither offered by a locus nor contradicted by one, and every step from
    # here on is a comparison of characters it has none of. Removing it leaves the observable half to be
    # judged normally; removing it and finding nothing left means nothing was seen at all.
    #
    # **Both sides, and the locus side is not optional.** A `*` in `alts` is what a joint-called VCF
    # writes, and left in the locus set it silently defeats `parsimony_reduce`: the shared flank cannot
    # be stripped past a member that has none, so `parsimony_reduce({'AGAG','AG','*'})` returns the set
    # unreduced and RM31's reconciliation collapses — `hosting_verdict('C/CAG', 'AGAG', 'AG')` is `True`
    # while `hosting_verdict('C/CAG', 'AGAG', 'AG,*')` was a confident `False`, refusing a correctly
    # transcribed indel under `--strict` and advising the author to "replace it with the alleles the
    # locus actually has". One rule on both sides is also the only consistent reading: an allele the
    # algebra must ignore cannot be one the algebra ignores in one direction.
    #
    # **Above the raw comparison, not below it**, which is the opposite of where the RM5 guard sits and
    # is forced rather than chosen: the whole point is that the *rest* of the call must still be matched
    # against the locus, and `*/G` at a `G>A` locus is exactly the case — `{'*','G'}` is not a subset,
    # so below the comparison it fell through to the substitution arm and came back a confident `False`.
    # It costs the raw comparison nothing, and provably rather than nearly: `called` is stripped first,
    # so `*` is not on the left of the subset test, and dropping it from the right therefore cannot
    # change the answer. Nothing reachable before RM59 could put one in a genotype at all.
    called = {a for a in called if not is_unobservable_allele(a)}
    locus = {a for a in locus if not is_unobservable_allele(a)}
    if not called or not locus:
        return None

    if called <= locus:
        return True

    # RM5. A symbolic allele names a variant whose sequence is deliberately unspelled, so from here on
    # every step is a comparison of *characters* and there are none to compare: `parsimony_reduce`
    # would treat `<DEL:1500>` as a nine-character sequence, `_indel_shaped` would read its bracket
    # count as an event length, and step 8 would then return a confident `False` on arithmetic over a
    # token. Undecided is the honest answer, and it only ever *adds* acceptances (`genotype_fits`
    # keeps an undecided locus), so no expansion and no digest moves for a module spelling in bases.
    if any(is_symbolic_allele(a) for a in locus | called):
        return None

    called_events = parsimony_reduce(called)
    locus_events = parsimony_reduce(locus)
    if len(called_events) > 1 and called_events <= locus_events:
        return True
    if not _indel_shaped(locus_events):
        return False
    if len(called) < 2:
        return None
    if not _indel_shaped(called_events):
        return False
    lengths = {len(event) for event in locus_events}
    if any(len(event) not in lengths for event in called_events - locus_events):
        return False
    return None

undecided_reason

undecided_reason(
    genotype: str, ref: str | None, alts: str | None
) -> str

Why hosting_verdict withheld — the clause a caller appends when the answer was None.

None has five causes and the message asserted one of them. Both reporting sites — this module's expansion warning and the enricher's twin — said "the two spellings describe events of the same size but different content, either one indel re-anchored inside a repeat or two different variants", which is step 9's cause and false for the other four: a symbolic allele was never compared at all (RM5), an all-* call observed nothing and an all-* locus offers nothing (RM59), and a homozygous call carries no frame. Stating a cause the tier did not establish is the same defect as the ([], None) collapse S20 was filed for — two ways of returning nothing rendered as one sentence — and here it sends the reader to check a reference sequence for a row where no reference could help.

Mirrors hosting_verdict's withholding branches in its order, so the two answer the same question the same way; a test walks every None-producing shape and asserts the pairing, which is what keeps a sixth branch from quietly inheriting a fifth's explanation.

The sets are built exactly as the predicate builds them — including dropping an empty ref rather than folding "" in, which would put a member in locus that is neither observable nor an allele and quietly defeat the branch below it. Keeping the two constructions identical is the entire point of the pair.

Source code in compiler/src/just_dna_compiler/resolution.py
def undecided_reason(genotype: str, ref: str | None, alts: str | None) -> str:
    """Why `hosting_verdict` withheld — the clause a caller appends when the answer was `None`.

    **`None` has five causes and the message asserted one of them.** Both reporting sites — this
    module's expansion warning and the enricher's twin — said *"the two spellings describe events of
    the same size but different content, either one indel re-anchored inside a repeat or two different
    variants"*, which is step 9's cause and false for the other four: a symbolic allele was never
    compared at all (RM5), an all-`*` call observed nothing and an all-`*` locus offers nothing (RM59),
    and a homozygous call carries no frame. Stating a cause the tier did not establish is the same
    defect as the `([], None)` collapse S20 was filed for — two ways of returning nothing rendered as
    one sentence — and here it sends the reader to check a reference sequence for a row where no
    reference could help.

    Mirrors `hosting_verdict`'s withholding branches **in its order**, so the two answer the same
    question the same way; a test walks every `None`-producing shape and asserts the pairing, which is
    what keeps a sixth branch from quietly inheriting a fifth's explanation.

    The sets are built exactly as the predicate builds them — including dropping an empty `ref` rather
    than folding `""` in, which would put a member in `locus` that is neither observable nor an allele
    and quietly defeat the branch below it. Keeping the two constructions identical is the entire point
    of the pair.
    """
    locus = {a.strip().upper() for a in (ref or "", *(alts or "").split(",")) if a.strip()}
    called = {a.upper() for a in split_genotype(genotype)}
    observable_called = {a for a in called if not is_unobservable_allele(a)}
    observable_locus = {a for a in locus if not is_unobservable_allele(a)}

    if not observable_called:
        return (
            "the call observed no allele at this position at all — every member is VCF's `*`, so "
            "there is nothing to match against the locus and nothing a reference could settle"
        )
    if not observable_locus:
        return "the locus records no allele that is not VCF's `*`, so it offers nothing to match"
    if any(is_symbolic_allele(a) for a in observable_locus | observable_called):
        return (
            "a symbolic/structural allele names a variant whose sequence is deliberately unspelled, "
            "so the two were never compared character by character — undecided, not a mismatch"
        )
    if len(observable_called) < 2:
        return (
            "the genotype names one distinct allele, so it carries no flank to be relative to and "
            "the spelling cannot be reconciled against an indel locus"
        )
    return (
        "the two spellings describe events of the same size but different content, which is either "
        "one indel re-anchored inside a repeat or two different variants, and telling those apart "
        "needs the reference sequence (run the enricher)"
    )

contradiction_reason

contradiction_reason(
    genotype: str, ref: str | None, alts: str | None
) -> str

Why hosting_verdict was a confident False — undecided_reason's twin, same contract (S85).

False has two causes and the enricher's message asserted one of them for both. Step 8 is the event-length arm — a 1 bp insertion cannot be a 2 bp deletion however it is spelled — and "the event sizes differ, which re-anchoring cannot change" is its reason. Step 6 is the substitution/MNV arm, where the sizes are identical and what makes the verdict confident is the absence of a flank: same-length alleles have no spelling freedom, so a genotype naming alleles the locus does not offer is a real contradiction. Saying "the event sizes differ" there is a false claim about two 1 bp substitutions, and it sends the author hunting a second variant sharing the rsID.

The case that made this worth splitting is a strand flip: a paper's supplementary published against hg19 spells the submitted strand, so an authored A/G meets GRCh38's C>T. The verdict is correct — on the strand it is written on, that genotype really cannot be hosted — and the remedy is a column the old sentence never mentioned. So the flip is named where it is established, and the step-6 fallback stays honest about what it did not establish rather than inventing a cause.

Mirrors hosting_verdict's False branches in its order, for the reason its twin does: a test walks every False-producing shape and asserts the pairing, so a third arm cannot silently inherit a second's explanation.

Source code in compiler/src/just_dna_compiler/resolution.py
def contradiction_reason(genotype: str, ref: str | None, alts: str | None) -> str:
    """Why `hosting_verdict` was a confident `False` — `undecided_reason`'s twin, same contract (S85).

    **`False` has two causes and the enricher's message asserted one of them for both.** Step 8 is the
    event-length arm — a 1 bp insertion cannot be a 2 bp deletion however it is spelled — and *"the
    event sizes differ, which re-anchoring cannot change"* is its reason. Step 6 is the substitution/MNV
    arm, where the sizes are identical and what makes the verdict confident is the *absence of a flank*:
    same-length alleles have no spelling freedom, so a genotype naming alleles the locus does not offer
    is a real contradiction. Saying "the event sizes differ" there is a false claim about two 1 bp
    substitutions, and it sends the author hunting a second variant sharing the rsID.

    The case that made this worth splitting is a strand flip: a paper's supplementary published against
    hg19 spells the submitted strand, so an authored `A/G` meets GRCh38's `C>T`. The verdict is correct
    — on the strand it is written on, that genotype really cannot be hosted — and the *remedy* is a
    column the old sentence never mentioned. So the flip is named where it is established, and the
    step-6 fallback stays honest about what it did not establish rather than inventing a cause.

    Mirrors `hosting_verdict`'s `False` branches in its order, for the reason its twin does: a test
    walks every `False`-producing shape and asserts the pairing, so a third arm cannot silently inherit
    a second's explanation.
    """
    if strand_flip_explains(genotype, ref, alts):
        return (
            f"the authored alleles are the reverse complement of this locus's — reading {genotype} on "
            "the other strand fits it exactly. The source HAS this variant; what does not match is the "
            "strand your alleles are written on, which is what a supplementary table published against "
            "an older assembly usually carries. Check the strand before looking for a second variant"
        )
    locus = {a.strip().upper() for a in (ref or "", *(alts or "").split(",")) if a.strip()}
    called = {a.upper() for a in split_genotype(genotype)}
    observable_locus = {a for a in locus if not is_unobservable_allele(a)}
    observable_called = {a for a in called if not is_unobservable_allele(a)}
    if not _indel_shaped(parsimony_reduce(observable_locus)):
        return (
            "the locus is a substitution or MNV, so its alleles have no shared flank to re-anchor on "
            "and there is no other spelling of them: the genotype names an allele this locus does not "
            "offer. Either it is a different variant sharing the rsID, or the alleles were transcribed "
            "from a record this one is not"
        )
    if not _indel_shaped(parsimony_reduce(observable_called)):
        return (
            "the genotype's alleles are all the same length while the locus's are not, so the two "
            "describe different events rather than two spellings of one"
        )
    return (
        "the event sizes differ, which re-anchoring cannot change, so this is a different variant "
        "sharing the rsID rather than another spelling"
    )

genotype_fits

genotype_fits(
    genotype: str, ref: str | None, alts: str | None
) -> bool

Whether a locus can host genotype, collapsing "cannot tell" into "keep it".

The boolean face of hosting_verdict, kept because both resolvers and three call sites read it, and because the collapse it performs is the module's existing doctrine: only a positive contradiction rejects. An undecidable spelling is therefore kept, exactly as a locus with no recorded alleles is. A caller that needs to report the difference asks hosting_verdict instead.

Source code in compiler/src/just_dna_compiler/resolution.py
def genotype_fits(genotype: str, ref: str | None, alts: str | None) -> bool:
    """Whether a locus can host `genotype`, collapsing "cannot tell" into "keep it".

    The boolean face of `hosting_verdict`, kept because both resolvers and three call sites read it, and
    because the collapse it performs is the module's existing doctrine: **only a positive contradiction
    rejects.** An undecidable spelling is therefore kept, exactly as a locus with no recorded alleles is.
    A caller that needs to *report* the difference asks `hosting_verdict` instead.
    """
    return hosting_verdict(genotype, ref, alts) is not False

spelling_caveat

spelling_caveat(ref: str | None, alts: str | None) -> str

The clause to append when a "cannot host" verdict is really about allele spelling.

Empty for a locus spelled in nucleotides, which is nearly all of them, so a caller can append it unconditionally.

hosting_verdict reaches False for a substitution locus by set difference, and it is right to: a substitution has no shared flank, so no spelling freedom, so A/G at a C>T locus is a real contradiction and must stay one — that is what makes the strand-flip check sharp. But it reaches the same False when the locus is spelled T>Y, and there the mismatch is between the cell and the nucleotide alphabet, not between the genotype and the variant. Reporting the generic message there sends the author to re-examine a genotype that was correct all along.

Three reasons reach this caveat, kept apart because what the author does next differs: an ambiguity code is an uncertainty that can never be expanded into definite alleles (expanding N to A,C,G,T asserts four alleles nobody stated), a well-formed symbolic allele is held by the grammar since RM5 and so is never a spelling defect at all, and everything else non-nucleotide is a grammar gap. None is repaired here — this tier reports.

* is the fourth classification and is excluded here (RM59), because hosting_verdict strips it from both sides before comparing, so it can no longer contribute to the False this sentence explains. Left in, it inverted the caveat's whole point: genotype=C/G at ref=A alts="T,*" is a real genotype error, and the author was told to read it as a spelling problem instead.

Source code in compiler/src/just_dna_compiler/resolution.py
def spelling_caveat(ref: str | None, alts: str | None) -> str:
    """The clause to append when a "cannot host" verdict is really about allele *spelling*.

    Empty for a locus spelled in nucleotides, which is nearly all of them, so a caller can append it
    unconditionally.

    `hosting_verdict` reaches `False` for a substitution locus by set difference, and it is right to: a
    substitution has no shared flank, so no spelling freedom, so `A/G` at a `C>T` locus is a real
    contradiction and must stay one — that is what makes the strand-flip check sharp. But it reaches the
    same `False` when the locus is spelled `T>Y`, and there the mismatch is between the *cell* and the
    nucleotide alphabet, not between the genotype and the variant. Reporting the generic message there
    sends the author to re-examine a genotype that was correct all along.

    Three reasons reach this caveat, kept apart because what the author does next differs: an ambiguity
    code is an uncertainty that can never be expanded into definite alleles (expanding `N` to `A,C,G,T`
    asserts four alleles nobody stated), a well-formed symbolic allele is *held* by the grammar since
    RM5 and so is never a spelling defect at all, and everything else non-nucleotide is a grammar gap.
    None is repaired here — this tier reports.

    **`*` is the fourth classification and is excluded here (RM59)**, because `hosting_verdict` strips
    it from both sides before comparing, so it can no longer contribute to the `False` this sentence
    explains. Left in, it inverted the caveat's whole point: `genotype=C/G` at `ref=A alts="T,*"` is a
    real genotype error, and the author was told to read it as a spelling problem instead.
    """
    offenders = {
        allele: reason
        for allele, reason in non_nucleotide_alleles(ref, alts).items()
        if not is_unobservable_allele(allele)
    }
    if not offenders:
        return ""
    return (
        " Read that as a spelling problem rather than a variant one: the locus records "
        + _spelling_clauses(offenders)
        + "."
    )