just_dna_enricher.resolver¶
just_dna_enricher.resolver ¶
Bidirectional rsid <-> position resolver using an Ensembl DuckDB (GRCh38).
Unlike the original pipelines resolver, this standalone version injects the Ensembl
reference: the caller passes either a prebuilt .duckdb file or a directory of Ensembl parquet
files, and the resolver builds an ephemeral ensembl_variations view over it. It never reaches
into just-dna-pipelines and never downloads anything — provisioning the reference is the
caller's responsibility (the marketplace pins one reference for the whole ecosystem).
EnsemblReferenceError ¶
Bases: FileNotFoundError
Raised when a provided Ensembl reference is neither a usable .duckdb nor a parquet dir.
PairCheck
dataclass
¶
PairCheck(
disagreements: list[str] = list(),
subjects: int = 0,
unknown: list[str] = list(),
undecided: list[str] = list(),
not_checked: str | None = None,
answered: int = 0,
)
What the rsid↔coordinate comparison did, not just what it found (RM45).
Same shape and same argument as sequences.RefCheck: the disagreements alone answer one of the
two questions a reader has, and an empty list says three different things — every pair agreed, the
reference had no record of any of them, or there was no reference at all. subjects is the
denominator the check really ran over, unknown names the pairs it could not put the question for
(an answered absence is not an agreement — S20), and not_checked carries a
VALID_VERIFICATION_SKIPS key when the whole pass could not run.
AlleleMismatch
dataclass
¶
AlleleMismatch(
rsid: str,
genotype: str,
loci: tuple[str, ...] = (),
offered: tuple[str, ...] = (),
strand_flip: bool = False,
)
An rsID the source has, whose loci cannot host the authored genotype (S85).
The state this records is not the one status="not_found" describes, and the two were
indistinguishable in the artifact until this existed: the source was asked, it answered, and every
locus it gave was rejected by the allele-aware filter. not_found says the source has no record
of the rsID, which is a different fact and — here — a false one. Same distinction
unreachable_rsids and unconsulted_rsids draw one branch over, arriving from a third direction:
the asking succeeded and the answer did not match.
Carried as a structured finding rather than spelled as a new VALID_RESOLUTION_STATUS member,
because that vocabulary is a wire contract every reader of a published resolution.csv shares,
and because the row itself is honestly unresolved either way — what was wrong was never the row's
existence, only the reason it gave. The row therefore stays exactly where it was, which also keeps
resolution_signature still (variant_key and rsid are fact fields; status is not).
strand_flip is the one cause this tier can settle from the allele strings alone, and it is by
far the most common: a paper's supplementary published against hg19 routinely spells the submitted
strand, so its G/A meets GRCh38's C/T. False means not established, never established
otherwise — an allele that cannot be complemented withholds it, and a reader must not invert it.
resolve_variants ¶
resolve_variants(
variants: list[VariantRow],
ensembl_cache: Path | None = None,
genome_build: str = "GRCh38",
) -> tuple[list[VariantRow], list[str]]
Fill in missing rsid or position from the injected Ensembl reference (GRCh38).
Variants that already carry both identifiers are returned unchanged. If no reference is available, resolution is skipped with a warning rather than raising.
One-to-many rsid → expansion. A no-coord rsid that maps to several loci is expanded into one
row per locus, each re-keyed to its coordinate (variant_key), so the N loci get N distinct
identities — a paralog/SV signal a consumer can count (data-agnostic). A 1:1 rsid just fills the
coordinate and keeps its rsid key. variant_key is frozen (base.derive_variant_key), so filling
a coord/rsid never re-keys a row (Principle 7); only expansion reassigns it.
GRCh38-bound. The reference is GRCh38, so resolution runs only for a GRCh38 module; a GRCh37/T2T build is skipped with a warning (positions are not re-resolved cross-build — RM15), rather than corrupting coordinates against the wrong assembly.
Bidirectional consistency (inject-only, no network). For rows that authored both an rsid and a coordinate, the same injected reference is used to check that the coordinate is among the rsid's loci and the rsid among the coordinate's ids; a disagreement is a warning (it may be a dbSNP merge/build difference — never fatal, matching the resolver's best-effort stance).
Source code in enricher/src/just_dna_enricher/resolver.py
97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 | |
lookup_loci ¶
lookup_loci(
reference: Path,
rsids: list[str],
positions: list[
tuple[
str | None, int | None, str | None, str | None
]
],
) -> tuple[
dict[str, list[dict]], dict[tuple, list[str]], list[str]
]
Public cache lookup the enricher uses to build a resolution table (0.5).
Returns (rsid -> [loci], (chrom,start,ref,alt) -> [candidate rsids], warnings) from an injected
Ensembl reference. The reverse map is allele-aware (matches the authored alt when given) and
returns all candidate rsIDs at that exact allele — never a silent first-pick — so the caller can
take a deterministic pick and flag a genuine multi-rsid allele as ambiguous. Inject-only.
Source code in enricher/src/just_dna_enricher/resolver.py
probe_table ¶
probe_table(
con: DuckDBPyConnection,
name: str,
columns: Sequence[tuple[str, str]],
rows: Sequence[tuple],
) -> None
Materialize rows into a TEMP table so a batch lookup can hash-join instead of OR-chaining.
DuckDB cannot fold a disjunction of equality conjunctions into a hash probe, so
WHERE (chrom=? AND start=? AND ref=? AND alt=?) OR … is evaluated against every row of the
reference: the cost grows with wanted × rows, i.e. quadratically in the module. Measured on the
4,431,781-record ClinVar snapshot with 5,000 alleles, same connection, identical output: 127 s
OR-chained against 0.13 s joined. A gene panel of any real size simply does not finish — this
is what made enrich() run two hours at 12% CPU with no I/O and look like a deadlock.
A single-column list does not need this: x IN (?, …) is hashed and is already fast
(_lookup_positions_by_rsid, clinvar.citations_for). What the planner does not do is fold
x = ? OR x = ? OR … into that IN — same 4,793 genes cost 15.22 s OR-chained and 1.02 s as
IN — so a single-column predicate is spelled IN, and only multi-column ones come here.
The rows are rendered as SQL literals, and that is deliberate — it is the whole speed-up.
The join itself is nearly free; what costs is handing 5,000 rows to DuckDB, and its Python
parameter binding is where the time goes. Measured on the same snapshot and sample, from
WHERE-clause to result: literal VALUES 0.21 s, a composite-key IN (?, …) with scalar
parameters 1.04 s, a parameterized UNNEST(?::VARCHAR[]) 3.51 s, executemany 8.6 s
— against 88 s for the OR-chain this replaces. So the obvious "fix" of parameterizing it back
gives up most of the win; do not make that change without re-measuring. Values are escaped the
same way clinvar._connect escapes the parquet path, and a numeric column is coerced through
int() rather than quoted, so nothing reaches SQL unexamined.
columns is (name, sql_type) pairs; an INTEGER/BIGINT column takes its value through
int(), everything else is quoted. The table is replaced on each call so one connection can serve
several probes. Callers keep their own ORDER BY: a join reorders nothing by itself, and emitted
row order is digest-visible (Principle 7).
Source code in enricher/src/just_dna_enricher/resolver.py
coordinate_verdict ¶
coordinate_verdict(
chrom: str | None,
start: int | None,
ref: str | None,
loci: Sequence[dict],
) -> bool | None
Does the reference place this rsID where the row says? Three-valued, hosting_verdict's shape.
The per-row verdict behind both of this tier's rsid↔coordinate checks — the one enrich() puts
against the injected snapshot and the deprecated DuckDB path's _check_rsid_coord_consistency
below — so the two cannot drift into two readings of one question.
True— one of the rsID's loci sits at the authored coordinate.False— none does, and every locus the reference gives is a substitution (and the row's ownref, where it states one, is too). A substitution has no shared flank, so its position is unambiguous and a mismatch can only be a different place. Same reasoning that keepshosting_verdict's confident negative sharp.None— none does, but an indel is involved. One deletion has several valid spellings and re-anchoring moves the coordinate — ClinVar'sX:634689 CAG>Cand Ensembl'sX:634690 AGAG>AGare the same 2 bp deletion (RM31) — so a differing position is not by itself a contradiction, and reporting one would be the false accusation that item exists to end.
Compared at chrom:start, not chrom:start:ref. Two reasons, and the first is what a
CPIC-drafted module runs into: ref is optional on HaplotypeRow/PharmVariantRow, and
pgx_draft writes exactly that shape, so keying on it renders 10:94781859:None and a legal row
can never match anything. The second is that a ref disagreeing with the reference is the
reference-allele check's finding (sequences.verify_reference_alleles) — reporting it here too
would give one defect two names, in the register the manifest publishes. The contig is normalized
on both sides for the same class of reason: chr10 and 10 are one place, and no PGx model runs
the chrom validator that would have folded them together.
Callers pass the loci the chain actually gathered and handle an empty list themselves: no record is a question never put, not an agreement, and collapsing the two is the S20 defect.
Source code in enricher/src/just_dna_enricher/resolver.py
coordinate_disagreement ¶
coordinate_disagreement(
rsid: str,
chrom: str | None,
start: int | None,
ref: str | None,
loci: Sequence[dict],
) -> str | None
The sentence for a False verdict above, or None for every other answer.
A convenience for the caller that only reports contradictions; anything needing to tell an
agreement from an undecided pair asks coordinate_verdict directly.
Source code in enricher/src/just_dna_enricher/resolver.py
disagreement_message ¶
disagreement_message(
rsid: str,
chrom: str | None,
start: int | None,
loci: Sequence[dict],
) -> str
The sentence a False verdict is reported as, in one place so the two callers cannot drift.
Source code in enricher/src/just_dna_enricher/resolver.py
check_rsid_coordinates ¶
check_rsid_coordinates(
pairs: Sequence[
tuple[str, str | None, int | None, str | None]
],
rsid_loci: Mapping[str, Sequence[dict]],
answered: Set[tuple[str, str]] = frozenset(),
) -> PairCheck
Compare each authored (rsid, chrom, start, ref) pair against the loci that rsid resolves to.
Takes the loci the caller's chain already gathered rather than opening anything, so the pass costs
no lookup of its own and no egress: enrich() widens the batch it was already sending.
A pair is compared, unplaceable, or undecided, and only the first is a subject. unknown is a
pair whose rsID the reference has no record of — a question never put — and undecided is one
where it has a record and the verdict is None (an indel, whose coordinate re-anchoring can move
legitimately). Both stay outside subjects, so findings/subjects describes exactly the pairs a
verdict was reached on, and both are named to the caller so what was not compared can be stated.
One direction only — is the authored coordinate among the rsid's loci — where the deprecated
path below also asks the converse (is the authored rsid among the ids at that position). The
converse needs a position→id lookup at chrom:start:ref granularity, and the reverse map
enrich() holds is allele-exact; asking it would report a disagreement whenever the authored
ALT is spelled differently from the reference's, which is a false accusation about a row rather
than a finding. The compiler's resolution._verify asks this same single direction over the
injected table, which is what makes the two halves one question.
answered is the author's overlay, and a pair it names is counted rather than reported
(RM136). Until 0.7 an author who corrected a resolution.csv coordinate through overrides.csv
— the mechanism RM124 built for exactly this — kept being told the same thing on every run, with no
way to clear it and nothing saying the correction had been honoured one tier over. A pair is
answered when the overlay updates that row's own coordinate cell, per field: correcting a
coordinate silences this check and leaves an unrelated finding standing. Per row was the cheaper
rule and was refused — it would silence findings the author never looked at, which is the
silent-suppress hole the overlay's own design calls its worst case.
Answered is not agreed, and the count is what keeps that honest. The pair leaves
disagreements and stays in subjects, with PairCheck.answered recording how many. A reader
still sees that the module and the source differ here; what changes is that the difference reads as
settled by the author rather than as work owed. Dropping it from the denominator would report a
cleaner module than there is — the silent-success shape this codebase keeps closing.