just_dna_enricher.alphagenome_check¶
just_dna_enricher.alphagenome_check ¶
The Atlas as a resolver — the three things the local artifact provably cannot do (RM193).
RM191's snapshot answers "what did AlphaGenome score this variant?" for 8.8 billion SNVs offline.
This pass exists for the questions it cannot answer, and deliberately for nothing else. Reports,
never repairs (@enrichment-is-validation).
(a) A knot the caller's threshold falls inside. The snapshot stores raw_score and reconstructs
PHRED through the knot table, which publishes an interval per printed score rather than a point.
Where a caller's threshold lands inside that interval the local data genuinely cannot decide, and
the Atlas can: its raw_score is a float32, ~7 significant digits against the file's 4, and at
the atom that causes every threshold-3 flip six rows the file prints identically as 0.00076 come
back distinct and reproduce the published PHRED exactly at five decimals (§ 4.7.1).
That scope is small and decidable before any request is made — from 466 KB of knots, without reading a data row. Genome-wide exactly one knot straddles any integer threshold from 1 to 50, so a module's candidate set is normally empty and occasionally a handful. The check refuses to run unbounded: rebuilding the column wholesale is 272 days and ~92 M RPCs at the measured rate, and prohibition 3 makes it the wrong shape of request anyway.
(b) A REF that disagrees with GRCh38. The Atlas validates it and names the real base —
"reference base does not match the expected reference base: G." A local file lookup cannot produce
that finding: it simply misses, and a miss is indistinguishable from an unscored position
(@va-omits-ref — a VA does not encode ref). So the finding carries the observed base, grouped by
reason (@ref-mismatch-causes).
(c) An answer that does not exist. An indel returns UNIMPLEMENTED, which is the third state
rather than a refusal: the request was legal and the Atlas has no score. It is recorded as
could-not-ask and never as a zero (@unreachable-not-absent), and --offline produces the same
third state rather than a silent skip — nobody-asked and asked-and-absent are different, and both
are different from scored-zero, of which the corpus holds 672,931.
It emits its own check member rather than reusing reference_allele. That check compares an
authored ref against the reference sequence and belongs to enrich; letting an Atlas outage
write a skip against it would make one registry's availability speak for another's question
(@one-registrys-outage-may-not-speak-for-another). Two sources, two checks, side by side.
VariantImpactError ¶
Bases: RuntimeError
The check could not run in a way the caller must see. Never raised for an absent answer.
ImpactFinding
dataclass
¶
One thing the Atlas said that the local artifact could not have said.
kind is the remediation key, not the emission site (@warning-code-names-the-finding): a
reader grouping by it gets one bucket per thing they can do about it.
VariantImpactResult
dataclass
¶
VariantImpactResult(
findings: list[ImpactFinding] = list(),
decided: list[str] = list(),
refined: list[str] = list(),
unanswered: list[tuple[str, str]] = list(),
straddling: list[str] = list(),
warnings: list[str] = list(),
mode: str = "best_effort",
threshold: float | None = None,
dataset: str | None = None,
)
What was decided, what was refined, and what could not be asked — the tri-state, reportable.
subjects
property
¶
The denominator the attestation publishes: every variant this pass had an opinion about.
Distinct variants, not len(decided) + len(unanswered), because a variant can be in
both: the snapshot decides it, and then a threshold falling inside its knot leaves it
unresolved. A three-variant module reported four subjects until this counted a set — an
attestation whose denominator exceeds the rows it was computed over is worse than no
denominator, since it reads as coverage nobody had.
LocalScore
dataclass
¶
One variant's answer from the snapshot, with the ambiguity the knot table records.
phred_lo/phred_hi are the interval, never a midpoint: where they bracket a caller's
threshold the local data cannot decide, and saying so is the whole point of publishing them.
above ¶
Only meaningful when the knot does not straddle — the caller checks that first.
load_module_variants ¶
The module's authored variants, or an empty list when it has none.
An empty list is not an error: a module with no variants.csv has posed no question the Atlas
could answer, and the caller returns without attesting rather than mining a nonce for a check
that did not apply.
Source code in enricher/src/just_dna_enricher/alphagenome_check.py
artifact_contig ¶
A module's chrom in the spelling the AVI artifact uses.
Two conventions meet here and neither is wrong. VariantRow normalizes through
vrs.normalize_chrom, which strips the prefix and stores 22; AlphaGenome ships UCSC-style
chr22, and its tabix index is keyed that way. Joining the module's spelling straight onto the
snapshot silently matches nothing — every row comes back absent_from_snapshot, which reads
exactly like an artifact that does not cover the variant, and is the kind of wrong answer this
tier's three-valued algebra exists to make impossible.
One direction only, and at the boundary: the snapshot's spelling is the source's, so the conversion belongs to the reader that crosses into it rather than to either model.
Delegated to atlas_client.wire_contig rather than duplicated: the RPC boundary needs the same
conversion, and two copies of one normalizer is the defect @one-normalizer-two-spellings names.
Kept as a named function here because the snapshot join is a different boundary from the wire,
and a reader of this module should find the rule where the join is.
Source code in enricher/src/just_dna_enricher/alphagenome_check.py
read_local_scores ¶
Look the module's variants up in the snapshot, keyed by _label.
A variant with no row is absent from the result, never present with a zero: the artifact
covers ~95% of the assembly and writes 672,931 genuine zeros, so the two must stay
distinguishable (@unreachable-not-absent).
Source code in enricher/src/just_dna_enricher/alphagenome_check.py
263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 | |
check_variant_impact ¶
check_variant_impact(
spec_dir: Path,
*,
reference: Path | None = None,
client=None,
threshold: float | None = None,
mode: str = "best_effort",
offline: bool = False,
write: bool = True,
refinement_cap: int = DEFAULT_REFINEMENT_CAP,
) -> VariantImpactResult
Compare a module's variants against AlphaGenome's AVI scores. Reports, never repairs.
threshold is what makes the network leg possible at all: without one there is no question the
local artifact cannot answer, so the pass stays entirely offline. With one, the knot table says
— before any request — which variants sit inside an interval spanning it, and only those are
asked about.
mode is carried for the report and is not a severity ladder here. A model's score
disagreeing with an author's expectation is not a strict matter; what strict still refuses is
structural, and that refuses in best_effort too.
Source code in enricher/src/just_dna_enricher/alphagenome_check.py
406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 | |
threshold_is_safe ¶
Whether any knot spans threshold, and how many rows sit behind the ones that do.
The offline half of this module, and the one a caller should reach for first: it costs a read of
466 KB and answers "can I trust a cut here at all" without touching the network. True means
every row in the corpus is decided at this threshold — which, genome-wide, is every integer from
1 to 50 except 3.
Source code in enricher/src/just_dna_enricher/alphagenome_check.py
score_to_threshold ¶
A raw_score as the integer the snapshot stores, for a caller comparing in the right domain.
Exposed because the alternative is what every consumer does otherwise: divide raw_score_e5 by
1e5 and compare floats, which disagrees with the printed value on 53% of rows. Comparing in the
integer domain has no such failure mode.