just_dna_enricher.clinvar_build¶
just_dna_enricher.clinvar_build ¶
ClinVar reference-snapshot builder — the [dev] half of the ClinVar link.
Turns the NCBI ClinVar GRCh38 VCF (clinvar.vcf.gz, ~200 MB gz, 4.4M records) into a
schema-shaped, per-chromosome parquet snapshot the resolver reads (clinvar/data/*.parquet) — the
same on-disk layout as the Ensembl snapshot, so one DuckDB view shape serves both. One row per ALT
allele; clin_sig is normalized into vocab.VALID_CLIN_SIG while clin_sig_raw keeps the verbatim
CLNSIG (lossless, auditable) — the fold itself lives in clin_sig.py since 0.7, because a second
source reporting a significance must reach the same map rather than grow its own. A release.json
beside the parquet records the provenance
(clinvar_file_date, source_url, source_sha256, record_count, ...) that will feed
GenePanelSpec.reference/reference_sha256 when RM4 lands.
Builder-only (just-dna-enricher[dev]): polars is a guarded import — the runtime resolver path
(clinvar.lookup_loci, duckdb) never needs it, keeping the core install polars-free (Goal 2). The
VCF-parsing idioms (_parse_info, _genes, _norm_chrom, the RS=→rs{n} rule, the ACGT/length
filter) are the ones proven in just-dna-lite's v1_port.clinvar, generalized here to keep every
record (this is a resolution reference, not a pathogenic gene panel).
ClinVarBuildError ¶
Bases: RuntimeError
A ClinVar artifact could not be built from the file given.
ClinVarUnavailable ¶
Bases: ClinVarBuildError
A ClinVar download did not answer — the fetch failed, not the file.
A subclass, so a caller catching ClinVarBuildError catches this too and one that wants to
tell the source was unreachable from the bytes were unreadable can ask for it by name. That
split is why it is not a flat sibling (@client-exception-contract): retrying is the right
response to one and pointless for the other.
BuildResult
dataclass
¶
BuildResult(
out_dir: Path,
parquet_files: list[Path],
record_count: int,
clinvar_file_date: str | None,
source_sha256: str | None,
chromosomes: list[str] = list(),
skipped_non_acgt: int = 0,
skipped_too_long: int = 0,
skipped_bad_chrom: int = 0,
)
Outcome of a snapshot build (paths + provenance + skip stats).
CitationsResult
dataclass
¶
CitationsResult(
row_count: int,
source_sha256: str | None = None,
release_updated: bool = False,
unusable_citations: int = 0,
)
What a citations build produced, and whether the snapshot now says so.
chrom_parquet_name ¶
The snapshot's filename for one chromosome — clinvar-chrMT.parquet and its 24 siblings.
A function rather than an f-string at the write site, because since RM171 there is a second reader: the derived MITOMAP-miss lane opens the chrMT parquet by name, and a naming convention two modules spell independently is one rename away from a lane that silently finds nothing.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
review_stars ¶
One CLNREVSTAT wording → ClinVar's 0-to-4 gold-star rating, or None for no wording at all.
Public since 0.6 (RM25), and tri-state since 0.6. It is the one place this workspace holds
ClinVar's review-status convention — Principle 2 keeps it out of the schema tier entirely, which
is why ClinicalAssertionRow stores the rating as a column instead of deriving it from the prose
beside it — so a second reader of that convention must reach this function rather than grow a copy.
None and 0 are different answers and the distinction is the whole point of the table this now
feeds: 0 is the rating ClinVar gives a submission with no assertion criteria, while None
means the record states no review status for anything to be rated. An unrecognized wording is
also None, not 0 — a wording this release does not model is an unknown, and answering 0 would
record a definite rating where none was read. Withhold rather than guess, as everywhere else.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
file_date_from_header ¶
The ##fileDate=YYYY-MM-DD value from VCF header text, or None when it states none.
Takes text rather than a path so a caller holding a decompressed prefix of the file can ask the same question as one holding the whole thing. Stops at the first data line, so handing it an entire VCF costs nothing beyond the header.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
download_clinvar_vcf ¶
Stream the NCBI ClinVar VCF to dest (atomic .part rename), returning the path.
Provisioning-only — the resolver never calls this; it reads a prebuilt cache.
This is the download RM187 was filed for. It raised httpx's own exception while
caches._rebuild_clinvar catches ClinVarBuildError, so when NCBI closed the connection
180,927,542 bytes into a 193,427,450-byte body on 2026-09-03 the lane could not report
built=False — and the traceback escaped rebuild_lane too, which in a full cache rebuild
takes every later lane with it. It also had no retry, on the largest single request this tier
makes. Both are net.stream_to_file's job now.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
download_var_citations ¶
Stream ClinVar's var_citations.txt to dest, the same way the VCF is fetched.
A separate download because ClinVar publishes it separately — the VCF carries no PMIDs. That
absence is why a drafted gene panel could not compile: studies.csv is mandatory ("grounding
evidence is mandatory") and the provider had nothing to ground rows with.
Returns (path, sha256). The hash was already being computed and then only logged; now that the
citations table is published with the snapshot, its provenance has to be recordable — otherwise
the artifact carries two ClinVar releases and release.json describes one of them.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
build_citations ¶
build_citations(
citations_txt: Path,
out_dir: Path,
*,
source_url: str = DEFAULT_CITATIONS_URL,
source_sha256: str | None = None,
) -> CitationsResult
var_citations.txt → out_dir/citations/citations.parquet, PubMed rows only.
Written beside an existing snapshot rather than into it: the per-chromosome parquet keeps its
bytes, so adding citations to a cache someone already built does not invalidate it. Only PubMed
sources are kept — the format's StudyRow.pmid is a PMID, and a curated module citing a source
the schema cannot express would be a row nobody can check.
Sorted by (variation_id, pmid) so a rebuild is byte-identical (Principle 7).
The provenance is merged into release.json, and that matters now that this table is published
with the snapshot. ClinVar publishes var_citations.txt on its own cadence, so the records and
the citations in one snapshot need not come from the same release — an artifact that ships both while
documenting only the VCF is mixed-vintage and silent about it. The block is merged, never
overwritten, so the VCF's own provenance survives; a snapshot from a builder that wrote no
release.json gets one with just this block, because partial provenance beats none.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 | |
build_snapshot ¶
build_snapshot(
vcf: Path,
out_dir: Path,
*,
source_url: str = DEFAULT_CLINVAR_URL,
) -> BuildResult
Convert a ClinVar VCF into the per-chromosome parquet snapshot + release.json.
One row per ACGT ALT allele (short indels kept, symbolic/structural alleles skipped and counted).
Rows are written sorted by (start, ref, alt, rsid, variation_id) per chromosome so a rebuild
is byte-identical (Principle 7 idempotency); release.json's built_at is the only
per-run-varying byte and lives outside the parquet.
Source code in enricher/src/just_dna_enricher/clinvar_build.py
461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 | |