just_dna_enricher.clinvar¶
just_dna_enricher.clinvar ¶
ClinVar resolver link — the core half of the ClinVar reference (no polars, duckdb only).
lookup_loci mirrors resolver.lookup_loci exactly (same signature, same determinism discipline) so
enrich() treats the ClinVar cache and the Ensembl cache identically — a rsid → [loci] /
pos → rsid batch lookup over an injected parquet directory. The resolver link reads only
chrom/start/ref/alt; the snapshot's annotation columns (clin_sig, gene, condition, ...) stay
out of resolution.csv, because they are annotation rather than resolution facts (orthogonal axes,
Principle 5). Inject-only: the caller provisions the reference (download.ensure_clinvar_snapshot or
a local cache); this never fetches.
lookup_clin_sig is the second, separate reader over the same view: it serves the clinical
cross-check in clinical.py, which compares a module's authored clin_sig against ClinVar's. Keeping
it here (data access) and the comparison there (judgement) is the same split sequences.py draws
between reading bases and deciding what a disagreement means.
The snapshot ships as parquet only (built by clinvar_build, [dev]); there is no prebuilt
.duckdb for ClinVar, so this connects a read_parquet view directly.
ClinVarReferenceError ¶
Bases: FileNotFoundError
Raised when a provided ClinVar reference has no usable parquet files.
clinvar_dataset_label ¶
Which ClinVar release a snapshot carries, as a SourceRow.dataset value (RM4).
One function, called by both sides. clinvar_draft writes this label onto the licence row it
records for the rows it copied out, and clinical.tautology_reason recomputes it from whatever
snapshot the check is about to read and compares. Shared rather than mirrored for the reason
hosting_verdict is shared: two spellings of one convention drift, and this drift would be
silent — a writer and a reader disagreeing about the label do not fail, they simply never match,
and a check quietly stops being skippable.
clinvar_file_date first, because that is what dataset is for ("which release the data came
from"). A snapshot built from a VCF whose header stated no file date still has the digest of the
bytes it was built from, which names the release exactly, so that is the fallback rather than a
gap. None when the snapshot cannot state its release at all — an unreadable or absent
release.json is an unknown, and an unknown is withheld rather than written as a label something
could match.
Source code in enricher/src/just_dna_enricher/clinvar.py
lookup_loci ¶
lookup_loci(
reference: Path,
rsids: list[str],
positions: list[
tuple[
str | None, int | None, str | None, str | None
]
],
) -> tuple[
dict[str, list[dict]], dict[tuple, list[str]], list[str]
]
(rsid -> [loci], (chrom,start,ref,alt) -> [candidate rsids], warnings) over ClinVar.
Signature-identical to resolver.lookup_loci, so the enrich chain calls either link with the same
code. The reverse map reuses resolver._lookup_rsid_candidates (allele-aware, all candidates) over
the clinvar view's rsid column — one implementation, no drift. Inject-only; never fetches.
Source code in enricher/src/just_dna_enricher/clinvar.py
lookup_clin_sig ¶
lookup_clin_sig(
reference: Path,
alleles: list[tuple[str, int, str, str]],
) -> dict[tuple[str, int, str, str], list[dict]]
(chrom, start, ref, alt) -> [{clin_sig, clin_sig_raw, review_status, review_stars, ...}].
Allele-exact by construction, and that is the whole point: an rsID is a position/multi-allelic
tag, so one rsID at one locus can carry a pathogenic, a benign and an uncertain allele
(rs33922842 in HBB does exactly that). Looking up clinical significance by rsID would therefore
manufacture disagreements out of ClinVar agreeing with itself.
A list rather than a single record because ClinVar genuinely holds several submissions for one
allele — different conditions, different review levels. Ordered by descending review_stars then
variation_id so the best-reviewed record leads and the order is stable (Principle 7).
Source code in enricher/src/just_dna_enricher/clinvar.py
select_by_gene ¶
select_by_gene(
reference: Path,
genes: list[str],
*,
clin_sig: frozenset[str] | None = None,
min_review_stars: int = 0,
) -> list[dict]
Rows for a gene panel: the third reader over the same view, serving the drafting provider.
Kept here with the other two because this module owns the snapshot's schema; deciding what a row
means for a module stays in clinvar_draft, the same data-access/judgement split lookup_loci
and clinical.py already draw.
min_review_stars is the honest quality dial for a panel — a 0-star "no assertion criteria"
submission and a 3-star expert-panel review are not the same evidence, and a panel that mixes them
silently is worse than one that says which floor it drew. Ordered deterministically so a re-draft
appends nothing new (Principle 7): the emitted row order is authored order, and authored order is
digest-visible.
Source code in enricher/src/just_dna_enricher/clinvar.py
citations_for ¶
variation_id -> [pmid, ...] from the optional citations table, or {} when absent.
Optional on purpose: a snapshot built before clinvar citations existed is still a perfectly good
resolver and cross-check reference, and refusing to read it would strand every existing cache. An
absent table means "no citations available", which the caller reports — never "this variant has no
literature", which would be a claim about ClinVar rather than about the file on disk.
Ordered by pmid so a re-draft emits the same rows in the same order (Principle 7).