Skip to content

just_dna_enricher.clinvar_build

just_dna_enricher.clinvar_build

ClinVar reference-snapshot builder — the [dev] half of the ClinVar link.

Turns the NCBI ClinVar GRCh38 VCF (clinvar.vcf.gz, ~200 MB gz, 4.4M records) into a schema-shaped, per-chromosome parquet snapshot the resolver reads (clinvar/data/*.parquet) — the same on-disk layout as the Ensembl snapshot, so one DuckDB view shape serves both. One row per ALT allele; clin_sig is normalized into vocab.VALID_CLIN_SIG while clin_sig_raw keeps the verbatim CLNSIG (lossless, auditable) — the fold itself lives in clin_sig.py since 0.7, because a second source reporting a significance must reach the same map rather than grow its own. A release.json beside the parquet records the provenance (clinvar_file_date, source_url, source_sha256, record_count, ...) that will feed GenePanelSpec.reference/reference_sha256 when RM4 lands.

Builder-only (just-dna-enricher[dev]): polars is a guarded import — the runtime resolver path (clinvar.lookup_loci, duckdb) never needs it, keeping the core install polars-free (Goal 2). The VCF-parsing idioms (_parse_info, _genes, _norm_chrom, the RS=→rs{n} rule, the ACGT/length filter) are the ones proven in just-dna-lite's v1_port.clinvar, generalized here to keep every record (this is a resolution reference, not a pathogenic gene panel).

ClinVarBuildError

Bases: RuntimeError

A ClinVar artifact could not be built from the file given.

ClinVarUnavailable

Bases: ClinVarBuildError

A ClinVar download did not answer — the fetch failed, not the file.

A subclass, so a caller catching ClinVarBuildError catches this too and one that wants to tell the source was unreachable from the bytes were unreadable can ask for it by name. That split is why it is not a flat sibling (@client-exception-contract): retrying is the right response to one and pointless for the other.

BuildResult dataclass

BuildResult(
    out_dir: Path,
    parquet_files: list[Path],
    record_count: int,
    clinvar_file_date: str | None,
    source_sha256: str | None,
    chromosomes: list[str] = list(),
    skipped_non_acgt: int = 0,
    skipped_too_long: int = 0,
    skipped_bad_chrom: int = 0,
)

Outcome of a snapshot build (paths + provenance + skip stats).

CitationsResult dataclass

CitationsResult(
    row_count: int,
    source_sha256: str | None = None,
    release_updated: bool = False,
    unusable_citations: int = 0,
)

What a citations build produced, and whether the snapshot now says so.

chrom_parquet_name

chrom_parquet_name(chrom: str) -> str

The snapshot's filename for one chromosome — clinvar-chrMT.parquet and its 24 siblings.

A function rather than an f-string at the write site, because since RM171 there is a second reader: the derived MITOMAP-miss lane opens the chrMT parquet by name, and a naming convention two modules spell independently is one rename away from a lane that silently finds nothing.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def chrom_parquet_name(chrom: str) -> str:
    """The snapshot's filename for one chromosome — `clinvar-chrMT.parquet` and its 24 siblings.

    A function rather than an f-string at the write site, because since RM171 there is a **second**
    reader: the derived MITOMAP-miss lane opens the chrMT parquet by name, and a naming convention
    two modules spell independently is one rename away from a lane that silently finds nothing.
    """
    return f"clinvar-chr{chrom}.parquet"

review_stars

review_stars(revstat: str | None) -> int | None

One CLNREVSTAT wording → ClinVar's 0-to-4 gold-star rating, or None for no wording at all.

Public since 0.6 (RM25), and tri-state since 0.6. It is the one place this workspace holds ClinVar's review-status convention — Principle 2 keeps it out of the schema tier entirely, which is why ClinicalAssertionRow stores the rating as a column instead of deriving it from the prose beside it — so a second reader of that convention must reach this function rather than grow a copy.

None and 0 are different answers and the distinction is the whole point of the table this now feeds: 0 is the rating ClinVar gives a submission with no assertion criteria, while None means the record states no review status for anything to be rated. An unrecognized wording is also None, not 0 — a wording this release does not model is an unknown, and answering 0 would record a definite rating where none was read. Withhold rather than guess, as everywhere else.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def review_stars(revstat: str | None) -> int | None:
    """One CLNREVSTAT wording → ClinVar's 0-to-4 gold-star rating, or `None` for no wording at all.

    **Public since 0.6 (RM25), and tri-state since 0.6.** It is the one place this workspace holds
    ClinVar's review-status convention — Principle 2 keeps it out of the schema tier entirely, which
    is why `ClinicalAssertionRow` stores the rating as a column instead of deriving it from the prose
    beside it — so a second reader of that convention must reach this function rather than grow a copy.

    `None` and `0` are different answers and the distinction is the whole point of the table this now
    feeds: `0` is the *rating* ClinVar gives a submission with no assertion criteria, while `None`
    means the record states no review status for anything to be rated. An **unrecognized** wording is
    also `None`, not `0` — a wording this release does not model is an unknown, and answering `0` would
    record a definite rating where none was read. Withhold rather than guess, as everywhere else.
    """
    text = (revstat or "").strip()
    if not text:
        return None
    return _REVIEW_STARS.get(text)

file_date_from_header

file_date_from_header(header: str) -> str | None

The ##fileDate=YYYY-MM-DD value from VCF header text, or None when it states none.

Takes text rather than a path so a caller holding a decompressed prefix of the file can ask the same question as one holding the whole thing. Stops at the first data line, so handing it an entire VCF costs nothing beyond the header.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def file_date_from_header(header: str) -> str | None:
    """The ``##fileDate=YYYY-MM-DD`` value from VCF header text, or `None` when it states none.

    Takes text rather than a path so a caller holding a decompressed *prefix* of the file can ask the
    same question as one holding the whole thing. Stops at the first data line, so handing it an
    entire VCF costs nothing beyond the header.
    """
    for line in header.splitlines():
        if not line.startswith("#"):
            break
        if line.startswith(FILE_DATE_HEADER):
            return line.split("=", 1)[1].strip() or None
    return None

download_clinvar_vcf

download_clinvar_vcf(
    dest: Path, url: str = DEFAULT_CLINVAR_URL
) -> Path

Stream the NCBI ClinVar VCF to dest (atomic .part rename), returning the path.

Provisioning-only — the resolver never calls this; it reads a prebuilt cache.

This is the download RM187 was filed for. It raised httpx's own exception while caches._rebuild_clinvar catches ClinVarBuildError, so when NCBI closed the connection 180,927,542 bytes into a 193,427,450-byte body on 2026-09-03 the lane could not report built=False — and the traceback escaped rebuild_lane too, which in a full cache rebuild takes every later lane with it. It also had no retry, on the largest single request this tier makes. Both are net.stream_to_file's job now.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def download_clinvar_vcf(dest: Path, url: str = DEFAULT_CLINVAR_URL) -> Path:
    """Stream the NCBI ClinVar VCF to ``dest`` (atomic ``.part`` rename), returning the path.

    Provisioning-only — the resolver never calls this; it reads a prebuilt cache.

    **This is the download RM187 was filed for.** It raised `httpx`'s own exception while
    `caches._rebuild_clinvar` catches `ClinVarBuildError`, so when NCBI closed the connection
    180,927,542 bytes into a 193,427,450-byte body on 2026-09-03 the lane could not report
    `built=False` — and the traceback escaped `rebuild_lane` too, which in a full `cache rebuild`
    takes every later lane with it. It also had no retry, on the largest single request this tier
    makes. Both are `net.stream_to_file`'s job now.
    """
    return stream_to_file(
        dest,
        url,
        error_cls=ClinVarUnavailable,
        what="the ClinVar VCF",
        remedy="Pass --source clinvar=<vcf> to build from a copy you already hold.",
    ).path

download_var_citations

download_var_citations(
    dest: Path, url: str = DEFAULT_CITATIONS_URL
) -> tuple[Path, str]

Stream ClinVar's var_citations.txt to dest, the same way the VCF is fetched.

A separate download because ClinVar publishes it separately — the VCF carries no PMIDs. That absence is why a drafted gene panel could not compile: studies.csv is mandatory ("grounding evidence is mandatory") and the provider had nothing to ground rows with.

Returns (path, sha256). The hash was already being computed and then only logged; now that the citations table is published with the snapshot, its provenance has to be recordable — otherwise the artifact carries two ClinVar releases and release.json describes one of them.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def download_var_citations(dest: Path, url: str = DEFAULT_CITATIONS_URL) -> tuple[Path, str]:
    """Stream ClinVar's `var_citations.txt` to `dest`, the same way the VCF is fetched.

    A *separate* download because ClinVar publishes it separately — the VCF carries no PMIDs. That
    absence is why a drafted gene panel could not compile: `studies.csv` is mandatory ("grounding
    evidence is mandatory") and the provider had nothing to ground rows with.

    Returns `(path, sha256)`. The hash was already being computed and then only logged; now that the
    citations table is *published* with the snapshot, its provenance has to be recordable — otherwise
    the artifact carries two ClinVar releases and `release.json` describes one of them.
    """
    streamed = stream_to_file(
        dest,
        url,
        error_cls=ClinVarUnavailable,
        what="the ClinVar citations table",
    )
    return streamed.path, streamed.sha256

build_citations

build_citations(
    citations_txt: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_CITATIONS_URL,
    source_sha256: str | None = None,
) -> CitationsResult

var_citations.txt → out_dir/citations/citations.parquet, PubMed rows only.

Written beside an existing snapshot rather than into it: the per-chromosome parquet keeps its bytes, so adding citations to a cache someone already built does not invalidate it. Only PubMed sources are kept — the format's StudyRow.pmid is a PMID, and a curated module citing a source the schema cannot express would be a row nobody can check.

Sorted by (variation_id, pmid) so a rebuild is byte-identical (Principle 7).

The provenance is merged into release.json, and that matters now that this table is published with the snapshot. ClinVar publishes var_citations.txt on its own cadence, so the records and the citations in one snapshot need not come from the same release — an artifact that ships both while documenting only the VCF is mixed-vintage and silent about it. The block is merged, never overwritten, so the VCF's own provenance survives; a snapshot from a builder that wrote no release.json gets one with just this block, because partial provenance beats none.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def build_citations(
    citations_txt: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_CITATIONS_URL,
    source_sha256: str | None = None,
) -> CitationsResult:
    """`var_citations.txt` → `out_dir/citations/citations.parquet`, PubMed rows only.

    Written **beside** an existing snapshot rather than into it: the per-chromosome parquet keeps its
    bytes, so adding citations to a cache someone already built does not invalidate it. Only PubMed
    sources are kept — the format's `StudyRow.pmid` is a PMID, and a curated module citing a source
    the schema cannot express would be a row nobody can check.

    Sorted by `(variation_id, pmid)` so a rebuild is byte-identical (Principle 7).

    **The provenance is merged into `release.json`, and that matters now that this table is published
    with the snapshot.** ClinVar publishes `var_citations.txt` on its own cadence, so the records and
    the citations in one snapshot need not come from the same release — an artifact that ships both while
    documenting only the VCF is mixed-vintage and silent about it. The block is merged, never
    overwritten, so the VCF's own provenance survives; a snapshot from a builder that wrote no
    `release.json` gets one with just this block, because partial provenance beats none.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the ClinVar citations table; install the publisher/dev "
            "surface with `pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    frame = pl.read_csv(
        citations_txt,
        separator="\t",
        has_header=True,
        infer_schema_length=0,
        truncate_ragged_lines=True,
    )
    columns = {name.lstrip("#").strip(): name for name in frame.columns}
    variation = columns.get("VariationID")
    source = columns.get("citation_source")
    citation = columns.get("citation_id")
    if not (variation and source and citation):
        raise ClinVarBuildError(
            f"unexpected var_citations.txt columns: {frame.columns}. Refusing rather than guessing — "
            f"a silently mis-parsed citations table would ground module rows in the wrong papers."
        )
    published = (
        frame.filter(pl.col(source) == "PubMed")
        .select(
            pl.col(variation).cast(pl.Utf8).alias("variation_id"),
            pl.col(citation).cast(pl.Utf8).alias("pmid"),
        )
        .unique()
        .sort(["variation_id", "pmid"])
    )
    # ClinVar files a handful of ids under `PubMed` that are not PMIDs — 218 of 3,952,341 rows in the
    # 2026-06-27 file, all nine digits (Variation 12606 cites `168335863`; PubMed is at eight). The
    # grammar is the format's own, not a second opinion restated here, so a future widening of
    # `StudyRow.pmid` reaches this filter too. They are dropped rather than stored, because a citation
    # nothing can key on grounds a module row in a paper nobody can look up — and because a snapshot
    # that carries one turns into an unhandled `ValidationError` in the middle of drafting a panel,
    # which is how this was found (one bad row killed a 297-gene draft).
    kept = published.filter(
        pl.col("pmid").map_elements(lambda v: bool(extract_pmids(v or "")), return_dtype=pl.Boolean)
    )
    unusable = published.height - kept.height
    if unusable:
        logger.warning(
            "Dropped %d citation(s) filed under PubMed whose id is not a PMID (kept %d).",
            unusable,
            kept.height,
        )
    target = Path(out_dir) / CITATIONS_DIRNAME
    target.mkdir(parents=True, exist_ok=True)
    kept.write_parquet(target / "citations.parquet")
    logger.info("Wrote %s PubMed citation link(s)", kept.height)
    # Hash the input when the caller did not already have it (the downloader computes it in-stream, so
    # it passes one in and this re-read is skipped). Recording `null` while the bytes sit on disk would
    # be an unknown we chose not to establish — and `source_sha256` is what RM4's `reference_sha256`
    # pins against, so an unhashed citations table is not pinnable.
    digest = source_sha256 or _sha256_file(Path(citations_txt))
    updated = _merge_release_block(
        Path(out_dir),
        "citations",
        {
            "source_url": source_url,
            "source_sha256": digest,
            "row_count": kept.height,
            "built_at": now_utc_iso(),
            "builder_version": _builder_version(),
        },
    )
    return CitationsResult(
        row_count=kept.height,
        source_sha256=digest,
        release_updated=updated,
        unusable_citations=unusable,
    )

build_snapshot

build_snapshot(
    vcf: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_CLINVAR_URL,
) -> BuildResult

Convert a ClinVar VCF into the per-chromosome parquet snapshot + release.json.

One row per ACGT ALT allele (short indels kept, symbolic/structural alleles skipped and counted). Rows are written sorted by (start, ref, alt, rsid, variation_id) per chromosome so a rebuild is byte-identical (Principle 7 idempotency); release.json's built_at is the only per-run-varying byte and lives outside the parquet.

Source code in enricher/src/just_dna_enricher/clinvar_build.py
def build_snapshot(
    vcf: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_CLINVAR_URL,
) -> BuildResult:
    """Convert a ClinVar VCF into the per-chromosome parquet snapshot + ``release.json``.

    One row per ACGT ALT allele (short indels kept, symbolic/structural alleles skipped and counted).
    Rows are written sorted by ``(start, ref, alt, rsid, variation_id)`` per chromosome so a rebuild
    is byte-identical (Principle 7 idempotency); ``release.json``'s ``built_at`` is the only
    per-run-varying byte and lives outside the parquet.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the ClinVar snapshot; install the publisher/dev surface "
            "with `pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    vcf = Path(vcf)
    if not vcf.exists():
        raise FileNotFoundError(f"ClinVar VCF not found at {vcf}")
    out_dir = Path(out_dir)
    data_dir = out_dir / "data"
    data_dir.mkdir(parents=True, exist_ok=True)

    by_chrom: dict[str, list[dict]] = {}
    skipped_non_acgt = skipped_too_long = skipped_bad_chrom = 0

    for chrom_raw, pos_s, vid, ref, alt_field, info in _iter_records(vcf):
        chrom = _norm_chrom(chrom_raw)
        if chrom is None:
            skipped_bad_chrom += 1
            continue
        info_d = _parse_info(info)
        rs = info_d.get("RS")
        rsid = f"rs{rs.split('|')[0]}" if rs else None
        gene_syms = _genes(info_d.get("GENEINFO"))
        clnsig_raw = info_d.get("CLNSIG")
        shared = {
            "chrom": chrom,
            "start": int(pos_s),
            "ref": ref.strip().upper(),
            "rsid": rsid,
            "variation_id": vid or None,
            "allele_id": info_d.get("ALLELEID"),
            "gene": gene_syms[0] if gene_syms else None,
            "genes": "|".join(gene_syms) if gene_syms else None,
            "clin_sig": normalize_clin_sig(clnsig_raw),
            "clin_sig_raw": clnsig_raw,
            "review_status": info_d.get("CLNREVSTAT"),
            "review_stars": review_stars(info_d.get("CLNREVSTAT")),
            "condition": _condition(info_d.get("CLNDN")),
            "molecular_consequence": _molecular_consequence(info_d.get("MC")),
            "variant_type": info_d.get("CLNVC"),
            "origin": info_d.get("ORIGIN"),
        }
        ref_u = shared["ref"]
        for alt in alt_field.split(","):
            alt_u = alt.strip().upper()
            if not (_ACGT_RE.match(ref_u) and _ACGT_RE.match(alt_u)) or ref_u == alt_u:
                skipped_non_acgt += 1
                continue
            if len(ref_u) > MAX_ALLELE_LEN or len(alt_u) > MAX_ALLELE_LEN:
                skipped_too_long += 1
                continue
            by_chrom.setdefault(chrom, []).append({**shared, "alt": alt_u})

    schema = _empty_schema()
    parquet_files: list[Path] = []
    record_count = 0
    written_chroms: list[str] = []
    # Karyotype order for files; unknown-but-valid chroms (shouldn't occur) appended sorted after.
    ordered = [c for c in _CHROM_ORDER if c in by_chrom] + sorted(
        c for c in by_chrom if c not in _CHROM_ORDER
    )
    for chrom in ordered:
        df = pl.DataFrame(by_chrom[chrom], schema=schema).sort(
            ["start", "ref", "alt", "rsid", "variation_id"], nulls_last=True
        )
        path = data_dir / chrom_parquet_name(chrom)
        df.write_parquet(path)
        parquet_files.append(path)
        record_count += df.height
        written_chroms.append(chrom)

    file_date = _read_file_date(vcf)
    source_sha256 = _sha256_file(vcf)
    _write_release_json(
        out_dir,
        clinvar_file_date=file_date,
        source_url=source_url,
        source_sha256=source_sha256,
        record_count=record_count,
    )
    logger.info(
        "Built ClinVar snapshot: %d rows across %d chromosomes → %s "
        "(skipped: non-ACGT %d, too-long %d, off-target chrom %d)",
        record_count,
        len(written_chroms),
        data_dir,
        skipped_non_acgt,
        skipped_too_long,
        skipped_bad_chrom,
    )
    return BuildResult(
        out_dir=out_dir,
        parquet_files=parquet_files,
        record_count=record_count,
        clinvar_file_date=file_date,
        source_sha256=source_sha256,
        chromosomes=written_chroms,
        skipped_non_acgt=skipped_non_acgt,
        skipped_too_long=skipped_too_long,
        skipped_bad_chrom=skipped_bad_chrom,
    )