Skip to content

just_dna_enricher.cpic_build

just_dna_enricher.cpic_build

Build the CPIC snapshot ([dev]) — the whole PostgREST database as five parquet tables (RM38).

CPIC is open and unauthenticated, so this cache is not about access. It is about a host not spending one shared per-IP allowance on every caller's request, and about the terms being accepted once, by the operator who built the snapshot, rather than implicitly on each incoming request. That is RM38's argument, and it is the same one for ClinPGx; PharmVar adds a third reason (a personal key) that these two do not have.

The whole database fits, comfortably. Probed live on 2026-08-07: 132 genes, 1,361 alleles, 1,337 allele definitions over 3,120 location values, 2,115 recommendations across 324 drugs, and 112,754 diplotypes — the last being the only large one, and four narrow string columns of it is under a megabyte of zstd parquet. There is no gene filter here for that reason: a snapshot that covers only the genes the operator thought of is a snapshot that silently answers "CPIC has nothing" for the next one.

Verbatim, and mapped at read time. The vocabulary translations (clinicalfunctionalstatus → VALID_FUNCTION_STATUS, classification → VALID_RECOMMENDATION_STRENGTH) are not applied here. They live in cpic.map_function_status / map_classification, which the snapshot client calls on the way out — so a mapping fix reaches an already-built snapshot, and a live answer and a snapshot answer are the same object by construction rather than by inspection. Storing a source's value verbatim is the house rule; the one documented exception is an encoding that lies about its own order (ClinGen's dosage codes), and CPIC's prose does not.

recommendation.phenotypes is a JSON map, so it is flattened one row per gene named. The live client keeps a recommendation only when it names exactly one gene — a row about CYP2C19 and CYP2D6 is not a statement about CYP2C19 alone — so gene_count travels with each flattened row and the reader applies the identical rule. Flattening without it would silently promote multi-gene rows.

Builder-only: polars is a guarded [dev] import, and the runtime pass reads the parquet through DuckDB, which is core. Builder in polars, runtime in duckdb — the house convention.

CpicBuildError

Bases: RuntimeError

The CPIC snapshot could not be built.

build_snapshot

build_snapshot(
    out_dir: Path,
    *,
    endpoint: str = DEFAULT_CPIC_ENDPOINT,
    client: CpicClient | None = None,
) -> CpicBuildResult

Fetch CPIC whole and write data/*.parquet + release.json under out_dir.

built_at is the only per-run byte and it lives in release.json, outside the parquet, exactly as in the ClinVar and constraint builders.

Source code in enricher/src/just_dna_enricher/cpic_build.py
def build_snapshot(
    out_dir: Path, *, endpoint: str = DEFAULT_CPIC_ENDPOINT, client: CpicClient | None = None
) -> CpicBuildResult:
    """Fetch CPIC whole and write `data/*.parquet` + `release.json` under `out_dir`.

    `built_at` is the only per-run byte and it lives in `release.json`, outside the parquet, exactly as
    in the ClinVar and constraint builders.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the CPIC snapshot; install the publisher/dev surface with "
            "`pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    out_dir = Path(out_dir)
    data_dir = out_dir / SNAPSHOT_DATA_DIRNAME
    data_dir.mkdir(parents=True, exist_ok=True)

    owned = client is None
    cpic = client or CpicClient(endpoint)
    try:
        genes = _fetch_all(cpic, "gene", "symbol,chr,ensemblid,hgncid,lookupmethod")
        alleles = _fetch_all(cpic, "allele", "genesymbol,name,activityvalue,clinicalfunctionalstatus")
        diplotypes = _fetch_all(cpic, "diplotype", "genesymbol,diplotype,generesult,totalactivityscore")
        definitions = _fetch_all(cpic, "allele_definition", "id,name,genesymbol")
        locations = _fetch_all(
            cpic,
            "allele_location_value",
            "alleledefinitionid,variantallele,sequence_location(genesymbol,dbsnpid,position)",
        )
        recommendations = _fetch_all(
            cpic,
            "recommendation",
            "drugid,phenotypes,implications,drugrecommendation,classification,population,activityscore",
        )
        drugs = _fetch_all(cpic, "drug", "drugid,name")
    finally:
        if owned:
            cpic.close()

    chrom_by_gene = {(row.get("symbol") or ""): normalize_chrom(row.get("chr")) for row in genes}
    gene_records = _sorted(
        [
            {
                "gene": row.get("symbol"),
                "chrom": chrom_by_gene.get(row.get("symbol") or ""),
                "ensembl_id": row.get("ensemblid") or None,
                "hgnc_id": row.get("hgncid") or None,
                "lookup_method": row.get("lookupmethod") or None,
            }
            for row in genes
            if row.get("symbol")
        ],
        ["gene"],
    )
    allele_records = _sorted(
        [
            {
                "gene": row.get("genesymbol"),
                "allele": row.get("name"),
                "activity_value": _float_or_none(row.get("activityvalue")),
                "clinical_function_status": row.get("clinicalfunctionalstatus") or None,
            }
            for row in alleles
            if row.get("genesymbol") and row.get("name")
        ],
        ["gene", "allele"],
    )
    diplotype_records = _sorted(
        [
            {
                "gene": row.get("genesymbol"),
                "diplotype": row.get("diplotype"),
                "phenotype": row.get("generesult") or None,
                "activity_score": row.get("totalactivityscore") or None,
            }
            for row in diplotypes
            if row.get("genesymbol") and row.get("diplotype")
        ],
        ["gene", "diplotype"],
    )

    # The `allele_definition → allele_location_value → sequence_location` join, flattened. The gene's
    # chromosome comes from the `gene` table, which is where CPIC keeps it (see `cpic.py`) — the
    # location row names only the symbol.
    allele_by_id = {row["id"]: row for row in definitions}
    defining_records: list[dict] = []
    for row in locations:
        definition = allele_by_id.get(row.get("alleledefinitionid"))
        if definition is None:
            continue
        location = row.get("sequence_location") or {}
        gene = location.get("genesymbol") or definition.get("genesymbol")
        defining_records.append(
            {
                "gene": gene,
                "allele": definition.get("name"),
                "rsid": location.get("dbsnpid") or None,
                "chrom": chrom_by_gene.get(gene or ""),
                "start": location.get("position"),
                "variant_allele": (row.get("variantallele") or "").strip().upper() or None,
            }
        )
    defining_records = _sorted(defining_records, ["gene", "allele", "start", "variant_allele"])

    drug_by_id = {row.get("drugid"): (row.get("name") or "").strip().lower() for row in drugs}
    recommendation_records: list[dict] = []
    for row in recommendations:
        phenotypes = row.get("phenotypes") or {}
        implications = row.get("implications") or {}
        scores = row.get("activityscore") or {}
        drug = drug_by_id.get(row.get("drugid"))
        if not drug:
            continue
        for gene, phenotype in phenotypes.items():
            if not phenotype:
                continue
            recommendation_records.append(
                {
                    "gene": gene,
                    "phenotype": phenotype,
                    "drug": drug,
                    "population": (row.get("population") or "").strip(),
                    "classification": row.get("classification") or None,
                    "recommendation": (row.get("drugrecommendation") or "").strip() or None,
                    "implication": (implications.get(gene) or "").strip() or None,
                    "activity_score": _score_text(scores.get(gene)),
                    "gene_count": len(phenotypes),
                }
            )
    recommendation_records = _sorted(recommendation_records, ["gene", "drug", "population", "phenotype"])

    schemas = _schemas()
    by_file = {
        GENES_PARQUET: gene_records,
        ALLELES_PARQUET: allele_records,
        DIPLOTYPES_PARQUET: diplotype_records,
        DEFINING_PARQUET: defining_records,
        RECOMMENDATIONS_PARQUET: recommendation_records,
    }
    counts: dict[str, int] = {}
    for name, records in by_file.items():
        pl.DataFrame(records, schema=schemas[name]).write_parquet(data_dir / name, compression="zstd")
        counts[name] = len(records)

    digest = _content_digest(by_file)
    dataset = _dataset_label(digest)
    release = {
        "source_url": endpoint,
        # CPIC publishes no release id, so the snapshot is identified by *content*: a digest over the
        # rows it holds. Deterministic (the records are sorted), and unlike a build date it changes
        # exactly when the data does — which is what makes `dataset` a fact a module can be compared on.
        "content_sha256": digest,
        "dataset": dataset,
        "row_counts": counts,
        "gene_count": len(gene_records),
        "built_at": now_utc_iso(),
        "builder_version": _builder_version(),
    }
    atomic_write_text((out_dir / RELEASE_FILENAME), json.dumps(release, indent=2, sort_keys=True) + "\n")
    logger.info(
        "CPIC snapshot: %d genes, %d alleles, %d diplotypes, %d defining variants, %d recommendations → %s",
        len(gene_records),
        len(allele_records),
        len(diplotype_records),
        len(defining_records),
        len(recommendation_records),
        data_dir,
    )
    return CpicBuildResult(out_dir=out_dir, row_counts=counts, gene_count=len(gene_records), dataset=dataset)