Skip to content

just_dna_enricher.clinpgx

just_dna_enricher.clinpgx

enrich-clinpgx — cross-check drug-response annotations against the ClinPGx snapshot (0.5).

Pass 6, and the offline-capable half of the pharmacogenomics work. pgx.py asks the two nomenclature authorities about star alleles over the network; this one asks ClinPGx about clinical annotations — which variant, which drug, at what evidence level — from a local snapshot, exactly as the ClinVar cross-check does.

The licence comes out of the snapshot, not out of a table. clinpgx_build extracts the LICENSE.txt ClinPGx bundles inside summaryAnnotations.zip and records its sha256 in release.json; this pass stamps that hash onto the SourceRow it emits. The recorded terms are therefore provably the ones shipped with the recorded data, which is the property a static source→licence map cannot offer — and two halves of such a map went stale inside one release.

What it checks, and what it deliberately does not. It compares the authored evidence_level against ClinPGx's current one for the same (variant, drug, genotype). That is a currency check, not an opinion: an evidence level is ClinPGx's own metadata about its own annotation, so ClinPGx is definitionally right about it and the severity follows the mode ladder. This is unlike the clin_sig and allele-function checks, which compare two expert judgements and therefore warn in both modes — here a disagreement means the module is stale, not that two panels differ.

Only for a row compared with the annotation it cites (RM297). A cited id the snapshot does not hold, and a row citing none, are reported in both modes as withheld and never refused: nothing per row says the row came from ClinPGx at all, which is RM298's.

It does not check the annotation text, and it does not write pharm_variants.csv. Those tables are authored _TABLE_KINDS; a network pass filling them would blur the authored/derived line.

ClinPgxEnrichmentError

Bases: RuntimeError

Raised in strict mode when an authored annotation disagrees with the snapshot.

EvidenceConflict dataclass

EvidenceConflict(
    rsid: str | None,
    drug: str,
    genotype: str | None,
    authored: str,
    reported: str,
)

An authored evidence level ClinPGx's own record does not support.

WithheldLevel dataclass

WithheldLevel(
    code: str,
    rsid: str,
    drug: str,
    genotype: str | None,
    authored: str,
    annotation_id: str | None,
    held: tuple[str, ...],
)

An authored evidence level this pass could not hold against the record it cites (RM297).

held is what the snapshot carries for the row's (rsid, drug, genotype) — narrowed to its category where the row states one — sorted, and empty when the snapshot holds nothing there.

load_snapshot

load_snapshot(reference: Path) -> tuple[list[dict], dict]

Read the snapshot parquet + its release.json. Returns (rows, release).

Read with duckdb, not polars: the convention is builder in polars, runtime pass in duckdb, which is what keeps the enricher's declared runtime dependency set honest about what it actually needs. (Not because polars would be missing — the compiler requires it unconditionally and the enricher requires the compiler, so it is installed either way. That justification was checked in the 0.5 audit and is false; the convention stands on its own.) clinvar.py reads its snapshot the same way.

The column list is the file's, not a subset chosen here. It was hand-written and omitted gene and annotation_text, both of which clinpgx_build writes for every record — and the omission then travelled outward as a claim about the source: clinpgx_draft refused --gene saying "the ClinPGx annotation snapshot carries no gene column" (it does, on 15,331 of 16,087 rows) and synthesized a conclusion restating the row's own key while the published sentence sat unread in annotation_text (16,087 of 16,087). A hand-kept projection is the same shape as a hand-kept column list anywhere else in this workspace, with the extra cost that what it drops is invisible to every reader downstream, who sees only a dict that does not have the key.

Source code in enricher/src/just_dna_enricher/clinpgx.py
def load_snapshot(reference: Path) -> tuple[list[dict], dict]:
    """Read the snapshot parquet + its `release.json`. Returns `(rows, release)`.

    Read with **duckdb**, not polars: the convention is builder in polars, runtime pass in duckdb,
    which is what keeps the enricher's declared runtime dependency set honest about what it actually
    needs. (Not because polars would be *missing* — the compiler requires it unconditionally and the
    enricher requires the compiler, so it is installed either way. That justification was checked in
    the 0.5 audit and is false; the convention stands on its own.) `clinvar.py` reads its snapshot the
    same way.

    **The column list is the file's, not a subset chosen here.** It was hand-written and omitted
    `gene` and `annotation_text`, both of which `clinpgx_build` writes for every record — and the
    omission then travelled outward as a claim about the *source*: `clinpgx_draft` refused `--gene`
    saying "the ClinPGx annotation snapshot carries no gene column" (it does, on 15,331 of 16,087
    rows) and synthesized a `conclusion` restating the row's own key while the published sentence sat
    unread in `annotation_text` (16,087 of 16,087). A hand-kept projection is the same shape as a
    hand-kept column list anywhere else in this workspace, with the extra cost that what it drops is
    invisible to every reader downstream, who sees only a dict that does not have the key.
    """
    reference = Path(reference)
    parquet = reference / SNAPSHOT_DATA_DIRNAME / "annotations.parquet"
    if not parquet.is_file():
        raise ClinPgxEnrichmentError(f"no ClinPGx snapshot at {parquet}")
    release_path = reference / RELEASE_FILENAME
    release = json.loads(release_path.read_text()) if release_path.is_file() else {}
    # Single-quote-escape defensively; the path comes from our own cache resolution, not user input.
    pattern = str(parquet).replace("'", "''")
    con = duckdb.connect(":memory:")
    try:
        cursor = con.execute(f"SELECT * FROM read_parquet('{pattern}')")
        columns = [d[0] for d in cursor.description]
        return [dict(zip(columns, row, strict=True)) for row in cursor.fetchall()], release
    finally:
        con.close()

enrich_clinpgx

enrich_clinpgx(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    declared_use: str = "unstated",
    snapshot: Path | None = None,
    offline: bool = False,
    download: bool = True,
    write: bool = True,
) -> ClinPgxResult

Cross-check pharm_variants.csv against the ClinPGx snapshot and record the terms.

Offline-capable by design: the snapshot is local, so unlike pgx.py this pass needs no network to read, and the declared-use gate still applies — the terms were accepted when the snapshot was built, and using it without a declaration is the same act.

Finding a snapshot, though, was the gap (RM38). The builder shipped a release ahead of any plumbing: there was no locations resolver and no ensure_*, so with no explicit snapshot= this pass skipped itself and said so — which on a hosted deployment is the check simply never running. The chain now is: explicit path → resolved cache → provisioned from HuggingFace. The last step is what offline turns off, and it is why this pass grew an offline parameter it was previously right not to have: it now has something to decline to do.

Unlike pgx, there is no live fallback and there cannot be — api.pharmgkb.org was retired on 2026-07-20 — which is exactly why provisioning is automatic here and a fallback elsewhere.

Source code in enricher/src/just_dna_enricher/clinpgx.py
def enrich_clinpgx(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    declared_use: str = "unstated",
    snapshot: Path | None = None,
    offline: bool = False,
    download: bool = True,
    write: bool = True,
) -> ClinPgxResult:
    """Cross-check `pharm_variants.csv` against the ClinPGx snapshot and record the terms.

    Offline-capable by design: the snapshot is local, so unlike `pgx.py` this pass needs no network to
    *read*, and the declared-use gate still applies — the terms were accepted when the snapshot was
    built, and using it without a declaration is the same act.

    **Finding a snapshot, though, was the gap (RM38).** The builder shipped a release ahead of any
    plumbing: there was no `locations` resolver and no `ensure_*`, so with no explicit `snapshot=` this
    pass skipped itself and said so — which on a hosted deployment is the check simply never running.
    The chain now is: explicit path → resolved cache → provisioned from HuggingFace. The last step is
    what `offline` turns off, and it is why this pass grew an `offline` parameter it was previously
    right not to have: it now has something to decline to do.

    Unlike `pgx`, there is no live fallback and there cannot be — `api.pharmgkb.org` was retired on
    2026-07-20 — which is exactly why provisioning is automatic here and a fallback elsewhere.
    """
    spec_dir = Path(spec_dir)
    result = ClinPgxResult(mode=mode, declared_use=declared_use)

    pharm_path = spec_dir / "pharm_variants.csv"
    if not pharm_path.exists():
        result.warnings.append("ClinPGx cross-check skipped: the module carries no pharm_variants.csv.")
        # **Not attested, and this is the one skip that must not be.** The others say "this check
        # applies to your module and did not run"; this one says the check does not apply at all, and
        # a module with no PGx table has no claim for it to have an opinion about. Recording it would
        # mine a nonce and create a `verification.json` on a module that never asked for one — a
        # file-creating side effect on a path that otherwise does nothing. `nothing_to_check` stays
        # reachable for a table that is present with no row in scope, which is a real answer.
        result.not_checked = "nothing_to_check"
        return result
    authored, errors, _ = load_csv_rows(pharm_path, PharmVariantRow, "pharm_variants.csv")
    if errors:
        raise ClinPgxEnrichmentError(f"pharm_variants.csv is invalid: {errors[0]}")

    declared_use, declared_from = effective_declared_use(spec_dir, CLINPGX_TERMS, declared_use)  # S105
    reason = check_declared_use(CLINPGX_TERMS, declared_use)  # raises on `commercial`
    if declared_from is not None:
        result.warnings.append(
            f"clinpgx: use {declared_use!r} read from {declared_from}, recorded by an earlier run; pass --use "
            f"to declare otherwise."
        )
    if reason is not None:
        result.warnings.append(reason)
        logger.warning("%s", reason)
        # `not_permitted`, never `offline`: what clears this is a *declaration*, not egress, and a
        # reader sent looking for a network problem would find none.
        result.not_checked = "not_permitted"
        return _attest(result, spec_dir, write=write)

    reference = Path(snapshot) if snapshot is not None else resolve_clinpgx_reference()
    if reference is None and not offline and download:
        # Same shape as `enrich()`: provision only when the local resolve came back empty, and degrade
        # to the honest skip on failure rather than raising — a snapshot that cannot be fetched is not
        # a finding about the module.
        try:
            reference = ensure_clinpgx_snapshot()
        except Exception as exc:
            logger.info("Could not provision the ClinPGx snapshot (%s).", exc)
            reference = None
    if reference is None:
        result.warnings.append(
            "ClinPGx cross-check skipped: no snapshot available"
            + (" and --offline" if offline else "")
            + ". Build one with `just-dna-enricher clinpgx build --out <dir>`, or point at it with "
            "$JUST_DNA_CLINPGX_CACHE."
        )
        result.not_checked = "offline" if offline else "no_reference"
        return _attest(result, spec_dir, write=write)

    snapshot_rows, release = load_snapshot(reference)
    result.dataset = release.get("dataset")

    # **The third instance of RM4's tautology, and the one whose marker was already being written**
    # (RM73). `clinpgx_draft` copies `evidence_level` straight out of this snapshot and stamps the
    # release onto the licence row; this check then compares that very column against the same
    # snapshot. The label had been recorded since 0.5.1 and read by nothing. Same conjunction as
    # `clinical.tautology_reason`: this release **and** an unmoved digest, either half missing runs
    # the check in full.
    recorded = [
        row
        for row in read_sources_file(spec_dir)
        if row.source == CLINPGX_TERMS.source and row.layer == "annotation"
    ]
    if (
        result.dataset
        and recorded
        and all((row.dataset or "").strip() == result.dataset for row in recorded)
        and drafted_unchanged(spec_dir, CLINPGX_TERMS.source, recorded)
    ):
        note = (
            f"ClinPGx evidence-level check not run: this module's licence row records that these "
            f"annotations were drafted from {result.dataset}, the snapshot this check reads, and "
            f"every authored evidence_level still hashes to what the drafter wrote — so each is a "
            f"copy of the value it would be compared against. Edit any of them and it runs again."
        )
        result.warnings.append(note)
        logger.info("%s", note)
        result.not_checked = "tautology"
        return _attest(result, spec_dir, write=write)

    # Index the snapshot. Drugs are `;`-separated upstream, so the cell is exploded — a row about
    # "simvastatin;simvastatin acid" is about both.
    #
    # **(rsid, drug, genotype) is NOT unique and keying on it alone produces false conflicts.** One
    # variant and one drug carry several distinct annotations: rs4149056 + simvastatin is
    # Metabolism/PK at 1A, Efficacy at 3 *and* Toxicity at 1A, each with the same three genotypes.
    # Comparing an authored Efficacy row against whichever annotation happened to be indexed first
    # reported all three of this example's correctly-authored levels as stale. So the annotation id
    # is the primary key, the category is the secondary one, and the bare triple is a last resort
    # that reports ambiguity instead of guessing.
    by_annotation: dict[tuple[str, str | None], str] = {}
    known_annotations: set[str] = set()
    by_category: dict[tuple[str, str, str | None, str | None], str] = {}
    by_triple: dict[tuple[str, str, str | None], set[str]] = {}
    for row in snapshot_rows:
        rsid, level = row["rsid"], row["evidence_level"]
        if not rsid or not level:
            continue
        genotype = _normalize_genotype(row["genotype"])
        category = _normalize_category(row["phenotype_category"])
        by_annotation.setdefault((row["annotation_id"], genotype), level)
        known_annotations.add(str(row["annotation_id"]))
        for drug in MULTI_SEP.split(row["drugs"] or ""):
            drug = drug.strip().lower()
            if not drug:
                continue
            by_category.setdefault((rsid, drug, genotype, category), level)
            by_triple.setdefault((rsid, drug, genotype), set()).add(level)

    for row in authored:
        if not row.rsid or not row.evidence_level:
            continue
        # Counted here rather than as `len(authored)`: a row with no level makes no claim to check, and
        # the denominator must be claims compared. The ambiguous-annotation branch below `continue`s
        # *after* this increment on purpose — that row was looked up and the answer was "cannot tell",
        # which is a comparison that ran and settled nothing, not one that never happened.
        result.compared += 1
        drug, genotype = row.drug.strip().lower(), _normalize_genotype(row.genotype)
        category = _normalize_category(row.phenotype_category)
        reported: str | None = None
        if row.annotation_id and row.annotation_id not in known_annotations:
            # RM297: the row cites a record this snapshot does not hold. Falling through compared it
            # against another annotation that happened to share its category (S122: a withdrawn 1A
            # read as "ClinPGx says 3") — so the lookup stops here and says what is held instead.
            result.withheld.append(
                WithheldLevel(
                    CLINPGX_ANNOTATION_NOT_IN_SNAPSHOT,
                    row.rsid,
                    row.drug,
                    row.genotype,
                    row.evidence_level,
                    row.annotation_id,
                    _held(by_category, by_triple, row.rsid, drug, genotype, category),
                )
            )
            continue
        if row.annotation_id:
            reported = by_annotation.get((row.annotation_id, genotype))
        if reported is None and category is not None:
            reported = by_category.get((row.rsid, drug, genotype, category))
        if reported is None:
            levels = by_triple.get((row.rsid, drug, genotype), set())
            if len(levels) > 1:
                # Several annotations, no way to tell which this row is about. Saying "stale" here
                # would be a coin flip; say what is actually known instead.
                result.warnings.append(
                    f"{row.rsid} + {row.drug} [{row.genotype}]: ClinPGx has {len(levels)} "
                    f"annotations at levels {sorted(levels)} and this row names no "
                    f"phenotype_category or annotation_id, so its level was not checked."
                )
                continue
            reported = next(iter(levels), None)
        if reported is None:
            result.unmatched.append(f"{row.rsid} + {row.drug}")
            continue
        if reported == row.evidence_level:
            continue
        if not row.annotation_id:
            # The row claimed no record, so the category/triple match is a neighbour, not a citation:
            # reported in both modes, never refused (RM297). A cited id present without this genotype
            # keeps the fall-through into `conflicts` below, as before.
            result.withheld.append(
                WithheldLevel(
                    CLINPGX_LEVEL_DIFFERS_FROM_UNCITED,
                    row.rsid,
                    row.drug,
                    row.genotype,
                    row.evidence_level,
                    None,
                    (reported,),
                )
            )
            continue
        result.conflicts.append(
            EvidenceConflict(row.rsid, row.drug, row.genotype, row.evidence_level, reported)
        )

    for conflict in result.conflicts:
        logger.warning("ClinPGx evidence-level difference — %s", conflict)
    for withheld in result.withheld:
        result.warnings.append(str(withheld))
        logger.warning("ClinPGx evidence level not compared — %s", withheld)
    if result.unmatched:
        result.warnings.append(
            f"{len(result.unmatched)} authored annotation(s) have no matching ClinPGx record: "
            f"{sorted(set(result.unmatched))[:5]}. A miss is not a defect — the snapshot is a "
            f"point-in-time slice and a curator may annotate what ClinPGx has not."
        )

    # Strict refuses a *stale* level, because that is a currency fact rather than an opinion — and only
    # one reached through the row's own citation. `withheld` never refuses (RM297).
    if mode == "strict" and result.conflicts:
        raise ClinPgxEnrichmentError(
            f"strict: {len(result.conflicts)} authored evidence level(s) disagree with ClinPGx "
            f"({result.conflicts[0]}). An evidence level is ClinPGx's own metadata about its own "
            f"annotation, so a difference means the module is stale."
        )

    # The licence hash comes from the snapshot's release.json, which the builder took from the
    # LICENSE.txt inside the very archive the data came from.
    row = CLINPGX_TERMS.row("annotation", declared_use=declared_use, dataset=result.dataset)
    recorded_hash = release.get("license_sha256")
    if recorded_hash:
        row = row.model_copy(update={"license_sha256": recorded_hash})
    result.rows = [row]

    if write:
        merge_sources_file(result.rows, spec_dir, error=ClinPgxEnrichmentError)
    return _attest(result, spec_dir, write=write)