Skip to content

just_dna_enricher.cpic

just_dna_enricher.cpic

Live CPIC queries — allele function, diplotype→phenotype, and star-allele definitions (0.5).

CPIC is a PostgREST service: open, unauthenticated, and queried with ?column=eq.value&select=… rather than a bespoke query language. It covers the layers PharmVar does not — a diplotype's metabolizer phenotype and the prescribing recommendations built on it — and reaches genes outside PharmVar's fifteen.

CPIC is not a licence escape hatch. It merged into ClinPGx, and cpicpgx.org/license/ now 302-redirects to the ClinPGx data usage policy, so its terms are ClinPGx's terms: CC BY-SA plus a bar on sale. Preferring PharmVar for star alleles is a data-authority choice (PharmVar is the naming authority), not a licensing one — neither source is sellable.

Three shapes here do not map cleanly onto this workspace's models, and are surfaced rather than coerced:

  • allele_location_value.variantallele uses IUPAC ambiguity codes (R for A-or-G at CYP2C19 rs58973490). HaplotypeRow.allele requires plain nucleotides, so an ambiguous definition is reported and skipped — expanding R to two rows would invent two defining variants where CPIC recorded one uncertainty.
  • The same column also carries deletion/insertion and repeat notations — DELTCT, AAAGGGGCG(2), GGA(1) in CYP2D6 — which are not ambiguity codes, and were reported as if they were until a real CYP2D6 draft showed them. They are a grammar gap (RM5) rather than an uncertainty CPIC recorded, so unusable_allele_reason names which of the two a value is and the two are reported separately.
  • gene_result.activityscore is an inequality string ("≥3.0", "n/a"), not a number, so it does not drop into MeasureBinRow's numeric measure_min/measure_max. The raw string is carried and the parsing left to a human, because guessing a bound from ≥3.0 means inventing the upper one.

Coordinates from sequence_location.position are GRCh38 and 1-based — checked against Ensembl for rs4986893 (chr10:94780653) — which is what this pipeline already stores. Do not convert.

CPIC does publish a chromosome; it is on gene, not on sequence_location. An earlier probe looked at sequence_location alone (genesymbol / dbsnpid / position / chromosomelocation — no chromosome), concluded CPIC has none, and the drafting provider therefore skipped every defining variant CPIC gives no rsID for: 18 in CYP2C9, 14 in TPMT, 4 in NUDT15. gene.chr carries chr10 for CYP2C9, chr6 for TPMT, chr13 for NUDT15 (probed 2026-08-07, 132 genes). Joining it onto sequence_location.genesymbol is a lookup in CPIC's own tables, not the inference the old comment rightly refused — a star allele's defining variant is in the gene the allele belongs to, and the location row already names that gene.

CpicError

Bases: RuntimeError

A CPIC request failed in a way the caller must see.

CpicAllele dataclass

CpicAllele(
    gene: str,
    allele: str,
    activity_value: float | None = None,
    function_status: str | None = None,
)

One star allele's function and activity value, as CPIC assigns them.

CpicDiplotype dataclass

CpicDiplotype(
    gene: str,
    diplotype: str,
    phenotype: str | None = None,
    activity_score: str | None = None,
)

One diplotype and the phenotype CPIC assigns it.

CpicDefiningVariant dataclass

CpicDefiningVariant(
    gene: str,
    allele: str,
    rsid: str | None,
    chrom: str | None,
    start: int | None,
    variant_allele: str | None,
    unusable: str | None = None,
)

One variant defining a CPIC allele, at a GRCh38 coordinate.

CpicRecommendation dataclass

CpicRecommendation(
    gene: str,
    phenotype: str,
    drug: str,
    population: str,
    classification: str | None = None,
    recommendation: str | None = None,
    implication: str | None = None,
    activity_score: str | None = None,
)

One CPIC prescribing recommendation: a (gene phenotype, drug, population) triple.

Keyed by phenotype, not by diplotype — one recommendation covers every diplotype that shares the phenotype, which is why drafting joins on it rather than fetching per pair.

CpicClient

CpicClient(
    endpoint: str = DEFAULT_CPIC_ENDPOINT,
    *,
    client: Client | None = None,
    timeout: float = 60.0,
)

PostgREST client for the CPIC API. No auth, so no key handling.

Source code in enricher/src/just_dna_enricher/cpic.py
def __init__(
    self,
    endpoint: str = DEFAULT_CPIC_ENDPOINT,
    *,
    client: httpx.Client | None = None,
    timeout: float = 60.0,
) -> None:
    self.endpoint = endpoint.rstrip("/")
    self._client = client or httpx.Client(timeout=timeout)
    self._owned = client is None

row_count

row_count(table: str) -> int | None

How many rows CPIC says a table holds, or None when it does not answer with a count.

PostgREST reports it in Content-Range when asked with Prefer: count=exact. Used by the snapshot builder to refuse a short read: the service imposes no default limit today, but that is a fact about a deployment rather than a contract, and a silently truncated snapshot reads as "CPIC has nothing for this gene".

Routed through _request (RM97), not self._client directly. It bypassed its own transport, so the method whose whole job is to catch a truncated read had no retry and no translation: a transport failure raised raw httpx from a method whose contract is CpicError and escaped cli.cpic_build_'s except (CpicError, CpicBuildError) as a traceback, and a 5xx returned None — read by the builder as "CPIC gave no count", a wrong answer rather than an error. A client that documents its contract in two places and violates it in a third is the argument for testing the contract instead of the method.

Source code in enricher/src/just_dna_enricher/cpic.py
def row_count(self, table: str) -> int | None:
    """How many rows CPIC says a table holds, or `None` when it does not answer with a count.

    PostgREST reports it in `Content-Range` when asked with `Prefer: count=exact`. Used by the
    snapshot builder to refuse a short read: the service imposes no default limit today, but that
    is a fact about a deployment rather than a contract, and a silently truncated snapshot reads
    as "CPIC has nothing for this gene".

    **Routed through `_request` (RM97), not `self._client` directly.** It bypassed its own
    transport, so the method whose whole job is to catch a truncated read had no retry and no
    translation: a transport failure raised raw `httpx` from a method whose contract is
    `CpicError` and escaped `cli.cpic_build_`'s `except (CpicError, CpicBuildError)` as a
    traceback, and a 5xx returned `None` — read by the builder as "CPIC gave no count", a wrong
    answer rather than an error. A client that documents its contract in two places and violates
    it in a third is the argument for testing the contract instead of the method.
    """
    try:
        response = self._request(
            table,
            {"select": "count"},
            headers={"Prefer": "count=exact", "Range-Unit": "items", "Range": "0-0"},
        )
    except (httpx.TransportError, httpx.HTTPStatusError) as exc:
        raise CpicError(f"CPIC {table} row count could not be read: {exc}") from exc
    header = response.headers.get("content-range", "")
    total = header.rsplit("/", 1)[-1] if "/" in header else ""
    return int(total) if total.isdigit() else None

fetch_table

fetch_table(
    table: str, select: str
) -> list[dict[str, Any]]

One whole CPIC table as raw rows — the snapshot builder's read.

Public because the builder is a separate module and reaching for _get across it is the private-symbol reach RM41 exists to stop doing.

Source code in enricher/src/just_dna_enricher/cpic.py
def fetch_table(self, table: str, select: str) -> list[dict[str, Any]]:
    """One whole CPIC table as raw rows — the snapshot builder's read.

    Public because the builder is a separate module and reaching for `_get` across it is the
    private-symbol reach RM41 exists to stop doing.
    """
    return self._get(table, {"select": select})

chrom_for_gene

chrom_for_gene(gene: str) -> str | None

The contig CPIC places a gene on, from its own gene table, or None if it names none.

Separate request rather than a PostgREST embed: sequence_location has no foreign key to gene (it joins on the symbol string), so there is nothing to embed through.

Source code in enricher/src/just_dna_enricher/cpic.py
def chrom_for_gene(self, gene: str) -> str | None:
    """The contig CPIC places a gene on, from its own `gene` table, or `None` if it names none.

    Separate request rather than a PostgREST embed: `sequence_location` has no foreign key to
    `gene` (it joins on the symbol string), so there is nothing to embed through.
    """
    rows = self._get("gene", {"symbol": f"eq.{gene}", "select": "symbol,chr"})
    return normalize_chrom(rows[0].get("chr")) if rows else None

alleles_for_gene

alleles_for_gene(gene: str) -> list[CpicAllele]

Allele function + activity value for one gene.

Source code in enricher/src/just_dna_enricher/cpic.py
def alleles_for_gene(self, gene: str) -> list[CpicAllele]:
    """Allele function + activity value for one gene."""
    rows = self._get(
        "allele",
        {
            "genesymbol": f"eq.{gene}",
            "select": "genesymbol,name,activityvalue,clinicalfunctionalstatus",
            "order": "name",
        },
    )
    return [
        CpicAllele(
            gene=r["genesymbol"],
            allele=r["name"],
            activity_value=_float_or_none(r.get("activityvalue")),
            function_status=map_function_status(r.get("clinicalfunctionalstatus")),
        )
        for r in rows
    ]

diplotypes_for_gene

diplotypes_for_gene(gene: str) -> list[CpicDiplotype]

Diplotype → metabolizer phenotype for one gene.

Source code in enricher/src/just_dna_enricher/cpic.py
def diplotypes_for_gene(self, gene: str) -> list[CpicDiplotype]:
    """Diplotype → metabolizer phenotype for one gene."""
    rows = self._get(
        "diplotype",
        {
            "genesymbol": f"eq.{gene}",
            "select": "genesymbol,diplotype,generesult,totalactivityscore",
            "order": "diplotype",
        },
    )
    return [
        CpicDiplotype(
            gene=r["genesymbol"],
            diplotype=r["diplotype"],
            phenotype=r.get("generesult") or None,
            activity_score=r.get("totalactivityscore") or None,
        )
        for r in rows
    ]

recommendations

recommendations(
    gene: str, drug: str
) -> list[CpicRecommendation]

CPIC's prescribing recommendations for one gene/drug pair, across every population.

Two hops, because CPIC keys recommendations by RxNorm id: drug resolves the name, then recommendation is filtered on it. phenotypes is a {gene: phenotype} map, so a multi-gene recommendation is kept only when it names the gene asked for — a row about CYP2C19 and CYP2D6 is not a statement about CYP2C19 alone.

Deterministically ordered (population, phenotype) so a re-draft emits the same rows in the same order (Principle 7).

Source code in enricher/src/just_dna_enricher/cpic.py
def recommendations(self, gene: str, drug: str) -> list[CpicRecommendation]:
    """CPIC's prescribing recommendations for one gene/drug pair, across every population.

    Two hops, because CPIC keys recommendations by RxNorm id: `drug` resolves the name, then
    `recommendation` is filtered on it. `phenotypes` is a `{gene: phenotype}` map, so a
    multi-gene recommendation is kept only when it names the gene asked for — a row about
    CYP2C19 *and* CYP2D6 is not a statement about CYP2C19 alone.

    Deterministically ordered (population, phenotype) so a re-draft emits the same rows in the
    same order (Principle 7).
    """
    found = self._get("drug", {"select": "drugid,name", "name": f"eq.{drug.strip().lower()}"})
    if not found:
        return []
    rows = self._get(
        "recommendation",
        {
            "select": (
                "drugid,phenotypes,implications,drugrecommendation,classification,"
                "population,activityscore"
            ),
            "drugid": f"eq.{found[0]['drugid']}",
        },
    )
    out: list[CpicRecommendation] = []
    for row in rows:
        phenotypes = row.get("phenotypes") or {}
        phenotype = phenotypes.get(gene)
        if not phenotype or len(phenotypes) != 1:
            continue
        implications = row.get("implications") or {}
        out.append(
            CpicRecommendation(
                gene=gene,
                phenotype=phenotype,
                drug=drug.strip().lower(),
                population=(row.get("population") or "").strip(),
                classification=map_classification(row.get("classification")),
                recommendation=(row.get("drugrecommendation") or "").strip() or None,
                implication=(implications.get(gene) or "").strip() or None,
                activity_score=(row.get("activityscore") or None),
            )
        )
    return sorted(out, key=lambda r: (r.population, r.phenotype))

knows_drug

knows_drug(drug: str) -> bool | None

Does CPIC list this drug at all — asked only to explain an empty recommendations.

Live, this is answerable: CPIC's drug table is the registry of every drug it names, so an empty answer really is "no such drug" (a typo) and a non-empty one really is "CPIC knows it and simply has no phenotype-keyed recommendation for the gene you asked about". Those are different situations for an author and the drafter used to report them with one sentence — --drug warfarin and --drug notarealdrugxyz produced byte-identical output, though CPIC's warfarin guideline is real and is merely not shaped as a per-phenotype recommendation.

Separate call rather than folded into recommendations, because it is only ever needed to explain a negative: the ordinary path already knows the drug exists, and paying a second request on every successful draft to answer a question nobody asked is the wrong trade.

Source code in enricher/src/just_dna_enricher/cpic.py
def knows_drug(self, drug: str) -> bool | None:
    """Does CPIC list this drug at all — asked only to explain an empty `recommendations`.

    Live, this is answerable: CPIC's `drug` table is the registry of every drug it names, so an
    empty answer really is "no such drug" (a typo) and a non-empty one really is "CPIC knows it
    and simply has no phenotype-keyed recommendation for the gene you asked about". Those are
    different situations for an author and the drafter used to report them with one sentence —
    `--drug warfarin` and `--drug notarealdrugxyz` produced byte-identical output, though CPIC's
    warfarin guideline is real and is merely not shaped as a per-phenotype recommendation.

    Separate call rather than folded into `recommendations`, because it is only ever needed to
    explain a *negative*: the ordinary path already knows the drug exists, and paying a second
    request on every successful draft to answer a question nobody asked is the wrong trade.
    """
    return bool(self._get("drug", {"select": "name", "name": f"eq.{drug.strip().lower()}"}))

partner_genes

partner_genes(gene: str, drug: str) -> list[str]

The genes CPIC keys this drug's recommendations on together with gene (S102).

recommendations keeps a row only when its phenotypes map names exactly one gene, so a drug whose every guideline row is keyed on a gene pair — thiopurines on TPMT and NUDT15, the tricyclics on CYP2D6 and CYP2C19 — comes back empty, and the drafter had only knows_drug to explain it: CPIC lists the drug but records no single-gene recommendation. True, and useless, because it reads as "nothing to draft" when the fact is that the recommendation exists and is about a subject this table cannot hold. This is the second half of that explanation: the partners, so the author knows which gene the row is keyed with.

Asked only to explain an empty answer, like knows_drug; sorted, so a re-draft says the same thing. Empty means no multi-gene row names gene for this drug, which is also what an unknown drug returns — knows_drug is the one that tells those apart.

Source code in enricher/src/just_dna_enricher/cpic.py
def partner_genes(self, gene: str, drug: str) -> list[str]:
    """The genes CPIC keys this drug's recommendations on *together with* `gene` (S102).

    `recommendations` keeps a row only when its `phenotypes` map names exactly one gene, so a
    drug whose every guideline row is keyed on a gene pair — thiopurines on TPMT **and** NUDT15,
    the tricyclics on CYP2D6 **and** CYP2C19 — comes back empty, and the drafter had only
    `knows_drug` to explain it: *CPIC lists the drug but records no single-gene recommendation*.
    True, and useless, because it reads as "nothing to draft" when the fact is that the
    recommendation exists and is about a subject this table cannot hold. This is the second
    half of that explanation: the partners, so the author knows which gene the row is keyed with.

    Asked only to explain an empty answer, like `knows_drug`; sorted, so a re-draft says the
    same thing. Empty means no multi-gene row names `gene` for this drug, which is also what an
    unknown drug returns — `knows_drug` is the one that tells those apart.
    """
    found = self._get("drug", {"select": "drugid,name", "name": f"eq.{drug.strip().lower()}"})
    if not found:
        return []
    rows = self._get("recommendation", {"select": "phenotypes", "drugid": f"eq.{found[0]['drugid']}"})
    partners: set[str] = set()
    for row in rows:
        phenotypes = row.get("phenotypes") or {}
        if gene in phenotypes and len(phenotypes) > 1:
            partners.update(g for g in phenotypes if g != gene)
    return sorted(partners)

defining_variants

defining_variants(
    gene: str,
) -> tuple[list[CpicDefiningVariant], list[str]]

Star-allele defining variants for one gene, plus warnings for what could not be used.

Two requests rather than a nested select, because PostgREST embedding across allele_definition → allele_location_value → sequence_location needs the definition ids first. Returns (variants, warnings); an allele the format cannot hold is reported by which of the two reasons applies (unusable_allele_reason), aggregated, and never coerced.

Source code in enricher/src/just_dna_enricher/cpic.py
def defining_variants(self, gene: str) -> tuple[list[CpicDefiningVariant], list[str]]:
    """Star-allele defining variants for one gene, plus warnings for what could not be used.

    Two requests rather than a nested select, because PostgREST embedding across
    `allele_definition → allele_location_value → sequence_location` needs the definition ids
    first. Returns `(variants, warnings)`; an allele the format cannot hold is reported by *which*
    of the two reasons applies (`unusable_allele_reason`), aggregated, and never coerced.
    """
    definitions = self._get(
        "allele_definition",
        {"genesymbol": f"eq.{gene}", "select": "id,name,genesymbol", "order": "name"},
    )
    if not definitions:
        return [], []
    # One extra request per gene, and it is what makes a coordinate-only defining variant usable:
    # `HaplotypeRow` wants rsid **or** chrom AND start, and 36 of CPIC's defining variants across
    # CYP2C9/TPMT/NUDT15 have a position and no rsID. See the module docstring.
    chrom = self.chrom_for_gene(gene)
    by_id = {d["id"]: d["name"] for d in definitions}
    ids = ",".join(str(i) for i in by_id)
    rows = self._get(
        "allele_location_value",
        {
            "alleledefinitionid": f"in.({ids})",
            "select": ("alleledefinitionid,variantallele,sequence_location(genesymbol,dbsnpid,position)"),
        },
    )
    out: list[CpicDefiningVariant] = []
    # Aggregated per reason, not one line per row: a real CYP2D6 draft hits 67 of these, which buries
    # every other finding in the run — the same collapse already applied to the activity scores and
    # the copy-number diplotypes. Grouped by *reason* because the two are different findings with
    # different fixes (an ambiguity is permanent, a notation is RM5).
    unusable: dict[str, list[str]] = {}
    for r in rows:
        location = r.get("sequence_location") or {}
        allele_value = (r.get("variantallele") or "").strip().upper()
        reason = unusable_allele_reason(allele_value)
        if reason is not None:
            label = by_id.get(r["alleledefinitionid"], "")
            unusable.setdefault(reason, []).append(f"{label}={allele_value}")
        out.append(
            CpicDefiningVariant(
                gene=location.get("genesymbol") or gene,
                allele=by_id.get(r["alleledefinitionid"], ""),
                rsid=location.get("dbsnpid") or None,
                # GRCh38, 1-based, matching Ensembl — no conversion. `chrom` comes from CPIC's own
                # `gene` table: `sequence_location` really does have no chromosome column, which
                # is what the 2026-08-03 probe saw, but `gene.chr` does and joining on the symbol
                # the location row already carries is a lookup rather than the inference that
                # probe rightly refused. Still `None` when CPIC leaves `chr` blank.
                chrom=chrom,
                start=location.get("position"),
                variant_allele=allele_value or None,
                unusable=reason,
            )
        )
    return out, _unusable_warnings(gene, unusable)

CpicSnapshotClient

CpicSnapshotClient(reference: Path)

The same four questions, answered from a built snapshot instead of the live service (RM38).

Duck-typed against CpicClient, deliberately, and not a subclass of it. enrich_pgx and pgx_draft.draft_gene already take an injected client, so a reader with the same four methods needs no branch in either — which is what keeps "snapshot or live" out of the passes entirely. It is not a subclass because it shares no implementation: there is no endpoint, no pacing, no retry.

The vocabulary mapping happens here, not in the builder. The parquet holds CPIC's own prose verbatim ("No function", "Strong") and map_function_status / map_classification run on the way out — the same functions the live client calls, so a snapshot answer and a live answer are the same object by construction, and a mapping fix reaches a snapshot built last month.

Read with duckdb, which is a core dependency, rather than polars, which is [dev] and only the builder needs. That is the house convention (builder in polars, runtime pass in duckdb), and it is what keeps the declared runtime dependency set honest about what the runtime actually requires.

Source code in enricher/src/just_dna_enricher/cpic.py
def __init__(self, reference: Path) -> None:
    self.reference = Path(reference)
    data_dir = self.reference / SNAPSHOT_DATA_DIRNAME
    if not data_dir.is_dir() or not any(data_dir.glob("*.parquet")):
        raise CpicError(f"no CPIC snapshot at {data_dir} — build one with `just-dna-enricher cpic build`")
    self._data_dir = data_dir
    release_path = self.reference / RELEASE_FILENAME
    self.release: dict[str, Any] = json.loads(release_path.read_text()) if release_path.is_file() else {}

dataset property

dataset: str | None

Which snapshot answered — recorded onto SourceRow.dataset, as the gnomAD routes are.

A consumer must be able to tell a pinned file from a live API: the two can differ by a release, and dataset is inside the fact set where source is not.

close

close() -> None

No-op, so a caller can close() this and a live client identically.

Source code in enricher/src/just_dna_enricher/cpic.py
def close(self) -> None:
    """No-op, so a caller can `close()` this and a live client identically."""

recommendations

recommendations(
    gene: str, drug: str
) -> list[CpicRecommendation]

One gene/drug pair's recommendations across every population.

gene_count = 1 reproduces the live client's len(phenotypes) != 1 filter exactly: the builder flattens a recommendation's phenotypes map to one row per gene and carries how many there were, because a row about CYP2C19 and CYP2D6 is not a statement about CYP2C19 alone.

Source code in enricher/src/just_dna_enricher/cpic.py
def recommendations(self, gene: str, drug: str) -> list[CpicRecommendation]:
    """One gene/drug pair's recommendations across every population.

    `gene_count = 1` reproduces the live client's `len(phenotypes) != 1` filter exactly: the builder
    flattens a recommendation's `phenotypes` map to one row per gene and carries how many there
    were, because a row about CYP2C19 *and* CYP2D6 is not a statement about CYP2C19 alone.
    """
    rows = self._rows(
        "recommendations.parquet",
        "gene = ? AND drug = ? AND gene_count = 1",
        [gene, drug.strip().lower()],
        "gene, phenotype, drug, population, classification, recommendation, implication, activity_score",
    )
    out = [
        CpicRecommendation(
            gene=r["gene"],
            phenotype=r["phenotype"],
            drug=r["drug"],
            population=(r["population"] or "").strip(),
            classification=map_classification(r["classification"]),
            recommendation=(r["recommendation"] or "").strip() or None,
            implication=(r["implication"] or "").strip() or None,
            activity_score=r["activity_score"] or None,
        )
        for r in rows
    ]
    # Same deterministic order the live client emits (Principle 7).
    return sorted(out, key=lambda r: (r.population, r.phenotype))

knows_drug

knows_drug(drug: str) -> bool | None

None — the snapshot cannot answer this, and saying False would be a false negative.

The live client reads CPIC's drug table, which lists every drug CPIC names. The snapshot carries recommendations.parquet and nothing else, so the only drug names in it are the ones that have a recommendation — which makes "absent from this file" and "absent from CPIC" exactly the two things the caller is trying to tell apart. Answering False here would turn a limitation of the snapshot into an assertion about the source, which is the shape S20 exists to prevent, so it withholds and the caller's message says which file it read.

Building the drug table into the snapshot is the fix if this ever needs a definite answer; it is not worth a column today, since the only consumer is one warning's wording.

Source code in enricher/src/just_dna_enricher/cpic.py
def knows_drug(self, drug: str) -> bool | None:
    """**`None` — the snapshot cannot answer this, and saying `False` would be a false negative.**

    The live client reads CPIC's `drug` table, which lists every drug CPIC names. The snapshot
    carries `recommendations.parquet` and nothing else, so the only drug names in it are the ones
    that *have* a recommendation — which makes "absent from this file" and "absent from CPIC"
    exactly the two things the caller is trying to tell apart. Answering `False` here would turn
    a limitation of the snapshot into an assertion about the source, which is the shape S20 exists
    to prevent, so it withholds and the caller's message says which file it read.

    Building the drug table into the snapshot is the fix if this ever needs a definite answer; it
    is not worth a column today, since the only consumer is one warning's wording.
    """
    return None

partner_genes

partner_genes(gene: str, drug: str) -> list[str]

The same question as the live client's, off the flattened table (S102).

The builder flattens one recommendation's phenotypes map to one row per gene and keeps no recommendation id, so the rows of one guideline cannot be grouped back together here. The answer is therefore every gene the drug's multi-gene rows are keyed on other than gene, which is exact while no drug is keyed on more than one pair — and that is CPIC today: measured on the 2026-09-02 snapshot, gene_count never exceeds 2 and no drug has both a single-gene and a multi-gene shape. A three-gene guideline would still be named correctly; two different pairs for one drug would be listed together, which is the one reading this cannot make and the builder would need a column to give it.

Source code in enricher/src/just_dna_enricher/cpic.py
def partner_genes(self, gene: str, drug: str) -> list[str]:
    """The same question as the live client's, off the flattened table (S102).

    The builder flattens one recommendation's `phenotypes` map to one row per gene and keeps
    no recommendation id, so the rows of one guideline cannot be grouped back together here.
    The answer is therefore *every* gene the drug's multi-gene rows are keyed on other than
    `gene`, which is exact while no drug is keyed on more than one pair — and that is CPIC
    today: measured on the 2026-09-02 snapshot, `gene_count` never exceeds 2 and no drug has
    both a single-gene and a multi-gene shape. A three-gene guideline would still be named
    correctly; two *different* pairs for one drug would be listed together, which is the one
    reading this cannot make and the builder would need a column to give it.
    """
    named = self._rows(
        "recommendations.parquet",
        "gene = ? AND drug = ? AND gene_count > 1",
        [gene, drug.strip().lower()],
        "gene",
    )
    if not named:
        return []
    others = self._rows(
        "recommendations.parquet",
        "drug = ? AND gene_count > 1 AND gene <> ?",
        [drug.strip().lower(), gene],
        "DISTINCT gene",
        "gene",
    )
    return [r["gene"] for r in others]

defining_variants

defining_variants(
    gene: str,
) -> tuple[list[CpicDefiningVariant], list[str]]

Star-allele defining variants, with the same aggregated unusable-allele warnings.

The classification runs here rather than being stored, for the reason the whole snapshot follows: unusable_allele_reason is a judgement this workspace makes about CPIC's value, so freezing it into the parquet would pin one release's opinion into every snapshot built under it.

Source code in enricher/src/just_dna_enricher/cpic.py
def defining_variants(self, gene: str) -> tuple[list[CpicDefiningVariant], list[str]]:
    """Star-allele defining variants, with the same aggregated unusable-allele warnings.

    The classification runs here rather than being stored, for the reason the whole snapshot
    follows: `unusable_allele_reason` is a *judgement this workspace makes* about CPIC's value, so
    freezing it into the parquet would pin one release's opinion into every snapshot built under it.
    """
    rows = self._rows(
        "allele_definitions.parquet",
        "gene = ?",
        [gene],
        "gene, allele, rsid, chrom, start, variant_allele",
        "allele, start, variant_allele",
    )
    out: list[CpicDefiningVariant] = []
    unusable: dict[str, list[str]] = {}
    for r in rows:
        allele_value = (r["variant_allele"] or "").strip().upper()
        reason = unusable_allele_reason(allele_value)
        if reason is not None:
            unusable.setdefault(reason, []).append(f"{r['allele']}={allele_value}")
        out.append(
            CpicDefiningVariant(
                gene=r["gene"] or gene,
                allele=r["allele"] or "",
                rsid=r["rsid"] or None,
                chrom=r["chrom"] or None,
                start=r["start"],
                variant_allele=allele_value or None,
                unusable=reason,
            )
        )
    return out, _unusable_warnings(gene, unusable)

unusable_allele_reason

unusable_allele_reason(value: str) -> str | None

Why variantallele cannot become a HaplotypeRow.allele, or None when it can.

Two distinct reasons, and calling both "an IUPAC ambiguity code" was wrong. Drafting real CYP2D6 turned up DELTCT, AAAGGGGCG(2) and GGA(1) beside the genuine codes (R, S), and the message announced all of them as ambiguity codes — a claim about the data that is simply false for a deletion or a repeat, and it points an author at the wrong thing. The distinction matters for what happens next, too: an ambiguous base is an uncertainty CPIC recorded and will never be expressible, while a structural notation is a grammar gap (RM5) that a future release may widen to hold.

  • "ambiguity" — every character is a nucleotide or an IUPAC ambiguity code (R = A or G).
  • "symbolic" — a symbolic/structural allele that states no length. RM5 made the well-formed ones authorable, so <DEL:1500> is now a perfectly good HaplotypeRow.allele and this function answers None for it; a lengthless <DEL> still cannot be used, because the compiler refuses one on haplotypes.csv in both modes. CPIC publishes neither spelling today (its own are DELTCT and AAAGGGGCG(2), which are "notation"), so the arm is dead against the live source — it exists because the classifier is shared, and because a caller that cannot explain a reason is a KeyError waiting for the first source that does emit one.
  • "notation" — a deletion/insertion or repeat notation, not a nucleotide string at all.

A classification is not a verdict, and this function owes the verdict. It delegates the classification to just_dna_format.alleles.non_nucleotide_reason — one definition, two callers, each keeping its own wording — but its own name asks whether the value can become a HaplotypeRow.allele, and since RM5 the answer for a well-formed symbolic allele is yes. Returning the bare classification made pgx_draft skip a defining variant the model would have accepted, while the message beside it said the value belongs in allele as written.

Source code in enricher/src/just_dna_enricher/cpic.py
def unusable_allele_reason(value: str) -> str | None:
    """Why `variantallele` cannot become a `HaplotypeRow.allele`, or `None` when it can.

    **Two distinct reasons, and calling both "an IUPAC ambiguity code" was wrong.** Drafting real CYP2D6
    turned up `DELTCT`, `AAAGGGGCG(2)` and `GGA(1)` beside the genuine codes (`R`, `S`), and the message
    announced all of them as ambiguity codes — a claim about the data that is simply false for a deletion
    or a repeat, and it points an author at the wrong thing. The distinction matters for what happens
    next, too: an ambiguous base is an *uncertainty CPIC recorded* and will never be expressible, while a
    structural notation is a **grammar gap** (RM5) that a future release may widen to hold.

    * `"ambiguity"` — every character is a nucleotide or an IUPAC ambiguity code (`R` = A or G).
    * `"symbolic"` — a symbolic/structural allele **that states no length**. RM5 made the well-formed
      ones authorable, so `<DEL:1500>` is now a perfectly good `HaplotypeRow.allele` and this function
      answers `None` for it; a lengthless `<DEL>` still cannot be used, because the compiler refuses one
      on `haplotypes.csv` in both modes. CPIC publishes neither spelling today (its own are `DELTCT` and
      `AAAGGGGCG(2)`, which are `"notation"`), so the arm is dead against the live source — it exists
      because the classifier is shared, and because a caller that cannot explain a reason is a
      `KeyError` waiting for the first source that does emit one.
    * `"notation"` — a deletion/insertion or repeat notation, not a nucleotide string at all.

    **A classification is not a verdict, and this function owes the verdict.** It delegates the
    *classification* to `just_dna_format.alleles.non_nucleotide_reason` — one definition, two callers,
    each keeping its own wording — but its own name asks whether the value can *become* a
    `HaplotypeRow.allele`, and since RM5 the answer for a well-formed symbolic allele is yes. Returning
    the bare classification made `pgx_draft` skip a defining variant the model would have accepted,
    while the message beside it said the value belongs in `allele` as written.
    """
    reason = non_nucleotide_reason(value)
    if reason != "symbolic":
        return reason
    return "symbolic" if symbolic_allele_defect(value) is not None else None

map_classification

map_classification(raw: str | None) -> str | None

CPIC's recommendation strength → this format's vocabulary, or None when it did not classify.

Source code in enricher/src/just_dna_enricher/cpic.py
def map_classification(raw: str | None) -> str | None:
    """CPIC's recommendation strength → this format's vocabulary, or None when it did not classify."""
    if not raw:
        return None
    return _CLASSIFICATION_MAP.get(raw.strip().lower())

map_function_status

map_function_status(raw: str | None) -> str | None

CPIC's prose → pgx.VALID_FUNCTION_STATUS, or None when it says something else.

The two possible … values collapse to uncertain_function: CPIC uses them for a call it is not confident in, which is exactly what uncertain means here, and inventing new vocabulary members for them would make every consumer's switch statement wrong.

Source code in enricher/src/just_dna_enricher/cpic.py
def map_function_status(raw: str | None) -> str | None:
    """CPIC's prose → `pgx.VALID_FUNCTION_STATUS`, or None when it says something else.

    The two `possible …` values collapse to `uncertain_function`: CPIC uses them for a call it is not
    confident in, which is exactly what `uncertain` means here, and inventing new vocabulary members
    for them would make every consumer's switch statement wrong.
    """
    if not raw:
        return None
    return _FUNCTION_MAP.get(raw.strip().lower())

normalize_chrom

normalize_chrom(value: str | None) -> str | None

CPIC's gene.chr (chr10, chrX) → this workspace's bare contig name (10, X).

None for a blank or an unplaced spelling, rather than a guess — the same rule chrom_from_accession follows in the PharmVar client.

Source code in enricher/src/just_dna_enricher/cpic.py
def normalize_chrom(value: str | None) -> str | None:
    """CPIC's `gene.chr` (`chr10`, `chrX`) → this workspace's bare contig name (`10`, `X`).

    `None` for a blank or an unplaced spelling, rather than a guess — the same rule
    `chrom_from_accession` follows in the PharmVar client.
    """
    text = (value or "").strip()
    if not text:
        return None
    return text.removeprefix("chr").removeprefix("CHR").strip() or None