Live CPIC queries — allele function, diplotype→phenotype, and star-allele definitions (0.5).
CPIC is a PostgREST service: open, unauthenticated, and queried with ?column=eq.value&select=…
rather than a bespoke query language. It covers the layers PharmVar does not — a diplotype's
metabolizer phenotype and the prescribing recommendations built on it — and reaches genes outside
PharmVar's fifteen.
CPIC is not a licence escape hatch. It merged into ClinPGx, and cpicpgx.org/license/ now
302-redirects to the ClinPGx data usage policy, so its terms are ClinPGx's terms: CC BY-SA plus a bar
on sale. Preferring PharmVar for star alleles is a data-authority choice (PharmVar is the naming
authority), not a licensing one — neither source is sellable.
Three shapes here do not map cleanly onto this workspace's models, and are surfaced rather than
coerced:
allele_location_value.variantallele uses IUPAC ambiguity codes (R for A-or-G at CYP2C19
rs58973490). HaplotypeRow.allele requires plain nucleotides, so an ambiguous definition is
reported and skipped — expanding R to two rows would invent two defining variants where CPIC
recorded one uncertainty.
- The same column also carries deletion/insertion and repeat notations —
DELTCT,
AAAGGGGCG(2), GGA(1) in CYP2D6 — which are not ambiguity codes, and were reported as if they
were until a real CYP2D6 draft showed them. They are a grammar gap (RM5) rather than an uncertainty
CPIC recorded, so unusable_allele_reason names which of the two a value is and the two are
reported separately.
gene_result.activityscore is an inequality string ("≥3.0", "n/a"), not a number, so it does
not drop into MeasureBinRow's numeric measure_min/measure_max. The raw string is carried and
the parsing left to a human, because guessing a bound from ≥3.0 means inventing the upper one.
Coordinates from sequence_location.position are GRCh38 and 1-based — checked against Ensembl
for rs4986893 (chr10:94780653) — which is what this pipeline already stores. Do not convert.
CPIC does publish a chromosome; it is on gene, not on sequence_location. An earlier probe
looked at sequence_location alone (genesymbol / dbsnpid / position / chromosomelocation — no
chromosome), concluded CPIC has none, and the drafting provider therefore skipped every defining
variant CPIC gives no rsID for: 18 in CYP2C9, 14 in TPMT, 4 in NUDT15. gene.chr carries chr10 for
CYP2C9, chr6 for TPMT, chr13 for NUDT15 (probed 2026-08-07, 132 genes). Joining it onto
sequence_location.genesymbol is a lookup in CPIC's own tables, not the inference the old comment
rightly refused — a star allele's defining variant is in the gene the allele belongs to, and the
location row already names that gene.
CpicError
Bases: RuntimeError
A CPIC request failed in a way the caller must see.
CpicAllele
dataclass
CpicAllele(
gene: str,
allele: str,
activity_value: float | None = None,
function_status: str | None = None,
)
One star allele's function and activity value, as CPIC assigns them.
CpicDiplotype
dataclass
CpicDiplotype(
gene: str,
diplotype: str,
phenotype: str | None = None,
activity_score: str | None = None,
)
One diplotype and the phenotype CPIC assigns it.
CpicDefiningVariant
dataclass
CpicDefiningVariant(
gene: str,
allele: str,
rsid: str | None,
chrom: str | None,
start: int | None,
variant_allele: str | None,
unusable: str | None = None,
)
One variant defining a CPIC allele, at a GRCh38 coordinate.
CpicRecommendation
dataclass
CpicRecommendation(
gene: str,
phenotype: str,
drug: str,
population: str,
classification: str | None = None,
recommendation: str | None = None,
implication: str | None = None,
activity_score: str | None = None,
)
One CPIC prescribing recommendation: a (gene phenotype, drug, population) triple.
Keyed by phenotype, not by diplotype — one recommendation covers every diplotype that shares
the phenotype, which is why drafting joins on it rather than fetching per pair.
CpicClient
CpicClient(
endpoint: str = DEFAULT_CPIC_ENDPOINT,
*,
client: Client | None = None,
timeout: float = 60.0,
)
PostgREST client for the CPIC API. No auth, so no key handling.
Source code in enricher/src/just_dna_enricher/cpic.py
| def __init__(
self,
endpoint: str = DEFAULT_CPIC_ENDPOINT,
*,
client: httpx.Client | None = None,
timeout: float = 60.0,
) -> None:
self.endpoint = endpoint.rstrip("/")
self._client = client or httpx.Client(timeout=timeout)
self._owned = client is None
|
row_count
row_count(table: str) -> int | None
How many rows CPIC says a table holds, or None when it does not answer with a count.
PostgREST reports it in Content-Range when asked with Prefer: count=exact. Used by the
snapshot builder to refuse a short read: the service imposes no default limit today, but that
is a fact about a deployment rather than a contract, and a silently truncated snapshot reads
as "CPIC has nothing for this gene".
Routed through _request (RM97), not self._client directly. It bypassed its own
transport, so the method whose whole job is to catch a truncated read had no retry and no
translation: a transport failure raised raw httpx from a method whose contract is
CpicError and escaped cli.cpic_build_'s except (CpicError, CpicBuildError) as a
traceback, and a 5xx returned None — read by the builder as "CPIC gave no count", a wrong
answer rather than an error. A client that documents its contract in two places and violates
it in a third is the argument for testing the contract instead of the method.
Source code in enricher/src/just_dna_enricher/cpic.py
| def row_count(self, table: str) -> int | None:
"""How many rows CPIC says a table holds, or `None` when it does not answer with a count.
PostgREST reports it in `Content-Range` when asked with `Prefer: count=exact`. Used by the
snapshot builder to refuse a short read: the service imposes no default limit today, but that
is a fact about a deployment rather than a contract, and a silently truncated snapshot reads
as "CPIC has nothing for this gene".
**Routed through `_request` (RM97), not `self._client` directly.** It bypassed its own
transport, so the method whose whole job is to catch a truncated read had no retry and no
translation: a transport failure raised raw `httpx` from a method whose contract is
`CpicError` and escaped `cli.cpic_build_`'s `except (CpicError, CpicBuildError)` as a
traceback, and a 5xx returned `None` — read by the builder as "CPIC gave no count", a wrong
answer rather than an error. A client that documents its contract in two places and violates
it in a third is the argument for testing the contract instead of the method.
"""
try:
response = self._request(
table,
{"select": "count"},
headers={"Prefer": "count=exact", "Range-Unit": "items", "Range": "0-0"},
)
except (httpx.TransportError, httpx.HTTPStatusError) as exc:
raise CpicError(f"CPIC {table} row count could not be read: {exc}") from exc
header = response.headers.get("content-range", "")
total = header.rsplit("/", 1)[-1] if "/" in header else ""
return int(total) if total.isdigit() else None
|
fetch_table
fetch_table(
table: str, select: str
) -> list[dict[str, Any]]
One whole CPIC table as raw rows — the snapshot builder's read.
Public because the builder is a separate module and reaching for _get across it is the
private-symbol reach RM41 exists to stop doing.
Source code in enricher/src/just_dna_enricher/cpic.py
| def fetch_table(self, table: str, select: str) -> list[dict[str, Any]]:
"""One whole CPIC table as raw rows — the snapshot builder's read.
Public because the builder is a separate module and reaching for `_get` across it is the
private-symbol reach RM41 exists to stop doing.
"""
return self._get(table, {"select": select})
|
chrom_for_gene
chrom_for_gene(gene: str) -> str | None
The contig CPIC places a gene on, from its own gene table, or None if it names none.
Separate request rather than a PostgREST embed: sequence_location has no foreign key to
gene (it joins on the symbol string), so there is nothing to embed through.
Source code in enricher/src/just_dna_enricher/cpic.py
| def chrom_for_gene(self, gene: str) -> str | None:
"""The contig CPIC places a gene on, from its own `gene` table, or `None` if it names none.
Separate request rather than a PostgREST embed: `sequence_location` has no foreign key to
`gene` (it joins on the symbol string), so there is nothing to embed through.
"""
rows = self._get("gene", {"symbol": f"eq.{gene}", "select": "symbol,chr"})
return normalize_chrom(rows[0].get("chr")) if rows else None
|
alleles_for_gene
alleles_for_gene(gene: str) -> list[CpicAllele]
Allele function + activity value for one gene.
Source code in enricher/src/just_dna_enricher/cpic.py
| def alleles_for_gene(self, gene: str) -> list[CpicAllele]:
"""Allele function + activity value for one gene."""
rows = self._get(
"allele",
{
"genesymbol": f"eq.{gene}",
"select": "genesymbol,name,activityvalue,clinicalfunctionalstatus",
"order": "name",
},
)
return [
CpicAllele(
gene=r["genesymbol"],
allele=r["name"],
activity_value=_float_or_none(r.get("activityvalue")),
function_status=map_function_status(r.get("clinicalfunctionalstatus")),
)
for r in rows
]
|
diplotypes_for_gene
diplotypes_for_gene(gene: str) -> list[CpicDiplotype]
Diplotype → metabolizer phenotype for one gene.
Source code in enricher/src/just_dna_enricher/cpic.py
| def diplotypes_for_gene(self, gene: str) -> list[CpicDiplotype]:
"""Diplotype → metabolizer phenotype for one gene."""
rows = self._get(
"diplotype",
{
"genesymbol": f"eq.{gene}",
"select": "genesymbol,diplotype,generesult,totalactivityscore",
"order": "diplotype",
},
)
return [
CpicDiplotype(
gene=r["genesymbol"],
diplotype=r["diplotype"],
phenotype=r.get("generesult") or None,
activity_score=r.get("totalactivityscore") or None,
)
for r in rows
]
|
recommendations
recommendations(
gene: str, drug: str
) -> list[CpicRecommendation]
CPIC's prescribing recommendations for one gene/drug pair, across every population.
Two hops, because CPIC keys recommendations by RxNorm id: drug resolves the name, then
recommendation is filtered on it. phenotypes is a {gene: phenotype} map, so a
multi-gene recommendation is kept only when it names the gene asked for — a row about
CYP2C19 and CYP2D6 is not a statement about CYP2C19 alone.
Deterministically ordered (population, phenotype) so a re-draft emits the same rows in the
same order (Principle 7).
Source code in enricher/src/just_dna_enricher/cpic.py
| def recommendations(self, gene: str, drug: str) -> list[CpicRecommendation]:
"""CPIC's prescribing recommendations for one gene/drug pair, across every population.
Two hops, because CPIC keys recommendations by RxNorm id: `drug` resolves the name, then
`recommendation` is filtered on it. `phenotypes` is a `{gene: phenotype}` map, so a
multi-gene recommendation is kept only when it names the gene asked for — a row about
CYP2C19 *and* CYP2D6 is not a statement about CYP2C19 alone.
Deterministically ordered (population, phenotype) so a re-draft emits the same rows in the
same order (Principle 7).
"""
found = self._get("drug", {"select": "drugid,name", "name": f"eq.{drug.strip().lower()}"})
if not found:
return []
rows = self._get(
"recommendation",
{
"select": (
"drugid,phenotypes,implications,drugrecommendation,classification,"
"population,activityscore"
),
"drugid": f"eq.{found[0]['drugid']}",
},
)
out: list[CpicRecommendation] = []
for row in rows:
phenotypes = row.get("phenotypes") or {}
phenotype = phenotypes.get(gene)
if not phenotype or len(phenotypes) != 1:
continue
implications = row.get("implications") or {}
out.append(
CpicRecommendation(
gene=gene,
phenotype=phenotype,
drug=drug.strip().lower(),
population=(row.get("population") or "").strip(),
classification=map_classification(row.get("classification")),
recommendation=(row.get("drugrecommendation") or "").strip() or None,
implication=(implications.get(gene) or "").strip() or None,
activity_score=(row.get("activityscore") or None),
)
)
return sorted(out, key=lambda r: (r.population, r.phenotype))
|
knows_drug
knows_drug(drug: str) -> bool | None
Does CPIC list this drug at all — asked only to explain an empty recommendations.
Live, this is answerable: CPIC's drug table is the registry of every drug it names, so an
empty answer really is "no such drug" (a typo) and a non-empty one really is "CPIC knows it
and simply has no phenotype-keyed recommendation for the gene you asked about". Those are
different situations for an author and the drafter used to report them with one sentence —
--drug warfarin and --drug notarealdrugxyz produced byte-identical output, though CPIC's
warfarin guideline is real and is merely not shaped as a per-phenotype recommendation.
Separate call rather than folded into recommendations, because it is only ever needed to
explain a negative: the ordinary path already knows the drug exists, and paying a second
request on every successful draft to answer a question nobody asked is the wrong trade.
Source code in enricher/src/just_dna_enricher/cpic.py
| def knows_drug(self, drug: str) -> bool | None:
"""Does CPIC list this drug at all — asked only to explain an empty `recommendations`.
Live, this is answerable: CPIC's `drug` table is the registry of every drug it names, so an
empty answer really is "no such drug" (a typo) and a non-empty one really is "CPIC knows it
and simply has no phenotype-keyed recommendation for the gene you asked about". Those are
different situations for an author and the drafter used to report them with one sentence —
`--drug warfarin` and `--drug notarealdrugxyz` produced byte-identical output, though CPIC's
warfarin guideline is real and is merely not shaped as a per-phenotype recommendation.
Separate call rather than folded into `recommendations`, because it is only ever needed to
explain a *negative*: the ordinary path already knows the drug exists, and paying a second
request on every successful draft to answer a question nobody asked is the wrong trade.
"""
return bool(self._get("drug", {"select": "name", "name": f"eq.{drug.strip().lower()}"}))
|
partner_genes
partner_genes(gene: str, drug: str) -> list[str]
The genes CPIC keys this drug's recommendations on together with gene (S102).
recommendations keeps a row only when its phenotypes map names exactly one gene, so a
drug whose every guideline row is keyed on a gene pair — thiopurines on TPMT and NUDT15,
the tricyclics on CYP2D6 and CYP2C19 — comes back empty, and the drafter had only
knows_drug to explain it: CPIC lists the drug but records no single-gene recommendation.
True, and useless, because it reads as "nothing to draft" when the fact is that the
recommendation exists and is about a subject this table cannot hold. This is the second
half of that explanation: the partners, so the author knows which gene the row is keyed with.
Asked only to explain an empty answer, like knows_drug; sorted, so a re-draft says the
same thing. Empty means no multi-gene row names gene for this drug, which is also what an
unknown drug returns — knows_drug is the one that tells those apart.
Source code in enricher/src/just_dna_enricher/cpic.py
| def partner_genes(self, gene: str, drug: str) -> list[str]:
"""The genes CPIC keys this drug's recommendations on *together with* `gene` (S102).
`recommendations` keeps a row only when its `phenotypes` map names exactly one gene, so a
drug whose every guideline row is keyed on a gene pair — thiopurines on TPMT **and** NUDT15,
the tricyclics on CYP2D6 **and** CYP2C19 — comes back empty, and the drafter had only
`knows_drug` to explain it: *CPIC lists the drug but records no single-gene recommendation*.
True, and useless, because it reads as "nothing to draft" when the fact is that the
recommendation exists and is about a subject this table cannot hold. This is the second
half of that explanation: the partners, so the author knows which gene the row is keyed with.
Asked only to explain an empty answer, like `knows_drug`; sorted, so a re-draft says the
same thing. Empty means no multi-gene row names `gene` for this drug, which is also what an
unknown drug returns — `knows_drug` is the one that tells those apart.
"""
found = self._get("drug", {"select": "drugid,name", "name": f"eq.{drug.strip().lower()}"})
if not found:
return []
rows = self._get("recommendation", {"select": "phenotypes", "drugid": f"eq.{found[0]['drugid']}"})
partners: set[str] = set()
for row in rows:
phenotypes = row.get("phenotypes") or {}
if gene in phenotypes and len(phenotypes) > 1:
partners.update(g for g in phenotypes if g != gene)
return sorted(partners)
|
defining_variants
defining_variants(
gene: str,
) -> tuple[list[CpicDefiningVariant], list[str]]
Star-allele defining variants for one gene, plus warnings for what could not be used.
Two requests rather than a nested select, because PostgREST embedding across
allele_definition → allele_location_value → sequence_location needs the definition ids
first. Returns (variants, warnings); an allele the format cannot hold is reported by which
of the two reasons applies (unusable_allele_reason), aggregated, and never coerced.
Source code in enricher/src/just_dna_enricher/cpic.py
| def defining_variants(self, gene: str) -> tuple[list[CpicDefiningVariant], list[str]]:
"""Star-allele defining variants for one gene, plus warnings for what could not be used.
Two requests rather than a nested select, because PostgREST embedding across
`allele_definition → allele_location_value → sequence_location` needs the definition ids
first. Returns `(variants, warnings)`; an allele the format cannot hold is reported by *which*
of the two reasons applies (`unusable_allele_reason`), aggregated, and never coerced.
"""
definitions = self._get(
"allele_definition",
{"genesymbol": f"eq.{gene}", "select": "id,name,genesymbol", "order": "name"},
)
if not definitions:
return [], []
# One extra request per gene, and it is what makes a coordinate-only defining variant usable:
# `HaplotypeRow` wants rsid **or** chrom AND start, and 36 of CPIC's defining variants across
# CYP2C9/TPMT/NUDT15 have a position and no rsID. See the module docstring.
chrom = self.chrom_for_gene(gene)
by_id = {d["id"]: d["name"] for d in definitions}
ids = ",".join(str(i) for i in by_id)
rows = self._get(
"allele_location_value",
{
"alleledefinitionid": f"in.({ids})",
"select": ("alleledefinitionid,variantallele,sequence_location(genesymbol,dbsnpid,position)"),
},
)
out: list[CpicDefiningVariant] = []
# Aggregated per reason, not one line per row: a real CYP2D6 draft hits 67 of these, which buries
# every other finding in the run — the same collapse already applied to the activity scores and
# the copy-number diplotypes. Grouped by *reason* because the two are different findings with
# different fixes (an ambiguity is permanent, a notation is RM5).
unusable: dict[str, list[str]] = {}
for r in rows:
location = r.get("sequence_location") or {}
allele_value = (r.get("variantallele") or "").strip().upper()
reason = unusable_allele_reason(allele_value)
if reason is not None:
label = by_id.get(r["alleledefinitionid"], "")
unusable.setdefault(reason, []).append(f"{label}={allele_value}")
out.append(
CpicDefiningVariant(
gene=location.get("genesymbol") or gene,
allele=by_id.get(r["alleledefinitionid"], ""),
rsid=location.get("dbsnpid") or None,
# GRCh38, 1-based, matching Ensembl — no conversion. `chrom` comes from CPIC's own
# `gene` table: `sequence_location` really does have no chromosome column, which
# is what the 2026-08-03 probe saw, but `gene.chr` does and joining on the symbol
# the location row already carries is a lookup rather than the inference that
# probe rightly refused. Still `None` when CPIC leaves `chr` blank.
chrom=chrom,
start=location.get("position"),
variant_allele=allele_value or None,
unusable=reason,
)
)
return out, _unusable_warnings(gene, unusable)
|
CpicSnapshotClient
CpicSnapshotClient(reference: Path)
The same four questions, answered from a built snapshot instead of the live service (RM38).
Duck-typed against CpicClient, deliberately, and not a subclass of it. enrich_pgx and
pgx_draft.draft_gene already take an injected client, so a reader with the same four methods needs
no branch in either — which is what keeps "snapshot or live" out of the passes entirely. It is not a
subclass because it shares no implementation: there is no endpoint, no pacing, no retry.
The vocabulary mapping happens here, not in the builder. The parquet holds CPIC's own prose
verbatim ("No function", "Strong") and map_function_status / map_classification run on the
way out — the same functions the live client calls, so a snapshot answer and a live answer are the
same object by construction, and a mapping fix reaches a snapshot built last month.
Read with duckdb, which is a core dependency, rather than polars, which is [dev] and only the
builder needs. That is the house convention (builder in polars, runtime pass in duckdb), and it is
what keeps the declared runtime dependency set honest about what the runtime actually requires.
Source code in enricher/src/just_dna_enricher/cpic.py
| def __init__(self, reference: Path) -> None:
self.reference = Path(reference)
data_dir = self.reference / SNAPSHOT_DATA_DIRNAME
if not data_dir.is_dir() or not any(data_dir.glob("*.parquet")):
raise CpicError(f"no CPIC snapshot at {data_dir} — build one with `just-dna-enricher cpic build`")
self._data_dir = data_dir
release_path = self.reference / RELEASE_FILENAME
self.release: dict[str, Any] = json.loads(release_path.read_text()) if release_path.is_file() else {}
|
dataset
property
Which snapshot answered — recorded onto SourceRow.dataset, as the gnomAD routes are.
A consumer must be able to tell a pinned file from a live API: the two can differ by a release,
and dataset is inside the fact set where source is not.
close
No-op, so a caller can close() this and a live client identically.
Source code in enricher/src/just_dna_enricher/cpic.py
| def close(self) -> None:
"""No-op, so a caller can `close()` this and a live client identically."""
|
recommendations
recommendations(
gene: str, drug: str
) -> list[CpicRecommendation]
One gene/drug pair's recommendations across every population.
gene_count = 1 reproduces the live client's len(phenotypes) != 1 filter exactly: the builder
flattens a recommendation's phenotypes map to one row per gene and carries how many there
were, because a row about CYP2C19 and CYP2D6 is not a statement about CYP2C19 alone.
Source code in enricher/src/just_dna_enricher/cpic.py
| def recommendations(self, gene: str, drug: str) -> list[CpicRecommendation]:
"""One gene/drug pair's recommendations across every population.
`gene_count = 1` reproduces the live client's `len(phenotypes) != 1` filter exactly: the builder
flattens a recommendation's `phenotypes` map to one row per gene and carries how many there
were, because a row about CYP2C19 *and* CYP2D6 is not a statement about CYP2C19 alone.
"""
rows = self._rows(
"recommendations.parquet",
"gene = ? AND drug = ? AND gene_count = 1",
[gene, drug.strip().lower()],
"gene, phenotype, drug, population, classification, recommendation, implication, activity_score",
)
out = [
CpicRecommendation(
gene=r["gene"],
phenotype=r["phenotype"],
drug=r["drug"],
population=(r["population"] or "").strip(),
classification=map_classification(r["classification"]),
recommendation=(r["recommendation"] or "").strip() or None,
implication=(r["implication"] or "").strip() or None,
activity_score=r["activity_score"] or None,
)
for r in rows
]
# Same deterministic order the live client emits (Principle 7).
return sorted(out, key=lambda r: (r.population, r.phenotype))
|
knows_drug
knows_drug(drug: str) -> bool | None
None — the snapshot cannot answer this, and saying False would be a false negative.
The live client reads CPIC's drug table, which lists every drug CPIC names. The snapshot
carries recommendations.parquet and nothing else, so the only drug names in it are the ones
that have a recommendation — which makes "absent from this file" and "absent from CPIC"
exactly the two things the caller is trying to tell apart. Answering False here would turn
a limitation of the snapshot into an assertion about the source, which is the shape S20 exists
to prevent, so it withholds and the caller's message says which file it read.
Building the drug table into the snapshot is the fix if this ever needs a definite answer; it
is not worth a column today, since the only consumer is one warning's wording.
Source code in enricher/src/just_dna_enricher/cpic.py
| def knows_drug(self, drug: str) -> bool | None:
"""**`None` — the snapshot cannot answer this, and saying `False` would be a false negative.**
The live client reads CPIC's `drug` table, which lists every drug CPIC names. The snapshot
carries `recommendations.parquet` and nothing else, so the only drug names in it are the ones
that *have* a recommendation — which makes "absent from this file" and "absent from CPIC"
exactly the two things the caller is trying to tell apart. Answering `False` here would turn
a limitation of the snapshot into an assertion about the source, which is the shape S20 exists
to prevent, so it withholds and the caller's message says which file it read.
Building the drug table into the snapshot is the fix if this ever needs a definite answer; it
is not worth a column today, since the only consumer is one warning's wording.
"""
return None
|
partner_genes
partner_genes(gene: str, drug: str) -> list[str]
The same question as the live client's, off the flattened table (S102).
The builder flattens one recommendation's phenotypes map to one row per gene and keeps
no recommendation id, so the rows of one guideline cannot be grouped back together here.
The answer is therefore every gene the drug's multi-gene rows are keyed on other than
gene, which is exact while no drug is keyed on more than one pair — and that is CPIC
today: measured on the 2026-09-02 snapshot, gene_count never exceeds 2 and no drug has
both a single-gene and a multi-gene shape. A three-gene guideline would still be named
correctly; two different pairs for one drug would be listed together, which is the one
reading this cannot make and the builder would need a column to give it.
Source code in enricher/src/just_dna_enricher/cpic.py
| def partner_genes(self, gene: str, drug: str) -> list[str]:
"""The same question as the live client's, off the flattened table (S102).
The builder flattens one recommendation's `phenotypes` map to one row per gene and keeps
no recommendation id, so the rows of one guideline cannot be grouped back together here.
The answer is therefore *every* gene the drug's multi-gene rows are keyed on other than
`gene`, which is exact while no drug is keyed on more than one pair — and that is CPIC
today: measured on the 2026-09-02 snapshot, `gene_count` never exceeds 2 and no drug has
both a single-gene and a multi-gene shape. A three-gene guideline would still be named
correctly; two *different* pairs for one drug would be listed together, which is the one
reading this cannot make and the builder would need a column to give it.
"""
named = self._rows(
"recommendations.parquet",
"gene = ? AND drug = ? AND gene_count > 1",
[gene, drug.strip().lower()],
"gene",
)
if not named:
return []
others = self._rows(
"recommendations.parquet",
"drug = ? AND gene_count > 1 AND gene <> ?",
[drug.strip().lower(), gene],
"DISTINCT gene",
"gene",
)
return [r["gene"] for r in others]
|
defining_variants
defining_variants(
gene: str,
) -> tuple[list[CpicDefiningVariant], list[str]]
Star-allele defining variants, with the same aggregated unusable-allele warnings.
The classification runs here rather than being stored, for the reason the whole snapshot
follows: unusable_allele_reason is a judgement this workspace makes about CPIC's value, so
freezing it into the parquet would pin one release's opinion into every snapshot built under it.
Source code in enricher/src/just_dna_enricher/cpic.py
| def defining_variants(self, gene: str) -> tuple[list[CpicDefiningVariant], list[str]]:
"""Star-allele defining variants, with the same aggregated unusable-allele warnings.
The classification runs here rather than being stored, for the reason the whole snapshot
follows: `unusable_allele_reason` is a *judgement this workspace makes* about CPIC's value, so
freezing it into the parquet would pin one release's opinion into every snapshot built under it.
"""
rows = self._rows(
"allele_definitions.parquet",
"gene = ?",
[gene],
"gene, allele, rsid, chrom, start, variant_allele",
"allele, start, variant_allele",
)
out: list[CpicDefiningVariant] = []
unusable: dict[str, list[str]] = {}
for r in rows:
allele_value = (r["variant_allele"] or "").strip().upper()
reason = unusable_allele_reason(allele_value)
if reason is not None:
unusable.setdefault(reason, []).append(f"{r['allele']}={allele_value}")
out.append(
CpicDefiningVariant(
gene=r["gene"] or gene,
allele=r["allele"] or "",
rsid=r["rsid"] or None,
chrom=r["chrom"] or None,
start=r["start"],
variant_allele=allele_value or None,
unusable=reason,
)
)
return out, _unusable_warnings(gene, unusable)
|
unusable_allele_reason
unusable_allele_reason(value: str) -> str | None
Why variantallele cannot become a HaplotypeRow.allele, or None when it can.
Two distinct reasons, and calling both "an IUPAC ambiguity code" was wrong. Drafting real CYP2D6
turned up DELTCT, AAAGGGGCG(2) and GGA(1) beside the genuine codes (R, S), and the message
announced all of them as ambiguity codes — a claim about the data that is simply false for a deletion
or a repeat, and it points an author at the wrong thing. The distinction matters for what happens
next, too: an ambiguous base is an uncertainty CPIC recorded and will never be expressible, while a
structural notation is a grammar gap (RM5) that a future release may widen to hold.
"ambiguity" — every character is a nucleotide or an IUPAC ambiguity code (R = A or G).
"symbolic" — a symbolic/structural allele that states no length. RM5 made the well-formed
ones authorable, so <DEL:1500> is now a perfectly good HaplotypeRow.allele and this function
answers None for it; a lengthless <DEL> still cannot be used, because the compiler refuses one
on haplotypes.csv in both modes. CPIC publishes neither spelling today (its own are DELTCT and
AAAGGGGCG(2), which are "notation"), so the arm is dead against the live source — it exists
because the classifier is shared, and because a caller that cannot explain a reason is a
KeyError waiting for the first source that does emit one.
"notation" — a deletion/insertion or repeat notation, not a nucleotide string at all.
A classification is not a verdict, and this function owes the verdict. It delegates the
classification to just_dna_format.alleles.non_nucleotide_reason — one definition, two callers,
each keeping its own wording — but its own name asks whether the value can become a
HaplotypeRow.allele, and since RM5 the answer for a well-formed symbolic allele is yes. Returning
the bare classification made pgx_draft skip a defining variant the model would have accepted,
while the message beside it said the value belongs in allele as written.
Source code in enricher/src/just_dna_enricher/cpic.py
| def unusable_allele_reason(value: str) -> str | None:
"""Why `variantallele` cannot become a `HaplotypeRow.allele`, or `None` when it can.
**Two distinct reasons, and calling both "an IUPAC ambiguity code" was wrong.** Drafting real CYP2D6
turned up `DELTCT`, `AAAGGGGCG(2)` and `GGA(1)` beside the genuine codes (`R`, `S`), and the message
announced all of them as ambiguity codes — a claim about the data that is simply false for a deletion
or a repeat, and it points an author at the wrong thing. The distinction matters for what happens
next, too: an ambiguous base is an *uncertainty CPIC recorded* and will never be expressible, while a
structural notation is a **grammar gap** (RM5) that a future release may widen to hold.
* `"ambiguity"` — every character is a nucleotide or an IUPAC ambiguity code (`R` = A or G).
* `"symbolic"` — a symbolic/structural allele **that states no length**. RM5 made the well-formed
ones authorable, so `<DEL:1500>` is now a perfectly good `HaplotypeRow.allele` and this function
answers `None` for it; a lengthless `<DEL>` still cannot be used, because the compiler refuses one
on `haplotypes.csv` in both modes. CPIC publishes neither spelling today (its own are `DELTCT` and
`AAAGGGGCG(2)`, which are `"notation"`), so the arm is dead against the live source — it exists
because the classifier is shared, and because a caller that cannot explain a reason is a
`KeyError` waiting for the first source that does emit one.
* `"notation"` — a deletion/insertion or repeat notation, not a nucleotide string at all.
**A classification is not a verdict, and this function owes the verdict.** It delegates the
*classification* to `just_dna_format.alleles.non_nucleotide_reason` — one definition, two callers,
each keeping its own wording — but its own name asks whether the value can *become* a
`HaplotypeRow.allele`, and since RM5 the answer for a well-formed symbolic allele is yes. Returning
the bare classification made `pgx_draft` skip a defining variant the model would have accepted,
while the message beside it said the value belongs in `allele` as written.
"""
reason = non_nucleotide_reason(value)
if reason != "symbolic":
return reason
return "symbolic" if symbolic_allele_defect(value) is not None else None
|
map_classification
map_classification(raw: str | None) -> str | None
CPIC's recommendation strength → this format's vocabulary, or None when it did not classify.
Source code in enricher/src/just_dna_enricher/cpic.py
| def map_classification(raw: str | None) -> str | None:
"""CPIC's recommendation strength → this format's vocabulary, or None when it did not classify."""
if not raw:
return None
return _CLASSIFICATION_MAP.get(raw.strip().lower())
|
map_function_status
map_function_status(raw: str | None) -> str | None
CPIC's prose → pgx.VALID_FUNCTION_STATUS, or None when it says something else.
The two possible … values collapse to uncertain_function: CPIC uses them for a call it is not
confident in, which is exactly what uncertain means here, and inventing new vocabulary members
for them would make every consumer's switch statement wrong.
Source code in enricher/src/just_dna_enricher/cpic.py
| def map_function_status(raw: str | None) -> str | None:
"""CPIC's prose → `pgx.VALID_FUNCTION_STATUS`, or None when it says something else.
The two `possible …` values collapse to `uncertain_function`: CPIC uses them for a call it is not
confident in, which is exactly what `uncertain` means here, and inventing new vocabulary members
for them would make every consumer's switch statement wrong.
"""
if not raw:
return None
return _FUNCTION_MAP.get(raw.strip().lower())
|
normalize_chrom
normalize_chrom(value: str | None) -> str | None
CPIC's gene.chr (chr10, chrX) → this workspace's bare contig name (10, X).
None for a blank or an unplaced spelling, rather than a guess — the same rule
chrom_from_accession follows in the PharmVar client.
Source code in enricher/src/just_dna_enricher/cpic.py
| def normalize_chrom(value: str | None) -> str | None:
"""CPIC's `gene.chr` (`chr10`, `chrX`) → this workspace's bare contig name (`10`, `X`).
`None` for a blank or an unplaced spelling, rather than a guess — the same rule
`chrom_from_accession` follows in the PharmVar client.
"""
text = (value or "").strip()
if not text:
return None
return text.removeprefix("chr").removeprefix("CHR").strip() or None
|