Skip to content

just_dna_enricher.pharmvar

just_dna_enricher.pharmvar

Live PharmVar queries — star-allele definitions and allele function (0.5).

PharmVar is the naming authority for the CYP star alleles, and the only PGx source here that is not part of the ClinPGx merger. It supplies two things this workspace has no other route to: haplotypes.csv material (which variants define *2) and allele_function.csv material (what that allele does), both with GRCh38 genomic coordinates and dbSNP ids already attached.

Access changed under us and the failure mode is silent. The API now requires a key, and every endpoint returns the same 401 {"errorMessage": "API Key is invalid or missing"} whether the key is absent, malformed, or passed under the wrong parameter name — so a wrong header looks exactly like a bad key. The parameter is Api-Key, a plain header with no X- prefix, documented per-endpoint in the service's own OpenAPI document (docs/vendor/pharmvar_api_docs.json) rather than in a securityDefinitions block, which is why it is easy to miss.

The key is personal. PharmVar's terms §2 make an account non-transferable, so the key is read from the environment and never written anywhere: not into a module, a recorded fixture, a log line, or a snapshot. PharmVarClient keeps it in the header only.

Rate limit: 2 requests/second. Low enough that per-allele fetching is hopeless and coarse endpoints are the only sane shape — /variants/gene/{symbol} returns a gene's whole variant set in one call. The unfiltered /alleles collection is ~25 MB and ignores a geneSymbol query parameter it does not define, so it is never used here; reference-sequence filtering is applied server-side instead.

Coordinates arrive as HGVS g. strings against several reference sequences at once — a transcript (NM_…:c.), a gene region (NG_…:g.) and the chromosome (NC_…:g.). Only the NC_ form is a genomic coordinate, and it is 1-based, which matches what this pipeline already stores (see the resolution table): the instinctive conversion to 0-based would introduce an off-by-one.

PharmVarError

Bases: RuntimeError

A PharmVar request failed in a way the caller must see (auth, or a broken query).

PharmVarVariant dataclass

PharmVarVariant(
    rsid: str | None,
    chrom: str | None,
    start: int | None,
    ref: str | None,
    alt: str | None,
)

One defining variant of a star allele, at a genomic coordinate.

usable property

usable: bool

Whether this carries a genomic coordinate we can key on.

PharmVarAllele dataclass

PharmVarAllele(
    gene: str,
    allele: str,
    function: str | None = None,
    activity_value: float | None = None,
    evidence_level: str | None = None,
    is_core: bool = True,
    variants: list[PharmVarVariant] = list(),
)

One star allele: its identity, its function, and the variants that define it.

PharmVarClient

PharmVarClient(
    endpoint: str = DEFAULT_PHARMVAR_ENDPOINT,
    *,
    api_key: str | None = None,
    client: Client | None = None,
    gate: PacingGate | None = None,
    timeout: float = 60.0,
)

Paced PharmVar REST client. The API key comes from the environment and is never persisted.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def __init__(
    self,
    endpoint: str = DEFAULT_PHARMVAR_ENDPOINT,
    *,
    api_key: str | None = None,
    client: httpx.Client | None = None,
    gate: PacingGate | None = None,
    timeout: float = 60.0,
) -> None:
    self.endpoint = endpoint.rstrip("/")
    # `.env` is where a developer actually keeps this key, and until now it only reached
    # `os.environ` as a side effect of some *other* call resolving a cache path. That worked for
    # `enrich_pgx` by accident and not at all for the snapshot builder, which resolves nothing —
    # `pharmvar build` reported "no PharmVar API key" on a machine that had one in `.env`. Loading
    # it where the key is read makes the credential path independent of unrelated calls.
    # An exported variable and a test's neutralizing `""` both win over the file, and nothing else
    # in the `.env` is exported (RM301).
    self._api_key = api_key or env_value(API_KEY_ENV)
    self._client = client or httpx.Client(timeout=timeout)
    self._owned = client is None
    self._gate = gate or PacingGate(interval=PHARMVAR_MIN_INTERVAL)

configured property

configured: bool

Whether a key is available. Absent is a skip, not an error — the pass degrades.

alleles_for_gene

alleles_for_gene(
    gene: str, *, include_sub_alleles: bool = False
) -> list[PharmVarAllele]

Every allele PharmVar defines for one gene, with its defining variants.

Sub-alleles (*2.001) are excluded by default: the core star is the identity this workspace keys on (AlleleFunctionRow.allele), and including them multiplies the table roughly threefold without changing any lookup.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def alleles_for_gene(self, gene: str, *, include_sub_alleles: bool = False) -> list[PharmVarAllele]:
    """Every allele PharmVar defines for one gene, with its defining variants.

    Sub-alleles (`*2.001`) are excluded by default: the core star is the identity this workspace
    keys on (`AlleleFunctionRow.allele`), and including them multiplies the table roughly threefold
    without changing any lookup.
    """
    payload = self._get(
        f"genes/{gene}",
        params={"exclude-sub-alleles": "true"} if not include_sub_alleles else {},
    )
    raw = payload.get("alleles") if isinstance(payload, dict) else payload
    alleles = [parse_allele(entry) for entry in (raw or [])]
    if not include_sub_alleles:
        alleles = [a for a in alleles if a.is_core]
    return alleles

all_genes

all_genes(
    *, include_sub_alleles: bool = False
) -> dict[str, list[PharmVarAllele]]

Every gene PharmVar defines, with its alleles — one request for the whole database.

/genes returns the same shape as /genes/{symbol} for all fifteen genes at once (1,173 core alleles, ~5 MB), which at 2 rps is the difference between one call and fifteen. That matters only for the snapshot builder, which is its only caller; the runtime passes stay gene-scoped because a module's own tables say which genes it is about.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def all_genes(self, *, include_sub_alleles: bool = False) -> dict[str, list[PharmVarAllele]]:
    """Every gene PharmVar defines, with its alleles — **one request** for the whole database.

    `/genes` returns the same shape as `/genes/{symbol}` for all fifteen genes at once (1,173 core
    alleles, ~5 MB), which at 2 rps is the difference between one call and fifteen. That matters
    only for the snapshot builder, which is its only caller; the runtime passes stay gene-scoped
    because a module's own tables say which genes it is about.
    """
    payload = self._get(
        "genes",
        params={"exclude-sub-alleles": "true"} if not include_sub_alleles else {},
    )
    entries = payload if isinstance(payload, list) else [payload]
    out: dict[str, list[PharmVarAllele]] = {}
    for entry in entries:
        symbol = str(entry.get("geneSymbol") or "")
        if not symbol:
            continue
        alleles = [parse_allele(a) for a in (entry.get("alleles") or [])]
        if not include_sub_alleles:
            alleles = [a for a in alleles if a.is_core]
        out[symbol] = alleles
    return out

PharmVarSnapshotClient

PharmVarSnapshotClient(reference: Path)

alleles_for_gene answered from an operator-built snapshot (RM38).

Duck-typed against PharmVarClient so enrich_pgx needs no branch, and configured is True because the question it answers — "can this leg run?" — is about reachability, not about a key. A snapshot exists or it does not, and locations.resolve_pharmvar_reference already decided that before this is constructed.

There is no all_genes here: that endpoint exists to build a snapshot, and a snapshot does not build itself. Read with duckdb (core) rather than polars ([dev], builder-only).

Source code in enricher/src/just_dna_enricher/pharmvar.py
def __init__(self, reference: Path) -> None:
    self.reference = Path(reference)
    data_dir = self.reference / SNAPSHOT_DATA_DIRNAME
    if not (data_dir / "alleles.parquet").is_file():
        raise PharmVarError(
            f"no PharmVar snapshot at {data_dir} — build one with "
            f"`just-dna-enricher pharmvar build` (it needs your own PHARMVAR_API_KEY; the "
            f"snapshot is never published, see the terms §2 note)"
        )
    self._data_dir = data_dir
    release_path = self.reference / RELEASE_FILENAME
    self.release: dict[str, Any] = json.loads(release_path.read_text()) if release_path.is_file() else {}

configured property

configured: bool

True: a resolved snapshot is the credential's stand-in, and it is already in hand.

dataset property

dataset: str | None

Which snapshot answered, for SourceRow.dataset.

genome_build property

genome_build: str

The assembly the stored coordinates are on. Recorded by the builder, never assumed.

close

close() -> None

No-op, so a caller can close() this and a live client identically.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def close(self) -> None:
    """No-op, so a caller can `close()` this and a live client identically."""

alleles_for_gene

alleles_for_gene(
    gene: str, *, include_sub_alleles: bool = False
) -> list[PharmVarAllele]

Every allele the snapshot holds for one gene, with its defining variants.

include_sub_alleles is accepted for signature parity and cannot be honoured: the builder fetches core alleles only, so asking for sub-alleles here would be answered with a subset while looking like a full answer. It raises instead — a wrong answer that looks right is the one outcome worth refusing.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def alleles_for_gene(self, gene: str, *, include_sub_alleles: bool = False) -> list[PharmVarAllele]:
    """Every allele the snapshot holds for one gene, with its defining variants.

    `include_sub_alleles` is accepted for signature parity and cannot be honoured: the builder
    fetches core alleles only, so asking for sub-alleles here would be answered with a subset while
    looking like a full answer. It raises instead — a wrong answer that looks right is the one
    outcome worth refusing.
    """
    if include_sub_alleles:
        raise PharmVarError(
            "the PharmVar snapshot holds core alleles only; sub-alleles need the live API. "
            "Rebuild with --sub-alleles, or run this leg online."
        )
    alleles_path = str(self._data_dir / "alleles.parquet").replace("'", "''")
    variants_path = str(self._data_dir / "variants.parquet").replace("'", "''")
    rows = self._query(
        f"SELECT gene, allele, function, activity_value, evidence_level "
        f"FROM read_parquet('{alleles_path}') WHERE gene = ? ORDER BY allele",
        [gene],
    )
    variants = self._query(
        f"SELECT allele, rsid, chrom, start, ref, alt FROM read_parquet('{variants_path}') "
        f"WHERE gene = ? ORDER BY allele, variant_index",
        [gene],
    )
    by_allele: dict[str, list[PharmVarVariant]] = {}
    for v in variants:
        by_allele.setdefault(v["allele"], []).append(
            PharmVarVariant(
                rsid=v["rsid"], chrom=v["chrom"], start=v["start"], ref=v["ref"], alt=v["alt"]
            )
        )
    return [
        PharmVarAllele(
            gene=r["gene"],
            allele=r["allele"],
            function=r["function"],
            activity_value=r["activity_value"],
            evidence_level=r["evidence_level"],
            is_core=True,
            variants=by_allele.get(r["allele"], []),
        )
        for r in rows
    ]

chrom_from_accession

chrom_from_accession(num: str) -> str | None

NC_000010 → 10, NC_000023 → X. None for anything off the primary assembly.

RefSeq numbers the chromosomes 1–22 then X (23), Y (24), MT (12920 as a special case). Returning None rather than guessing keeps an unplaced contig out of the coordinate space entirely.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def chrom_from_accession(num: str) -> str | None:
    """`NC_000010` → `10`, `NC_000023` → `X`. None for anything off the primary assembly.

    RefSeq numbers the chromosomes 1–22 then X (23), Y (24), MT (12920 as a special case). Returning
    None rather than guessing keeps an unplaced contig out of the coordinate space entirely.
    """
    index = int(num)
    if 1 <= index <= 22:
        return str(index)
    return {23: "X", 24: "Y"}.get(index)

parse_genomic_variant

parse_genomic_variant(
    payload: dict[str, Any],
    *,
    build: str = PHARMVAR_GENOME_BUILD,
) -> PharmVarVariant

One variants[] entry → a coordinate, when it names one on build.

A variant appears several times per allele, once per reference sequence. Only the NC_ genomic row yields a position; the transcript rows still carry the rsID, so the caller merges by rsID rather than discarding them.

And only the NC_ row whose referenceCollections names build does — PharmVar publishes both assemblies and lists GRCh37 first, so accepting any NC_ row stores the wrong coordinate for most variants. See PHARMVAR_GENOME_BUILD.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def parse_genomic_variant(payload: dict[str, Any], *, build: str = PHARMVAR_GENOME_BUILD) -> PharmVarVariant:
    """One `variants[]` entry → a coordinate, when it names one **on `build`**.

    A variant appears several times per allele, once per reference sequence. Only the `NC_` genomic
    row yields a position; the transcript rows still carry the rsID, so the caller merges by rsID
    rather than discarding them.

    And only the `NC_` row whose `referenceCollections` names `build` does — PharmVar publishes both
    assemblies and lists GRCh37 first, so accepting any `NC_` row stores the wrong coordinate for most
    variants. See `PHARMVAR_GENOME_BUILD`.
    """
    rsid = payload.get("rsId") or None
    match = _NC_HGVS.match(str(payload.get("hgvs") or payload.get("position") or ""))
    collections = payload.get("referenceCollections") or []
    if match is None or build not in collections:
        return PharmVarVariant(rsid=rsid, chrom=None, start=None, ref=None, alt=None)
    return PharmVarVariant(
        rsid=rsid,
        chrom=chrom_from_accession(match.group("num")),
        # 1-based, matching Ensembl and what resolution.csv stores. Do NOT subtract one.
        start=int(match.group("pos")),
        ref=match.group("ref"),
        alt=match.group("alt"),
    )

parse_allele

parse_allele(
    payload: dict[str, Any],
    *,
    build: str = PHARMVAR_GENOME_BUILD,
) -> PharmVarAllele

One /alleles entry → a PharmVarAllele.

activityScore is absent for most alleles and non-numeric for some ("n/a"), so it is parsed defensively rather than trusted — a non-numeric score becomes None instead of failing the allele.

Source code in enricher/src/just_dna_enricher/pharmvar.py
def parse_allele(payload: dict[str, Any], *, build: str = PHARMVAR_GENOME_BUILD) -> PharmVarAllele:
    """One `/alleles` entry → a `PharmVarAllele`.

    `activityScore` is absent for most alleles and non-numeric for some ("n/a"), so it is parsed
    defensively rather than trusted — a non-numeric score becomes None instead of failing the allele.
    """
    raw_score = payload.get("activityScore")
    try:
        activity = float(raw_score) if raw_score not in (None, "", "n/a") else None
    except (TypeError, ValueError):
        activity = None
    return PharmVarAllele(
        gene=str(payload.get("geneSymbol") or ""),
        allele=str(payload.get("alleleName") or ""),
        function=payload.get("function") or None,
        activity_value=activity,
        evidence_level=payload.get("evidenceLevel") or None,
        is_core=(payload.get("alleleType") or "").lower() != "sub",
        variants=_merge_variants(list(payload.get("variants") or []), build=build),
    )