just_dna_enricher.clinpgx_build¶
just_dna_enricher.clinpgx_build ¶
Build the ClinPGx summary-annotation snapshot ([dev]).
Turns the ClinPGx summaryAnnotations.zip bulk download into a flat parquet snapshot the pass reads
offline, following the ClinVar builder exactly: stream the archive, emit deterministically-sorted
parquet, and write a release.json recording provenance beside it.
One thing here is not in the ClinVar builder, and it is the point of the whole licensing design.
ClinPGx ships a LICENSE.txt inside the same archive as the data, so this builder extracts it and
records both the text and its sha256 into release.json. The pass then stamps that hash onto every
SourceRow it emits, which makes the recorded terms provably contemporaneous with the recorded data —
strictly stronger than a lookup in a table that was true once. Both halves of such a table went stale
inside a single release: api.pharmgkb.org was retired on 2026-07-20 and CPIC's licence page moved
when it merged into ClinPGx.
The grain is (annotation, genotype), not annotation. A ClinPGx summary annotation names a variant
and a drug in summary_annotations.tsv, then gives one row per genotype in
summary_ann_alleles.tsv — the large majority carry exactly three, and the calls can be opposed
(rs4149056/simvastatin reads "decreased response" for CC and CT, "increased" for TT). Flattening to
the annotation would throw away the axis the module keys on, so the two files are joined here. How
many carry three is a property of today's download, so it is counted per build and never written down
(the count this docstring used to carry outlived the file it was measured on by fourteen months).
The archive is identified before it is read, and the retired spelling is refused (RM175). PharmGKB
became ClinPGx on 2025-07-29 and renamed clinical annotations to summary annotations; the archive
followed, clinical_ann* becoming summary_ann* and Clinical Annotation ID becoming Summary
Annotation ID. clinicalAnnotations.zip was last written on 2025-07-05, is on no downloads page, and
the API still answers it 200 through a 303 to that frozen object — so a stale URL in a config
cannot be told from a live one at the HTTP layer, and every earlier build of this lane came out of a
snapshot fourteen months old. require_current_archive therefore refuses an archive carrying the old
member names by name, rather than letting a plausible parquet come out of it.
CREATED_*.txt is the release id. ClinPGx publishes no version number, and the archives are not
refreshed in lockstep — relationships.zip was a year newer than clinicalAnnotations.zip when this
was written, which RM175 later explained: that archive had stopped being rebuilt at all — so the
per-archive creation date is the only honest dataset label.
ClinPgxArchiveError ¶
Bases: RuntimeError
The archive handed to the builder is not the one ClinPGx publishes today.
ArchiveVintage
dataclass
¶
One published spelling of the annotation archive: its zip name, its members, its id column.
Two exist because the source renamed all four in one go (RM175), and a reader that knows only the current spelling can say "member missing" but not why it is missing.
ClinPgxUnavailable ¶
Bases: ClinPgxArchiveError
The ClinPGx download did not answer — the fetch failed, not the archive.
Worth the subclass here more than anywhere else: this lane's other failure mode is a retired filename that still answers 200 (RM175), so "could not reach it" and "reached the wrong thing" must not arrive as one type.
download_clinpgx_zip ¶
Stream the ClinPGx bulk archive to dest, returning (path, sha256).
Through net.stream_to_file since RM187, so the atomic rename, the retry and the translation are
the same rule every bulk download in this package obeys.
Source code in enricher/src/just_dna_enricher/clinpgx_build.py
read_license ¶
The LICENSE.txt ClinPGx bundles with the data, or None if the archive has none.
Read from the archive rather than fetched separately on purpose: the terms that govern these bytes are the ones shipped alongside them.
A member that is present and blank answers None, the same as no member at all: both builders
hash whatever this returns into license_sha256 and write it beside the data, and an empty file
would pin the terms to the hash of the empty string and hand the drafter a licence to read that
says nothing. Decided here rather than in the two callers, so the rule cannot hold in one build
and not the other.
Source code in enricher/src/just_dna_enricher/clinpgx_build.py
read_created_date ¶
The CREATED_<date>.txt marker — ClinPGx's only release identifier.
Source code in enricher/src/just_dna_enricher/clinpgx_build.py
require_current_archive ¶
CURRENT_ARCHIVE if this zip is the one ClinPGx publishes today, else a refusal that says why.
Three arms, three diagnoses (@answered-is-not-absent): the current spelling, the retired one,
and an archive that is neither. The middle arm is the reason this function exists — the retired
clinicalAnnotations.zip still answers 200 and still parses, so without a check by name a stale
URL builds a plausible parquet out of the database as it stood before the 2025-07-29 rename
(@specific-rejection: a generic "member missing" is a dead end where naming the rename is a fix).
Source code in enricher/src/just_dna_enricher/clinpgx_build.py
build_snapshot ¶
build_snapshot(
zip_path: Path,
out_dir: Path,
*,
source_url: str = DEFAULT_CLINPGX_URL,
source_sha256: str | None = None,
) -> ClinPgxBuildResult
summaryAnnotations.zip → clinpgx/data/annotations.parquet + release.json + LICENSE.txt.
Joins the summary table to its per-genotype child so the snapshot's grain is
(annotation, genotype) — the grain PharmVariantRow keys on. Rows are sorted by
(annotation_id, genotype) so a rebuild is byte-identical; built_at is the only per-run byte
and lives in release.json, outside the parquet.
Raises ClinPgxArchiveError on the retired clinicalAnnotations.zip, which parses fine and is
fourteen months stale (RM175) — see require_current_archive.
Source code in enricher/src/just_dna_enricher/clinpgx_build.py
233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 | |