just_dna_enricher.clinpgx_draft¶
just_dna_enricher.clinpgx_draft ¶
Draft pharm_variants.csv rows from the ClinPGx snapshot (0.5, RM26) — the second provider.
The clean contrast to pgx_draft: every column PharmVariantRow requires is published, so this
provider builds real rows and hands them to draft.append_rows unchanged. Nothing is stubbed and
nothing is invented. Where pgx_draft had to skip what CPIC's grammar could not express, the work
here is almost entirely re-spelling.
What it fills, and what it deliberately does not. It fills rsid, gene, genotype, drug,
phenotype_category, annotation_id, evidence_level and a transcribed conclusion. It does
not fill chrom/start/ref: the snapshot carries no coordinate, and even if it did, a
coordinate authored here would be compared by resolution._verify against the table that supplied
it. Resolution puts coordinates in resolution.csv, which is where they belong — and
PharmVariantRow.variant_key is a @property over them, so filling them would make a row's identity
depend on the enricher.
The key is all five parts. (variant_key, drug, genotype, phenotype_category, annotation_id),
which is draft.natural_key's own answer for this model, so a re-run can never append a row the
compiler would then reject. Indexing ClinPGx by the bare (variant, drug) triple is a real bug that
has already been made once in this package.
One annotation names several drugs. drugs is ;-joined (antidepressants;citalopram;paroxetine
is one annotation), and drug is singular, so one snapshot record becomes one row per drug. They
share an annotation_id and key distinctly, which is correct: PharmGKB really is saying the same
thing about three drugs.
One annotation also names several genes, and that one cannot become several rows. gene is
;-joined by the same dialect (PRSS53;VKORC1) but sits outside the dedup key, so copies would
collide where the drug copies do not. --gene matches per member — it used to test the whole cell,
silently dropping the 3 VKORC1 rows hiding inside PRSS53;VKORC1 — and the written cell is the
member the request selects, or empty when nothing selects one. _authored_gene carries the
argument.
Skipped, with a warning rather than a coercion: haplotype-keyed genotypes (*1, *1/*1) belong on
DiplotypeRow, and symbolic alleles (del/del) carry no length. Both are the policy pgx_draft
already set. The second reason changed in 0.6 and the distinction is worth keeping straight: the
grammar now holds <DEL:1500>, so the block is no longer "the format cannot spell it" but "ClinPGx
does not publish the length, and a lengthless symbolic allele is a rule the compiler drops". A
provider must not write rows the next command in the documented workflow discards.
ClinPgxDraftResult
dataclass
¶
ClinPgxDraftResult(
reports: list[DraftReport] = list(),
warnings: list[str] = list(),
skipped: bool = False,
)
What a draft run did.
draft_pharm_variants ¶
draft_pharm_variants(
spec_dir: Path,
*,
snapshot: Path,
genes: Sequence[str] = (),
drugs: Sequence[str] = (),
min_evidence_level: str | None = None,
declared_use: str = "unstated",
dry_run: bool = False,
) -> ClinPgxDraftResult
Append ClinPGx annotations into pharm_variants.csv, never rewriting a row that is there.
Inject-only: snapshot is a path this function reads, never downloads (build it with
just-dna-enricher clinpgx build). Re-runnable — narrow by --drug and run again as a module
grows; a row already present is reported, not replaced.
Source code in enricher/src/just_dna_enricher/clinpgx_draft.py
364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 | |