Skip to content

just_dna_enricher.clinpgx_draft

just_dna_enricher.clinpgx_draft

Draft pharm_variants.csv rows from the ClinPGx snapshot (0.5, RM26) — the second provider.

The clean contrast to pgx_draft: every column PharmVariantRow requires is published, so this provider builds real rows and hands them to draft.append_rows unchanged. Nothing is stubbed and nothing is invented. Where pgx_draft had to skip what CPIC's grammar could not express, the work here is almost entirely re-spelling.

What it fills, and what it deliberately does not. It fills rsid, gene, genotype, drug, phenotype_category, annotation_id, evidence_level and a transcribed conclusion. It does not fill chrom/start/ref: the snapshot carries no coordinate, and even if it did, a coordinate authored here would be compared by resolution._verify against the table that supplied it. Resolution puts coordinates in resolution.csv, which is where they belong — and PharmVariantRow.variant_key is a @property over them, so filling them would make a row's identity depend on the enricher.

The key is all five parts. (variant_key, drug, genotype, phenotype_category, annotation_id), which is draft.natural_key's own answer for this model, so a re-run can never append a row the compiler would then reject. Indexing ClinPGx by the bare (variant, drug) triple is a real bug that has already been made once in this package.

One annotation names several drugs. drugs is ;-joined (antidepressants;citalopram;paroxetine is one annotation), and drug is singular, so one snapshot record becomes one row per drug. They share an annotation_id and key distinctly, which is correct: PharmGKB really is saying the same thing about three drugs.

One annotation also names several genes, and that one cannot become several rows. gene is ;-joined by the same dialect (PRSS53;VKORC1) but sits outside the dedup key, so copies would collide where the drug copies do not. --gene matches per member — it used to test the whole cell, silently dropping the 3 VKORC1 rows hiding inside PRSS53;VKORC1 — and the written cell is the member the request selects, or empty when nothing selects one. _authored_gene carries the argument.

Skipped, with a warning rather than a coercion: haplotype-keyed genotypes (*1, *1/*1) belong on DiplotypeRow, and symbolic alleles (del/del) carry no length. Both are the policy pgx_draft already set. The second reason changed in 0.6 and the distinction is worth keeping straight: the grammar now holds <DEL:1500>, so the block is no longer "the format cannot spell it" but "ClinPGx does not publish the length, and a lengthless symbolic allele is a rule the compiler drops". A provider must not write rows the next command in the documented workflow discards.

ClinPgxDraftResult dataclass

ClinPgxDraftResult(
    reports: list[DraftReport] = list(),
    warnings: list[str] = list(),
    skipped: bool = False,
)

What a draft run did.

draft_pharm_variants

draft_pharm_variants(
    spec_dir: Path,
    *,
    snapshot: Path,
    genes: Sequence[str] = (),
    drugs: Sequence[str] = (),
    min_evidence_level: str | None = None,
    declared_use: str = "unstated",
    dry_run: bool = False,
) -> ClinPgxDraftResult

Append ClinPGx annotations into pharm_variants.csv, never rewriting a row that is there.

Inject-only: snapshot is a path this function reads, never downloads (build it with just-dna-enricher clinpgx build). Re-runnable — narrow by --drug and run again as a module grows; a row already present is reported, not replaced.

Source code in enricher/src/just_dna_enricher/clinpgx_draft.py
def draft_pharm_variants(
    spec_dir: Path,
    *,
    snapshot: Path,
    genes: Sequence[str] = (),
    drugs: Sequence[str] = (),
    min_evidence_level: str | None = None,
    declared_use: str = "unstated",
    dry_run: bool = False,
) -> ClinPgxDraftResult:
    """Append ClinPGx annotations into `pharm_variants.csv`, never rewriting a row that is there.

    Inject-only: `snapshot` is a path this function reads, never downloads (build it with
    `just-dna-enricher clinpgx build`). Re-runnable — narrow by `--drug` and run again as a module
    grows; a row already present is reported, not replaced.
    """
    declared_use, declared_from = effective_declared_use(spec_dir, CLINPGX_TERMS, declared_use)  # S105
    skip_reason = check_declared_use(CLINPGX_TERMS, declared_use)
    if skip_reason:
        # Acquisition-time refusal: the terms are accepted by taking the data, so nothing is read.
        return ClinPgxDraftResult(warnings=[skip_reason], skipped=True)

    records, release = load_snapshot(snapshot)
    rows, warnings = _rows_from_snapshot(
        records, genes=genes, drugs=drugs, min_evidence_level=min_evidence_level
    )
    if not rows:
        return ClinPgxDraftResult(warnings=warnings + ["nothing matched; no rows drafted"])

    # Hoisted above the append so the licence closure can carry it (RM232): the row has to land
    # inside the table's commit, so everything the row states must be known before the write. The
    # read is pure; the warning it can raise stays below, gated on `not dry_run` as it was.
    license_path = Path(snapshot) / SNAPSHOT_LICENSE_FILENAME
    license_text = license_path.read_text(encoding="utf-8") if license_path.is_file() else None
    if not (license_text or "").strip():
        license_text = None
    commit_licence = licence_commit(
        sources=[CLINPGX_TERMS.source],
        spec_dir=spec_dir,
        dataset=release.get("dataset"),
        declared_use=declared_use,
        error=ClinPgxEnrichmentError,
        license_texts={CLINPGX_TERMS.source: license_text} if license_text else None,
    )
    reports = [
        append_rows(spec_dir, "pharm_variants.csv", rows, dry_run=dry_run, before_commit=commit_licence)
    ]
    if not dry_run:
        # A pass that consults a source must WRITE its SourceRow: the compile gate and
        # `manifest.sources` read sources.csv and nothing else, so a row that is only returned is a
        # source the module cannot account for.
        # **The terms are pinned to the same moment as the data, which is the whole point of
        # `license_sha256` and this call was not doing it (S44).** ClinPGx bundles its `LICENSE.txt`
        # inside `summaryAnnotations.zip` and `clinpgx_build` extracts it beside the parquet
        # precisely so a holder of the snapshot can read the terms without the archive — but the row
        # passed only `declared_use` and `dataset`, so a share-alike source was recorded with a null
        # hash and nothing tied the recorded terms to the text that governed the bytes. Read from the
        # snapshot rather than from `release.json`'s stated hash: the file is what the module is
        # actually claiming, and hashing it here means a tampered or truncated copy cannot pin to a
        # value it does not have. Absent stays `None` (the tri-state rule) — an older snapshot built
        # before the extractor has no licence file, and inventing a hash for it would be worse than
        # the null this fixes.
        # Present-but-blank reads as absent here too, and the warning is the reason this arm is not
        # left to `SourceTerms.row`'s normalization alone: an empty file is what a provisioning run
        # used to leave behind, and a drafter that silently recorded no hash for it would say nothing
        # about a snapshot whose terms cannot be pinned.
        if license_text is None:
            warnings.append(
                f"no readable {SNAPSHOT_LICENSE_FILENAME} in the snapshot, so the recorded ClinPGx "
                f"terms are not pinned to the text that governed them (license_sha256 stays empty). "
                f"Rebuild the snapshot with `just-dna-enricher clinpgx build` to extract it."
            )
        # The release label alone cannot see a cell edited after the draft, and this provider writes
        # `evidence_level` straight out of the snapshot that `clinpgx` then compares it against —
        # RM4's tautology, one source over (RM73). The digest restamp is driven by this provider's
        # `kind` now, not by remembering to call it.
        #
        # **`withdraw_stale_dataset` is new here (RM228).** This drafter recorded a `dataset` and
        # never withdrew a stale one, so widening a module from a newer ClinPGx snapshot left the
        # licence row naming the older release — a false claim, because `merge_sources_file` is
        # never-clobber and the row cannot name two releases. Every snapshot-drafting provider but
        # this one and `pgx_draft` already did it.
        warnings.extend(
            record_draft_provenance(
                provider=_PROVIDER,
                sources=[CLINPGX_TERMS.source],
                spec_dir=spec_dir,
                dataset=release.get("dataset"),
                covered=True,
                drafted=any(report.added for report in reports),
                declared_use=declared_use,
                error=ClinPgxEnrichmentError,
                license_texts={CLINPGX_TERMS.source: license_text} if license_text else None,
            )
        )
    return ClinPgxDraftResult(reports=reports, warnings=warnings)