Skip to content

just_dna_enricher.drug_labels_build

just_dna_enricher.drug_labels_build

Build the regulator drug-label snapshot the label cross-check reads ([dev], RM166).

ClinPGx publishes drugLabels.zip on the same endpoint summaryAnnotations.zip comes from, and clinpgx_build downloads exactly one of the twelve archives that endpoint serves. This is the second: 59 KB holding LICENSE.txt, README.pdf, drugLabels.tsv (1,433 rows when probed on 2026-08-05) and drugLabels.byGene.tsv (238 rows, a pivot of the same labels by gene symbol — it carries no fact the label table does not, so nothing here reads it).

Its own release.json, never the annotation lane's. clinpgx_build's module docstring records that relationships.zip was a year newer than clinicalAnnotations.zip, so assuming ClinPGx's archives refresh in lockstep is a mistake this lane has already made once — and RM175 found the reason that particular gap was so wide, which was that ClinPGx had stopped rebuilding the annotation archive under that name altogether. drugLabels.zip carries its own CREATED_<date>.txt and it is what labels this snapshot — clinpgx_drug_labels_<date>, distinct from the annotation lane's clinpgx_<date> because the two archives are two surfaces with two denominators (@two-surfaces-two-denominators).

Every cell is stored verbatim. Seven of the fifteen columns are flag-shaped — a blank or one constant string (Prescribing Info, Cancer Genome) — and coercing them to booleans would have this builder decide that Biomarker Flag is one, which it is not: it carries three values, On FDA Biomarker List and Formerly on FDA Biomarker List beside the blank. So nothing is coerced, a blank becomes None, and the reader decides what a cell means (@verbatim-except-order).

No --offline, and the licence gate lives at the CLI. A builder's off-switch is passing the local archive instead of downloading one, and check_declared_use gates the fetch — which is the command, not this module (@acquisition-gate-is-not-a-read-gate). clinpgx build puts the gate in the command body for the same reason and this follows it rather than inventing a third answer.

DrugLabelBuildResult dataclass

DrugLabelBuildResult(
    out_dir: Path,
    parquet_path: Path,
    release_file: Path,
    label_count: int,
    created_date: str | None,
    dataset: str | None,
    source_url: str,
    source_sha256: str,
    license_sha256: str | None,
    regulators: list[str] = list(),
    testing_levels: list[str] = list(),
)

Where the snapshot landed, what it holds, and what it came from.

download_drug_labels_zip

download_drug_labels_zip(
    dest: Path, url: str = DEFAULT_DRUG_LABELS_URL
) -> tuple[Path, str]

Stream drugLabels.zip to dest (atomic .part rename), returning (path, sha256).

Core httpx with the hash taken while streaming, the shape every bulk builder here uses — download.py is HuggingFace snapshot provisioning and net.py is pacing for live API clients, so neither is what a one-file download reaches for.

httpx's exceptions do not leave this function (@client-exception-contract): a retired endpoint is a 404, and it must reach the CLI as DRUG-LABEL BUILD FAILED: … rather than as a raw HTTPStatusError traceback. The half-written .part goes with it.

Source code in enricher/src/just_dna_enricher/drug_labels_build.py
def download_drug_labels_zip(dest: Path, url: str = DEFAULT_DRUG_LABELS_URL) -> tuple[Path, str]:
    """Stream `drugLabels.zip` to `dest` (atomic `.part` rename), returning `(path, sha256)`.

    Core `httpx` with the hash taken while streaming, the shape every bulk builder here uses —
    `download.py` is HuggingFace snapshot provisioning and `net.py` is pacing for live API clients,
    so neither is what a one-file download reaches for.

    **`httpx`'s exceptions do not leave this function** (`@client-exception-contract`): a retired
    endpoint is a 404, and it must reach the CLI as `DRUG-LABEL BUILD FAILED: …` rather than as a raw
    `HTTPStatusError` traceback. The half-written `.part` goes with it.
    """
    streamed = stream_to_file(
        dest,
        url,
        error_cls=DrugLabelUnavailable,
        what="the ClinPGx drug-label archive",
        remedy="Pass --zip <archive> to build from a copy you already hold.",
    )
    return streamed.path, streamed.sha256

build_drug_label_snapshot

build_drug_label_snapshot(
    zip_path: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_DRUG_LABELS_URL,
    source_sha256: str | None = None,
) -> DrugLabelBuildResult

drugLabels.zip → data/drug_labels.parquet + LICENSE.txt + release.json.

Rows are sorted by label_id so a rebuild is byte-identical (Principle 7); built_at is the only per-run byte and it lives in release.json, outside the parquet.

Source code in enricher/src/just_dna_enricher/drug_labels_build.py
def build_drug_label_snapshot(
    zip_path: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_DRUG_LABELS_URL,
    source_sha256: str | None = None,
) -> DrugLabelBuildResult:
    """`drugLabels.zip` → `data/drug_labels.parquet` + `LICENSE.txt` + `release.json`.

    Rows are sorted by `label_id` so a rebuild is byte-identical (Principle 7); `built_at` is the only
    per-run byte and it lives in `release.json`, outside the parquet.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the drug-label snapshot; install the dev surface with "
            "`pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    csv.field_size_limit(_CSV_FIELD_LIMIT)
    zip_path, out_dir = Path(zip_path), Path(out_dir)
    if not zip_path.is_file():
        raise DrugLabelError(f"no drug-label archive at {zip_path}")

    try:
        with zipfile.ZipFile(zip_path) as archive:
            license_text = read_license(archive)
            created = read_created_date(archive)
            records = _rows(archive)
    except zipfile.BadZipFile as exc:
        raise DrugLabelError(f"{zip_path} is not a readable zip archive: {exc}") from exc
    if not records:
        raise DrugLabelError(f"{source_url} parsed to zero labels; refusing to record it as a snapshot")

    data_dir = out_dir / SNAPSHOT_DATA_DIRNAME
    data_dir.mkdir(parents=True, exist_ok=True)
    frame = pl.DataFrame(records, schema=_schema()).sort("label_id", nulls_last=True)
    parquet_path = data_dir / LABELS_PARQUET
    frame.write_parquet(parquet_path, compression="zstd")

    license_sha: str | None = None
    if license_text is not None:
        # Written before it is hashed, and hashed from what was written: `clinpgx_draft` reads this
        # file rather than `release.json`'s stated hash, so the two must describe the same bytes.
        license_path = out_dir / SNAPSHOT_LICENSE_FILENAME
        atomic_write_text(license_path, license_text)
        license_sha = "sha256:" + hashlib.sha256(license_path.read_bytes()).hexdigest()

    regulators = sorted({r["regulator"] for r in records if r["regulator"]})
    levels = sorted({r["testing_level"] for r in records if r["testing_level"]})
    digest = source_sha256 or _sha256_file(zip_path)
    dataset = f"{SOURCE_NAME}_drug_labels_{created}" if created else None
    release_file = _write_release_json(
        out_dir,
        source_url=source_url,
        source_sha256=digest,
        created_date=created,
        dataset=dataset,
        license_sha256=license_sha,
        label_count=len(records),
        regulators=regulators,
        testing_levels=levels,
    )

    logger.info(
        "Drug-label snapshot: %d label(s) from %d regulator(s) (%s) → %s",
        len(records),
        len(regulators),
        created or "undated",
        parquet_path,
    )
    return DrugLabelBuildResult(
        out_dir=out_dir,
        parquet_path=parquet_path,
        release_file=release_file,
        label_count=len(records),
        created_date=created,
        dataset=dataset,
        source_url=source_url,
        source_sha256=digest,
        license_sha256=license_sha,
        regulators=regulators,
        testing_levels=levels,
    )