Skip to content

just_dna_enricher.pubmind_build

just_dna_enricher.pubmind_build

Build the PubMind snapshot ([dev]) — derived, operator-built, and never published (RM134 § A).

PubMind (Wang & Wang, Nat Commun, doi:10.1038/s41467-026-76834-4) extracts variant–disease pathogenicity assertions from the literature with an LLM. It is a source, of the same kind as ClinVar: an authoritative annotation source. Nothing it produces may enter resolution.csv — its coordinates are PyEnsembl back-mappings of extracted text, and resolution.csv's authority column is a different word for a different thing (@source-vs-authority).

There is exactly one per-variant channel and it is the ANNOVAR-redistributed bulk table. The web API has two endpoints, neither takes a variant, and both state that per-record detail is withheld. So hg38_pubmind_db.txt.gz is the input, and its columns are VCF-style despite the ANNOVAR packaging: there is not a single - allele in the 2026-08-24 file, a one-base deletion appears as 1 1014264 1014265 CC C with the anchor base retained, and Start is the 1-based POS (@start-1based). A join therefore needs no coordinate translation. Whether the indels are left-normalized is not established, which is why they are stamped rather than silently mixed in.

Two thirds of the file is not a genotypable position, and every dropped row is counted. When PubMind recovers only a protein change from the text it back-maps through the transcript and writes out every codon that could encode it. A codon block differing at exactly one base is a single substitution wearing three letters, so it is decomposed onto that base and stamped derivation=codon; a block needing two or three simultaneous substitutions asserts a change to the protein, not to a position, and is dropped. Silent truncation reads as full coverage, so the drop counters and the kept count sum to the input row count and all of them land in release.json (@dont-discard-computed).

A contested coordinate keeps every PVID as its own row, and that is the finding. Consolidation into a PVID is keyed on the text the model extracted — gene symbol plus cDNA or protein change — never on a coordinate, so one physical variant fragments into many PVIDs whose verdicts disagree. Collapsing them would mean choosing a winner by an ordering nobody defined, which is mode() over an unsorted group and is what the deterministic-ordering rule bans outright. release.json records how many keys are contested and how far the multiplicity goes.

pubmind publish refuses, on the PharmVar precedent (@gated-source-caches). The ANNOVAR-shipped table publishes no data terms of its own: the software licence covers the software, the paper is CC BY-NC-ND, and only CHOP can say what the bytes are under. Unknown is not permissive (@no-named-licence), and a bulk file arriving under terms we cannot establish is not a file we may pass on. The command exists and refuses with the reason, because a missing command reads as an oversight somebody will helpfully add.

Builder-only: polars is a guarded [dev] import, exactly as in the sibling builders.

PubMindBuildError

Bases: RuntimeError

A PubMind snapshot could not be built from the table given.

PubMindUnavailable

Bases: PubMindBuildError

The ANNOVAR-distributed table could not be reached.

A subclass, so except PubMindBuildError keeps catching everything it did while a caller who wants to tell "the source is down" from "your table is malformed" can. That distinction has to be carried by the type: neither exc.__cause__ nor a message is pinned as an API, so a reword would flip a consumer's verdict from unchecked to "your data is wrong".

The subclassing makes a caller's except order load-bearing — list this arm before its parent or it is dead code (@client-exception-contract).

PubMindDownload dataclass

PubMindDownload(
    path: Path,
    sha256: str | None,
    url: str | None = None,
    etag: str | None = None,
    last_modified: str | None = None,
)

What a download established about the bytes — each half None when the server did not say.

All three of sha256, etag and last_modified are recorded because all three are available, and because an upstream revision then becomes a finding rather than a silent change of answer.

PubMindBuildResult dataclass

PubMindBuildResult(
    out_dir: Path,
    parquet_file: Path,
    input_rows: int,
    record_count: int,
    dropped: dict[str, int] = dict(),
    derivations: dict[str, int] = dict(),
    allele_keys: int = 0,
    multi_pvid_keys: int = 0,
    max_pvids_per_key: int = 0,
    contested_keys: int = 0,
    unparsable_score: int = 0,
    unparsable_confidence: int = 0,
    source_sha256: str | None = None,
    dataset: str | None = None,
)

Outcome of a build: the paths, the counts kept, and every count dropped.

download_pubmind_table

download_pubmind_table(
    dest: Path, url: str = DEFAULT_PUBMIND_URL
) -> PubMindDownload

Stream the ANNOVAR-redistributed table to dest (atomic .part rename).

Mirrors clinvar_build.download_clinvar_vcf, and additionally keeps the ETag and Last-Modified the server sends: PubMind publishes no version string of its own, so those two headers plus the sha256 are the whole of what pins which bytes a snapshot was built from.

Source code in enricher/src/just_dna_enricher/pubmind_build.py
def download_pubmind_table(dest: Path, url: str = DEFAULT_PUBMIND_URL) -> PubMindDownload:
    """Stream the ANNOVAR-redistributed table to `dest` (atomic `.part` rename).

    Mirrors `clinvar_build.download_clinvar_vcf`, and additionally keeps the `ETag` and
    `Last-Modified` the server sends: PubMind publishes no version string of its own, so those two
    headers plus the sha256 are the whole of what pins which bytes a snapshot was built from.
    """
    # **This handler used to leave its `.part` behind** — the only one of the eleven that forgot,
    # which is the drift a shared body ends (`@a-failed-fetch-is-not-a-no-op`).
    streamed = stream_to_file(
        dest,
        url,
        error_cls=PubMindUnavailable,
        what="the PubMind table",
        remedy=(
            "It is a single dated bulk file on a third party's server, so a move or a rotation "
            "looks exactly like this — pass `--table` with a copy you already hold."
        ),
    )
    return PubMindDownload(
        path=streamed.path,
        sha256=streamed.sha256,
        url=url,
        etag=streamed.etag,
        last_modified=streamed.last_modified,
    )

build_snapshot

build_snapshot(
    table: Path,
    out_dir: Path,
    *,
    source_url: str | None = None,
    source_sha256: str | None = None,
    source_etag: str | None = None,
    source_last_modified: str | None = None,
) -> PubMindBuildResult

Convert the ANNOVAR-redistributed PubMind table into data/pubmind.parquet + release.json.

Rows are emitted sorted by (chrom in karyotype order, start, ref, alt, pvid), so a rebuild from the same input is byte-identical (Principle 7); release.json's built_at is the only per-run-varying byte and lives outside the parquet.

Every provenance argument defaults to None rather than to the module's constants, and source_url in particular: only a caller that actually fetched can say where the bytes came from. Defaulting it to DEFAULT_PUBMIND_URL would have written the ANNOVAR URL into the release.json of a snapshot built from a file on local disk — asserting a provenance the build never established, in the one file whose whole job is to pin which bytes it was built from.

Source code in enricher/src/just_dna_enricher/pubmind_build.py
def build_snapshot(
    table: Path,
    out_dir: Path,
    *,
    source_url: str | None = None,
    source_sha256: str | None = None,
    source_etag: str | None = None,
    source_last_modified: str | None = None,
) -> PubMindBuildResult:
    """Convert the ANNOVAR-redistributed PubMind table into `data/pubmind.parquet` + `release.json`.

    Rows are emitted sorted by `(chrom in karyotype order, start, ref, alt, pvid)`, so a rebuild from
    the same input is byte-identical (Principle 7); `release.json`'s `built_at` is the only
    per-run-varying byte and lives outside the parquet.

    **Every provenance argument defaults to `None` rather than to the module's constants**, and
    `source_url` in particular: only a caller that actually fetched can say where the bytes came from.
    Defaulting it to `DEFAULT_PUBMIND_URL` would have written the ANNOVAR URL into the `release.json`
    of a snapshot built from a file on local disk — asserting a provenance the build never established,
    in the one file whose whole job is to pin which bytes it was built from.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the PubMind snapshot; install the publisher/dev surface "
            "with `pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    table = Path(table)
    if not table.exists():
        raise FileNotFoundError(f"PubMind table not found at {table}")
    out_dir = Path(out_dir)
    data_dir = out_dir / SNAPSHOT_DATA_DIRNAME
    data_dir.mkdir(parents=True, exist_ok=True)

    dropped = dict.fromkeys(PUBMIND_DROP_REASONS, 0)
    input_rows = 0
    unparsable_score = unparsable_confidence = 0
    records: list[dict] = []

    for row in _open_table(table):
        input_rows += 1
        chrom = (row[_COL_CHROM] or "").strip().removeprefix("chr")
        if chrom in ("M", "chrM"):
            chrom = "MT"
        if chrom not in _VALID_CHROMS:
            dropped["off_target_chrom"] += 1
            continue
        ref = (row[_COL_REF] or "").strip().upper()
        alt = (row[_COL_ALT] or "").strip().upper()
        if not (_ACGT_RE.match(ref) and _ACGT_RE.match(alt)):
            dropped["non_acgt"] += 1
            continue
        if ref == alt:
            dropped["ref_equals_alt"] += 1
            continue
        pvid = (row[_COL_PVID] or "").strip()
        if not pvid:
            dropped["no_pvid"] += 1
            continue
        position = (row[_COL_START] or "").strip()
        if not position.isdigit():
            dropped["unparsable_position"] += 1
            continue
        start = int(position)
        if len(ref) != len(alt):
            derivation = "indel"
        elif len(ref) == 1:
            derivation = "direct"
        else:
            differing = [i for i in range(len(ref)) if ref[i] != alt[i]]
            if len(differing) != 1:
                dropped["multi_substitution"] += 1
                continue
            index = differing[0]
            start, ref, alt = start + index, ref[index], alt[index]
            derivation = "codon"
        raw_sig = row[_COL_SIG] or None
        score, bad_score = _score(row.get(_COL_SCORE))
        confidence, bad_confidence = _confidence(row.get(_COL_CONFIDENCE))
        unparsable_score += bad_score
        unparsable_confidence += bad_confidence
        records.append(
            {
                "chrom": chrom,
                "start": start,
                "ref": ref,
                "alt": alt,
                "pvid": pvid,
                "clin_sig": normalize_clin_sig(raw_sig),
                "clin_sig_raw": raw_sig,
                "pathogenicity_score": score,
                "confidence": confidence,
                "derivation": derivation,
            }
        )

    # Two source rows can decompose onto one base with identical values in every column (the model
    # enumerated two ref codons for one amino-acid change). Collapse only on the *whole* row: a pair
    # differing anywhere — a different verdict, a different score — is two claims and stays two rows,
    # because the dedup key decides which columns may become several rows (`@dedup-key-decides-rows`).
    seen: set[tuple] = set()
    unique: list[dict] = []
    for record in records:
        signature = tuple(record.values())
        if signature in seen:
            dropped["identical_duplicate"] += 1
            continue
        seen.add(signature)
        unique.append(record)

    unique.sort(key=lambda r: (_CHROM_INDEX[r["chrom"]], r["start"], r["ref"], r["alt"], r["pvid"]))
    frame = pl.DataFrame(unique, schema=_schema())
    parquet_file = data_dir / PUBMIND_PARQUET
    frame.write_parquet(parquet_file, compression="zstd")

    derivations = dict.fromkeys(sorted(PUBMIND_DERIVATIONS), 0)
    by_key: dict[tuple[str, int, str, str], set[str | None]] = {}
    sigs_by_key: dict[tuple[str, int, str, str], set[str]] = {}
    for record in unique:
        derivations[record["derivation"]] += 1
        key = (record["chrom"], record["start"], record["ref"], record["alt"])
        by_key.setdefault(key, set()).add(record["pvid"])
        sigs_by_key.setdefault(key, set()).add(record["clin_sig"])

    result = PubMindBuildResult(
        out_dir=out_dir,
        parquet_file=parquet_file,
        input_rows=input_rows,
        record_count=frame.height,
        dropped=dropped,
        derivations=derivations,
        allele_keys=len(by_key),
        multi_pvid_keys=sum(1 for pvids in by_key.values() if len(pvids) > 1),
        max_pvids_per_key=max((len(p) for p in by_key.values()), default=0),
        contested_keys=sum(1 for sigs in sigs_by_key.values() if len(sigs) > 1),
        unparsable_score=unparsable_score,
        unparsable_confidence=unparsable_confidence,
        source_sha256=source_sha256 if source_sha256 is not None else _sha256_file(table),
    )
    result.dataset = None if result.source_sha256 is None else f"pubmind_{result.source_sha256[:12]}"
    _write_release_json(
        out_dir,
        result,
        source_url=source_url,
        source_etag=source_etag,
        source_last_modified=source_last_modified,
    )
    logger.info(
        "Built the PubMind snapshot: %d of %d row(s) kept → %s (dropped: %s; %d contested key(s) of "
        "%d, worst %d PVIDs)",
        result.record_count,
        result.input_rows,
        parquet_file,
        ", ".join(f"{name} {count}" for name, count in dropped.items()),
        result.contested_keys,
        result.allele_keys,
        result.max_pvids_per_key,
    )
    return result