Skip to content

just_dna_enricher.pubmind_draft

just_dna_enricher.pubmind_draft

Draft variants.csv rows from the PubMind snapshot (RM134 § C) — draft-panel --source pubmind.

A flag on the existing command, never a second command. draft-pubmind would write the same tables from the same gene argument, so it would carry its own copy of the genotype worklist, the placeholder guard, the dedup-against-the-file pass and the refusal summary — the four parts that are hard to get right and the four that have each been fixed here at least once. draft-clinpgx is a separate command because it writes different tables; this does not. So the machinery below is imported from clinvar_draft rather than copied, including the worklist seam whose once-only scoping was the bug RM71 removed: a second copy of that rule is the copy that goes stale.

PubMind publishes no gene column, and this module does not invent one. The snapshot is (chrom, start, ref, alt) and nothing else locational — a verdict about a position, keyed to the text an LLM extracted rather than to a locus. Turning --gene BRCA1 into a set of positions therefore needs a gene→locus map, and this repo deliberately holds none: the compiler's own gene/locus check is chromosome-granular precisely because gene boundaries are not ours to state. So the map is ClinVar's own per-record gene attribution, matched at the exact position — every position ClinVar records for the gene, with no clinical or review filter, since the map is a locus universe and not a selection. A min/max span over those positions would be a boundary nobody defined, and would write a gene cell that is a false claim wherever two genes overlap.

The consequence is a scope, stated rather than counted: a PubMind verdict at a position ClinVar has no record for cannot be reached by gene at all, and that class is not countable — attributing it to a gene is the very thing there is no map for. It includes the codon-decomposed offsets, which are PubMind's own back-mapping of a protein change and frequently land where nobody has filed a record.

Identity is the whole coordinate or nothing (@identity-whole-or-none). Most PubMind rows carry no rsID — the snapshot has no rsID column at all — so chrom/start/ref/alts go in together, and the row matches on the same five columns a ClinVar-drafted row does, so the two providers dedupe against each other at one event rather than writing it twice.

What is withheld is named (@unreachable-not-absent). Four classes never become a row — a contested key, a length-changing row, a call outside --clin-sig, and a confidence below the floor or not stated at all — and each is counted under exactly one reason and reported with that reason. PUBMIND_WITHHELD_REASONS is walked rather than restated, so candidates == drafted + Σ withheld is an equality over the set (@registry-completeness).

A contested key is withheld, not resolved. Where the PVIDs at one coordinate disagree about the call, choosing one means an ordering nobody defined — mode() over an unsorted group, which the deterministic-ordering rule bans outright. The multiplicity is the finding: the coordinate, its PVIDs and its competing calls are all reported, and no row is written. Contestation is decided over every PVID at the key before --clin-sig and --min-confidence are applied, because a filter that removes the dissenting record has picked the winner just as surely as mode() would.

Terms are unknown, and unknown warns rather than gates (@no-named-licence). check_declared_use is a gate on fetching, and its unknown branch skips — correctly, for a pass that would go and get data whose terms nobody can state. Nothing is fetched here: there is no ensure_pubmind_snapshot by design, the operator built the snapshot themselves with pubmind build, and refusing to read it would make that command's output a file nothing may consume. So the reason is reported in the source's own words and the draft proceeds, exactly as the GWAS Catalog pass does with the same null terms. What the unknown answer gates is publishing a module carrying those bytes, which is the redistribution axis nothing in this repo designs yet.

What it will not fill, each a rule rather than an omission:

  • clinvar, pathogenic, benign — all three are ClinVar flags by their own field descriptions, and a position appearing in ClinVar's gene map says nothing about whether this allele is in ClinVar (@field-description-is-a-claim).
  • phenotype — the ANNOVAR-redistributed channel carries no condition, and PubMind's per-record detail is withheld from it.
  • a study row — the same channel carries no PMID, so a PubMind draft grounds nothing and says so.
  • weight, direction, trait_efo_id, curator, method — as for ClinVar, and for the same reasons.

PubMindDraftError

Bases: RuntimeError

A PubMind panel draft could not be completed.

PubMindDraftResult dataclass

PubMindDraftResult(
    reports: list[DraftReport] = list(),
    warnings: list[str] = list(),
    candidates: int = 0,
    drafted: int = 0,
    withheld: dict[str, int] = dict(),
    mapped_positions: int = 0,
    spoken_positions: int = 0,
)

What a PubMind panel draft did — and, in equal detail, what it did not write.

accounts_for_every_candidate

accounts_for_every_candidate() -> bool

candidates == drafted + Σ withheld — the equality PUBMIND_WITHHELD_REASONS exists for.

Walked rather than listed, so a sixth reason cannot arrive without joining the sum.

Source code in enricher/src/just_dna_enricher/pubmind_draft.py
def accounts_for_every_candidate(self) -> bool:
    """`candidates == drafted + Σ withheld` — the equality `PUBMIND_WITHHELD_REASONS` exists for.

    Walked rather than listed, so a sixth reason cannot arrive without joining the sum.
    """
    return self.candidates == self.drafted + sum(self.withheld.values())

select_by_positions

select_by_positions(
    reference: Path, positions: Sequence[tuple[str, int]]
) -> list[dict]

Every PubMind record at one of these exact positions — every PVID, never a winner.

A range query per contig, narrowed to the exact positions afterwards: the alternative is an IN list of tens of thousands of positions, and the alternative to that is a span, which is the one thing this module refuses to infer. Ordered deterministically so a re-draft offers the same rows in the same order (Principle 7).

Source code in enricher/src/just_dna_enricher/pubmind_draft.py
def select_by_positions(reference: Path, positions: Sequence[tuple[str, int]]) -> list[dict]:
    """Every PubMind record at one of these exact positions — **every** PVID, never a winner.

    A range query per contig, narrowed to the exact positions afterwards: the alternative is an `IN`
    list of tens of thousands of positions, and the alternative to *that* is a span, which is the one
    thing this module refuses to infer. Ordered deterministically so a re-draft offers the same rows
    in the same order (Principle 7).
    """
    wanted = {(normalize_chrom(str(chrom)), int(start)) for chrom, start in positions}
    if not wanted:
        return []
    by_chrom: dict[str, list[int]] = {}
    for chrom, start in sorted(wanted):
        by_chrom.setdefault(chrom, []).append(start)
    # The reader is shared with the concordance pass, so it raises the reader's own
    # `PubMindReferenceError`. A pass owes its caller its OWN type (RM101): translated here rather
    # than left to leak, which is what the two exception-contract tests walk the package for.
    try:
        con = _connect(reference)
    except PubMindReferenceError as exc:
        raise PubMindDraftError(
            f"{exc} Draft with `--source pubmind` needs the snapshot this reads."
        ) from exc
    rows: list[dict] = []
    try:
        for chrom in sorted(by_chrom):
            starts = by_chrom[chrom]
            cursor = con.execute(
                f"SELECT {_SELECT} FROM pubmind WHERE chrom = ? AND start BETWEEN ? AND ? "
                f"ORDER BY start, ref, alt, pvid",
                [chrom, min(starts), max(starts)],
            )
            columns = [d[0] for d in cursor.description]
            rows.extend(
                record
                for record in (dict(zip(columns, row, strict=True)) for row in cursor.fetchall())
                if (normalize_chrom(str(record["chrom"])), int(record["start"])) in wanted
            )
    finally:
        con.close()
    return rows

gene_positions

gene_positions(
    reference: Path, genes: Sequence[str]
) -> dict[tuple[str, int], set[str]]

(chrom, start) → the requested gene(s) ClinVar attributes a record at that position to.

No clinical or review filter, deliberately: this is a locus universe rather than a selection, and the dials an author sets are about which PubMind verdicts to take. Filtering the map instead would silently narrow which positions PubMind is even asked about, and the narrower map is invisible in the output.

A position ClinVar attributes to two of the requested genes keeps both, so nothing here picks a gene either.

Source code in enricher/src/just_dna_enricher/pubmind_draft.py
def gene_positions(reference: Path, genes: Sequence[str]) -> dict[tuple[str, int], set[str]]:
    """`(chrom, start)` → the requested gene(s) ClinVar attributes a record at that position to.

    **No clinical or review filter**, deliberately: this is a locus universe rather than a selection,
    and the dials an author sets are about which *PubMind* verdicts to take. Filtering the map instead
    would silently narrow which positions PubMind is even asked about, and the narrower map is
    invisible in the output.

    A position ClinVar attributes to two of the requested genes keeps both, so nothing here picks a
    gene either.
    """
    mapping: dict[tuple[str, int], set[str]] = {}
    for record in select_by_gene(reference, list(genes), clin_sig=None, min_review_stars=0):
        chrom, start = record.get("chrom"), record.get("start")
        gene = (record.get("gene") or "").strip()
        if chrom is None or start is None or not gene:
            continue
        # Both sides through the same normalizer: the snapshots agree today (each builder strips a
        # `chr` prefix and folds M to MT), and a map keyed on the raw cell would silently find
        # nothing the day one of them stops agreeing.
        mapping.setdefault((normalize_chrom(str(chrom)), int(start)), set()).add(gene)
    return mapping

draft_gene_panel_from_pubmind

draft_gene_panel_from_pubmind(
    spec_dir: Path,
    genes: Sequence[str],
    *,
    snapshot: Path | None = None,
    pubmind_snapshot: Path | None = None,
    clin_sig: frozenset[str] = DEFAULT_CLIN_SIG,
    min_confidence: int = DEFAULT_MIN_CONFIDENCE,
    declared_use: str = "unstated",
    offline: bool = False,
    download: bool = True,
    dry_run: bool = False,
) -> PubMindDraftResult

Draft variants.csv rows for one or more genes from PubMind, leaving genotype to the human.

Two snapshots, and each is doing a different job: snapshot is the ClinVar one, read for its per-record gene attribution and nothing else, and pubmind_snapshot is the operator-built PubMind one the verdicts come from. Neither is optional, and the messages say which is missing — an author who has one and not the other should not have to guess.

Re-runnable and additive, exactly as the ClinVar path is: a key already in the file — stub or filled — is reported rather than re-added, and the open-stub worklist is read off the file rather than off this run, so asking twice answers twice.

Source code in enricher/src/just_dna_enricher/pubmind_draft.py
def draft_gene_panel_from_pubmind(
    spec_dir: Path,
    genes: Sequence[str],
    *,
    snapshot: Path | None = None,
    pubmind_snapshot: Path | None = None,
    clin_sig: frozenset[str] = DEFAULT_CLIN_SIG,
    min_confidence: int = DEFAULT_MIN_CONFIDENCE,
    declared_use: str = "unstated",
    offline: bool = False,
    download: bool = True,
    dry_run: bool = False,
) -> PubMindDraftResult:
    """Draft `variants.csv` rows for one or more genes from PubMind, leaving `genotype` to the human.

    Two snapshots, and each is doing a different job: `snapshot` is the **ClinVar** one, read for its
    per-record gene attribution and nothing else, and `pubmind_snapshot` is the operator-built PubMind
    one the verdicts come from. Neither is optional, and the messages say which is missing — an
    author who has one and not the other should not have to guess.

    Re-runnable and additive, exactly as the ClinVar path is: a key already in the file — stub or
    filled — is reported rather than re-added, and the open-stub worklist is read off the file rather
    than off this run, so asking twice answers twice.
    """
    build_warning = source_build_mismatch(spec_dir, "the PubMind snapshot", PUBMIND_GENOME_BUILD)
    warnings: list[str] = []

    if pubmind_snapshot is not None:
        reference = Path(pubmind_snapshot)
    else:
        resolved = resolve_pubmind_reference()
        if resolved is None:
            raise PubMindDraftError(
                "no PubMind snapshot found. It is operator-built and never published — the terms of "
                "the ANNOVAR-redistributed table could not be established, so nothing here passes it "
                "on. Build one with `just-dna-enricher pubmind build --download` and point "
                "$JUST_DNA_PUBMIND_CACHE at it, or pass --pubmind-cache PATH."
            )
        reference = Path(resolved)

    # A pass owes its caller its own exception type (`@client-exception-contract`), and the ClinVar
    # ladder raises ClinVar's. Translated rather than left to leak, with the reason the second
    # snapshot is wanted at all — an author told "no ClinVar snapshot" by a PubMind command has been
    # handed a puzzle rather than a diagnosis.
    try:
        clinvar_reference, provisioning = _resolve_snapshot(snapshot, offline=offline, download=download)
    except ClinVarDraftError as exc:
        raise PubMindDraftError(
            f"{exc} A ClinVar snapshot is needed even for a PubMind draft: PubMind's table names no "
            f"gene, so a coordinate is attributed to one through ClinVar's own per-record gene "
            f"attribution and through nothing else."
        ) from exc
    warnings.extend(provisioning)
    if build_warning:
        warnings.append(build_warning)
    # Unknown terms warn and never gate (`@no-named-licence`) — see the module docstring for why
    # `check_declared_use`, which is a gate on *fetching*, is not the right question for a snapshot
    # the operator built themselves.
    warnings.append(
        f"{PUBMIND_SOURCE}: the terms of the ANNOVAR-redistributed table could not be established, "
        f"so every licence cell on this module's {PUBMIND_SOURCE} row is null. Unknown is neither "
        f"permission nor refusal: the compile does not refuse a null, and what the answer would "
        f"govern is publishing a module carrying these values."
    )

    positions = gene_positions(clinvar_reference, genes)
    if not positions:
        return PubMindDraftResult(
            warnings=warnings
            + [
                f"ClinVar records no position for {', '.join(genes)}, so there is no gene map to ask "
                f"PubMind with. PubMind's table names no gene, so a coordinate can only be attributed "
                f"to one through a source that does — check the symbol, or draft by hand."
            ],
            withheld=dict.fromkeys(PUBMIND_WITHHELD_REASONS, 0),
        )

    keys = _group_by_key(select_by_positions(reference, list(positions)))
    withheld = dict.fromkeys(PUBMIND_WITHHELD_REASONS, 0)
    contested: list[_Key] = []
    indels: list[_Key] = []
    ambiguous_gene: list[str] = []
    partials: list[PartialRow] = []
    record_by_signature: dict[tuple[str, ...], dict] = {}
    stubbed_by_signature: dict[tuple[str, ...], tuple[str, ...]] = {}
    for key in keys:
        reason = _withhold_reason(key, clin_sig=clin_sig, min_confidence=min_confidence)
        if reason is not None:
            withheld[reason] += 1
            if reason == "contested_key":
                contested.append(key)
            elif reason == "indel_derivation":
                indels.append(key)
            continue
        genes_here = positions[(key.chrom, key.start)]
        if len(genes_here) > 1:
            ambiguous_gene.append(f"{key.label} ({'/'.join(sorted(genes_here))})")
        cells = _row_cells(key, genes_here)
        # `state` is stubbed for the same reason it is on the ClinVar path: PubMind states a call, and
        # `VALID_STATES` has no member meaning "undecided", so a call the fold does not cover is the
        # human's. Derived from the cells rather than listed, so a column this provider filled — the
        # haploid genotype — is not then stubbed over the top of itself.
        stubbed = tuple(column for column in ("genotype", "state") if column not in cells)
        signature = _signature(cells)
        record_by_signature[signature] = {
            "chrom": key.chrom,
            "start": key.start,
            "ref": key.ref,
            "alt": key.alt,
            "clin_sig": key.calls[0],
        }
        stubbed_by_signature[signature] = stubbed
        partials.append(
            PartialRow(
                model=VariantRow,
                cells=cells,
                stubbed=stubbed,
                # The same five columns the ClinVar provider matches on, so a coordinate PubMind and
                # ClinVar both speak about is one row in the file rather than two.
                match_on=_MATCH_ON,
            )
        )

    result = PubMindDraftResult(
        warnings=warnings,
        candidates=len(keys),
        drafted=len(partials),
        withheld=withheld,
        mapped_positions=len(positions),
        spoken_positions=len({(k.chrom, k.start) for k in keys}),
    )
    result.warnings.extend(_withheld_warnings(withheld, contested, min_confidence, clin_sig, indels))
    if ambiguous_gene:
        result.warnings.append(
            f"{len(ambiguous_gene)} coordinate(s) are attributed to more than one of the genes you "
            f"asked for, so their `gene` cell is empty rather than naming one or naming both: "
            f"{examples(ambiguous_gene)}. Both readings are wrong to write — a joined pair is not a "
            f"symbol any lookup resolves, and choosing between them is the gene model this pass goes "
            f"to ClinVar precisely to avoid inventing. The row is drafted; the coordinate is its "
            f"identity, and `gene` is yours to fill if you want one."
        )
    # Absence is not disagreement and is not silence: a position PubMind says nothing about means no
    # paper in the corpus survived their triage, never that the literature is quiet.
    if result.spoken_positions < result.mapped_positions:
        result.warnings.append(
            f"PubMind states nothing at {result.mapped_positions - result.spoken_positions} of the "
            f"{result.mapped_positions} position(s) ClinVar attributes to {', '.join(genes)}. That is "
            f"an absence in their corpus, not a benign call and not a disagreement — and it is only "
            f"the positions ClinVar records: a PubMind verdict somewhere ClinVar has no record for "
            f"cannot be reached by gene at all, because nothing here maps a bare coordinate to one."
        )
    if not partials:
        result.warnings.append("nothing matched; no rows drafted")
        return result

    # Read here rather than at the tail (RM232), for the reason `clinvar_draft` states: the licence
    # row lands inside the table's commit, so the row's contents precede the write. The warning for a
    # snapshot that cannot state its release stays below.
    dataset = pubmind_dataset_label(reference)
    commit_licence = licence_commit(
        sources=[PUBMIND_SOURCE],
        spec_dir=spec_dir,
        dataset=dataset,
        declared_use=declared_use,
        error=PubMindDraftError,
    )
    report = append_partial_rows(
        spec_dir,
        "variants.csv",
        partials,
        group_by=("gene",),
        dry_run=dry_run,
        before_commit=commit_licence,
    )
    result.reports.append(report)
    result.warnings.extend(_refusal_summary(report.invalid))
    # PubMind's own channel carries no citation, so a drafted panel is ungrounded by construction and
    # has to say so: `studies.csv` is mandatory, and a module that cannot compile without telling the
    # author why is the defect the ClinVar path already fixed once.
    result.warnings.append(
        f"no studies.csv rows were drafted: the ANNOVAR-redistributed {PUBMIND_LABEL} table carries "
        f"no PMID, and their API withholds per-record detail. Grounding evidence is mandatory, so add "
        f"it by hand, or draft the same genes from ClinVar, which publishes its literature links."
    )

    # ── What is still open, scoped to the FILE rather than to this run ──────────────────────────
    #
    # `_open_stubs` is imported rather than reimplemented: scoping this to `report.added` is the
    # once-only defect RM71 removed on the ClinVar path, where a second run added nothing and
    # therefore said nothing, and a second copy of the rule is the copy that goes stale.
    genotype_stubs = _open_stubs(report, "genotype", stubbed_by_signature)
    if genotype_stubs:
        result.warnings.append(
            f"{len(genotype_stubs)} row(s) carry an unreplaced genotype placeholder and will not "
            f"compile until you decide the zygosity each finding is about."
        )
        result.warnings.extend(
            _genotype_worklist(
                [
                    record_by_signature[signature]
                    for signature, _ in genotype_stubs
                    if signature in record_by_signature
                ],
                source=PUBMIND_LABEL,
            )
        )
        withheld_alleles = [
            signature for signature, _ in genotype_stubs if signature not in record_by_signature
        ]
        if withheld_alleles:
            result.warnings.append(
                f"{len(withheld_alleles)} row(s) of that list carry alleles this run cannot state, "
                f"because nothing it selected covers them "
                f"({examples([':'.join(p for p in s if p) for s in withheld_alleles])}). The alleles "
                f"are withheld rather than guessed: draft the gene each row records, or "
                f"`hint variant`, will state them."
            )
    result.warnings.extend(
        _state_stub_warnings(
            _open_stubs(report, "state", stubbed_by_signature),
            record_by_signature,
            source=PUBMIND_LABEL,
        )
    )

    if dataset is None:
        result.warnings.append(
            f"this snapshot does not say which {PUBMIND_LABEL} bytes it carries (no readable "
            f"{RELEASE_FILENAME}), so the licence row records no dataset: nothing downstream can tell "
            f"these rows were copied out of it. Rebuild it with `just-dna-enricher pubmind build`, "
            f"which writes the release file this reads."
        )
    if not dry_run:
        # `covered=True` reproduces this provider exactly; like `clinvar_draft` it writes the row on
        # any non-dry run rather than on having covered something. Recorded, not changed, under a
        # behaviour-preserving migration (RM228).
        result.warnings.extend(
            record_draft_provenance(
                provider=_PROVIDER,
                sources=[PUBMIND_SOURCE],
                spec_dir=spec_dir,
                dataset=dataset,
                covered=True,
                drafted=bool(report.added),
                declared_use=declared_use,
                error=PubMindDraftError,
                stale_warning=lambda superseded, ds: (
                    f"this module already recorded rows drafted from {superseded}, and these came "
                    f"from {ds or 'a snapshot that does not state its release'} — so the licence "
                    f"row's dataset has been cleared rather than re-labelled: it cannot name two "
                    f"releases, and naming one would be a claim about rows that did not come from it."
                ),
            )
        )
    return result