just_dna_enricher.pubmind¶
just_dna_enricher.pubmind ¶
PubMind snapshot reader — the runtime half of the source pubmind_build writes (RM134 § B).
The split is the one clinvar.py draws against clinvar_build.py: the builder is polars and lives
behind the [dev] extra, the runtime pass is duckdb and ships in the ordinary install
(@duckdb-vs-polars). Nothing here builds, downloads or writes.
This reader answers one question and deliberately not the other one. lookup_pubmind_calls
serves the concordance check in clinical.py — what does PubMind say about this allele — and there
is no lookup_loci beside it, unlike ClinVar's. That absence is structural rather than unfinished
work: PubMind's coordinates are PyEnsembl back-mappings of text an LLM extracted, so they are
annotation and never resolution, and resolution.csv's authority column is a different word for a
different thing (@source-vs-authority).
One allele can carry several records, and they are all returned. Consolidation upstream is keyed
on the text PubMind extracted — a gene symbol plus a cDNA or protein change — so one physical
allele fragments into several PVIDs whose verdicts genuinely disagree; the snapshot measured 72,121
allele keys carrying more than one, 35,742 of them disagreeing. Collapsing them here would be
mode() over an unsorted group, which the deterministic-ordering rule bans outright. The
multiplicity is a finding, and what to do with it is clinical._fold_pubmind_records' judgement —
this half only reads.
PubMindReferenceError ¶
Bases: FileNotFoundError
Raised when a provided PubMind reference has no usable parquet files.
A FileNotFoundError subclass for the reason ClinVarReferenceError is one: a caller bracketing
a snapshot open with except OSError keeps working, and a caller that wants to tell "no snapshot
here" from "this snapshot will not answer" has the type to do it with.
pubmind_dataset_label ¶
Which PubMind snapshot this reference carries, as a SourceRow.dataset value.
PubMind publishes no version string of its own, so pubmind_build derives the label from the
sha256 of the bytes it was built from and records it in release.json. Read back rather than
recomputed: recomputing would need the source table, which a deployment holding only the snapshot
does not have.
None when the snapshot cannot state its release at all — an absent or unreadable release.json,
or one recording no dataset. An unknown is withheld rather than written as a label something could
match, because a fabricated label is what would make a tautology guard skip a real comparison.
Source code in enricher/src/just_dna_enricher/pubmind.py
lookup_pubmind_calls ¶
lookup_pubmind_calls(
reference: Path,
alleles: list[tuple[str, int, str, str]],
) -> dict[tuple[str, int, str, str], list[dict]]
(chrom, start, ref, alt) -> [{pvid, clin_sig, clin_sig_raw, confidence, ...}].
Allele-exact, never by rsID, and the snapshot carries no rsID column to be tempted by: PubMind keys on extracted text and back-maps to coordinates, so a position-level tag would pool verdicts about different alleles at one locus into a single answer.
A list rather than one record, and every element of it is kept — see the module docstring. Ordered
by (chrom, start, ref, alt, pvid), which is the snapshot's own emitted order, so a caller
folding the list reads the same records in the same sequence on every run (Principle 7).