just_dna_enricher.pubmind_build¶
just_dna_enricher.pubmind_build ¶
Build the PubMind snapshot ([dev]) — derived, operator-built, and never published (RM134 § A).
PubMind (Wang & Wang, Nat Commun, doi:10.1038/s41467-026-76834-4) extracts variant–disease
pathogenicity assertions from the literature with an LLM. It is a source, of the same kind as
ClinVar: an authoritative annotation source. Nothing it produces may enter resolution.csv — its
coordinates are PyEnsembl back-mappings of extracted text, and resolution.csv's authority column
is a different word for a different thing (@source-vs-authority).
There is exactly one per-variant channel and it is the ANNOVAR-redistributed bulk table. The web
API has two endpoints, neither takes a variant, and both state that per-record detail is withheld. So
hg38_pubmind_db.txt.gz is the input, and its columns are VCF-style despite the ANNOVAR packaging:
there is not a single - allele in the 2026-08-24 file, a one-base deletion appears as
1 1014264 1014265 CC C with the anchor base retained, and Start is the 1-based POS
(@start-1based). A join therefore needs no coordinate translation. Whether the indels are
left-normalized is not established, which is why they are stamped rather than silently mixed in.
Two thirds of the file is not a genotypable position, and every dropped row is counted. When
PubMind recovers only a protein change from the text it back-maps through the transcript and writes
out every codon that could encode it. A codon block differing at exactly one base is a single
substitution wearing three letters, so it is decomposed onto that base and stamped derivation=codon;
a block needing two or three simultaneous substitutions asserts a change to the protein, not to a
position, and is dropped. Silent truncation reads as full coverage, so the drop counters and the kept
count sum to the input row count and all of them land in release.json (@dont-discard-computed).
A contested coordinate keeps every PVID as its own row, and that is the finding. Consolidation
into a PVID is keyed on the text the model extracted — gene symbol plus cDNA or protein change —
never on a coordinate, so one physical variant fragments into many PVIDs whose verdicts disagree.
Collapsing them would mean choosing a winner by an ordering nobody defined, which is mode() over an
unsorted group and is what the deterministic-ordering rule bans outright. release.json records how
many keys are contested and how far the multiplicity goes.
pubmind publish refuses, on the PharmVar precedent (@gated-source-caches). The ANNOVAR-shipped
table publishes no data terms of its own: the software licence covers the software, the paper is
CC BY-NC-ND, and only CHOP can say what the bytes are under. Unknown is not permissive
(@no-named-licence), and a bulk file arriving under terms we cannot establish is not a file we may
pass on. The command exists and refuses with the reason, because a missing command reads as an
oversight somebody will helpfully add.
Builder-only: polars is a guarded [dev] import, exactly as in the sibling builders.
PubMindBuildError ¶
Bases: RuntimeError
A PubMind snapshot could not be built from the table given.
PubMindUnavailable ¶
Bases: PubMindBuildError
The ANNOVAR-distributed table could not be reached.
A subclass, so except PubMindBuildError keeps catching everything it did while a caller who
wants to tell "the source is down" from "your table is malformed" can. That distinction has to be
carried by the type: neither exc.__cause__ nor a message is pinned as an API, so a reword
would flip a consumer's verdict from unchecked to "your data is wrong".
The subclassing makes a caller's except order load-bearing — list this arm before its parent
or it is dead code (@client-exception-contract).
PubMindDownload
dataclass
¶
PubMindDownload(
path: Path,
sha256: str | None,
url: str | None = None,
etag: str | None = None,
last_modified: str | None = None,
)
What a download established about the bytes — each half None when the server did not say.
All three of sha256, etag and last_modified are recorded because all three are available,
and because an upstream revision then becomes a finding rather than a silent change of answer.
PubMindBuildResult
dataclass
¶
PubMindBuildResult(
out_dir: Path,
parquet_file: Path,
input_rows: int,
record_count: int,
dropped: dict[str, int] = dict(),
derivations: dict[str, int] = dict(),
allele_keys: int = 0,
multi_pvid_keys: int = 0,
max_pvids_per_key: int = 0,
contested_keys: int = 0,
unparsable_score: int = 0,
unparsable_confidence: int = 0,
source_sha256: str | None = None,
dataset: str | None = None,
)
Outcome of a build: the paths, the counts kept, and every count dropped.
download_pubmind_table ¶
Stream the ANNOVAR-redistributed table to dest (atomic .part rename).
Mirrors clinvar_build.download_clinvar_vcf, and additionally keeps the ETag and
Last-Modified the server sends: PubMind publishes no version string of its own, so those two
headers plus the sha256 are the whole of what pins which bytes a snapshot was built from.
Source code in enricher/src/just_dna_enricher/pubmind_build.py
build_snapshot ¶
build_snapshot(
table: Path,
out_dir: Path,
*,
source_url: str | None = None,
source_sha256: str | None = None,
source_etag: str | None = None,
source_last_modified: str | None = None,
) -> PubMindBuildResult
Convert the ANNOVAR-redistributed PubMind table into data/pubmind.parquet + release.json.
Rows are emitted sorted by (chrom in karyotype order, start, ref, alt, pvid), so a rebuild from
the same input is byte-identical (Principle 7); release.json's built_at is the only
per-run-varying byte and lives outside the parquet.
Every provenance argument defaults to None rather than to the module's constants, and
source_url in particular: only a caller that actually fetched can say where the bytes came from.
Defaulting it to DEFAULT_PUBMIND_URL would have written the ANNOVAR URL into the release.json
of a snapshot built from a file on local disk — asserting a provenance the build never established,
in the one file whose whole job is to pin which bytes it was built from.
Source code in enricher/src/just_dna_enricher/pubmind_build.py
309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 | |