just_dna_enricher.gene_validity¶
just_dna_enricher.gene_validity ¶
enrich-gene-validity — the module's genes in, gene_validity.csv out (0.6, RM24).
The question gene_metrics.csv cannot answer. Constraint says how intolerant of variation a gene
looks; dosage sensitivity says whether losing a copy causes disease. Neither says whether variation
in this gene causes this disease, which is what a curated gene-validity assertion is for — and it
is the claim a clinical module most often rests on without recording.
Two submitters ship here, and they are different kinds of thing rather than two copies of one:
- ClinGen publishes expert-panel curations — one assertion per (gene, disease, mode of inheritance), each from a named Gene Curation Expert Panel working to a numbered SOP.
- GenCC publishes an aggregate of nineteen submitters (ClinGen among them, plus Orphanet,
PanelApp, several laboratories). The same gene–disease pair routinely carries several submitters at
different strengths, and that disagreement is the data — so
submitteris part of the key, not a note. Reducing it to one row would be the bare-triple mistake the ClinPGx cross-check paid for once.
Both vocabularies are mapped at this boundary, never stored verbatim: ClinGen writes Disputed
where GenCC writes Disputed Evidence, and ClinGen writes AD where GenCC writes
Autosomal dominant. A consumer filtering on one spelling would silently miss the other's rows. The
verbatim wording survives in classification_raw, so the mapping stays auditable — the same shape
clin_sig/clin_sig_raw already has, and the same builders-store-verbatim/readers-map rule ClinGen's
dosage codes follow.
Neither has an offline snapshot, so --offline makes this pass a no-op with a warning (ClinGen's
file is ~1 MB, GenCC's ~28 MB — small enough to fetch whole and too incidental to publish a snapshot
for). clingen.enrich_dosage_sensitivity is the shape copied, injected text and all.
HPO ships no route here, and both reasons were established by probe rather than assumed. Its
release declares terms:license https://hpo.jax.org/app/license; that URL answers HTTP 404 with a
JavaScript shell, and OBO Foundry records the licence as a bare label hpo with no SPDX id — so its
terms cannot be established from any machine-readable source, and an unestablished permission is not a
permission. Separately, genes_to_phenotype.txt is gene × HP feature × frequency, a different grain
from this table entirely, and genes_to_disease.txt's association_type
(MENDELIAN/POLYGENIC/UNKNOWN, 8,288 of 15,944 rows UNKNOWN) is a mechanism class, not an evidence
grade — putting it in classification would overload the axis. The row shape holds an HPO row
perfectly well; what is missing is a link this tier may take the data over.
GeneValidityError ¶
Bases: RuntimeError
A gene-validity fetch or parse failed in a way the caller must see.
GeneValidityUnavailable ¶
Bases: GeneValidityError
A submitter's export could not be fetched, so it was never actually asked (RM101).
The same split ClinGenUnavailable draws, and it is here because this module has the identical
shape: GeneValidityError covers both "could not fetch the gene-validity export" and "the
existing gene_validity.csv is invalid", and only the first means a source was asked. Found by
walking the passes rather than by a report — S37 named ClinGen's instance and this one is its
twin, which is why the repair is a walked registry and not two hand-picked sites.
ValidityAssertion
dataclass
¶
ValidityAssertion(
gene: str,
gene_id: str | None = None,
disease_id: str | None = None,
disease_label: str | None = None,
moi: str | None = None,
classification: str | None = None,
classification_raw: str | None = None,
classification_date: str | None = None,
submitter: str | None = None,
assertion_id: str | None = None,
report_url: str | None = None,
)
One parsed assertion, before it is scoped to a module's genes.
map_classification ¶
A submitter's own classification wording → a VALID_GENE_VALIDITY member, or None.
None for blank and for a wording this release does not model; the caller collects the unknown
into unmapped so it is reported once with a count rather than per row.
Source code in enricher/src/just_dna_enricher/gene_validity.py
read_curation_date ¶
A submitter's curation date → the canonical UTC spelling, or None when it cannot be read.
The two submitters spell one instant two ways — ClinGen 2024-03-14T16:00:00.000Z, GenCC
2018-03-30 13:31:56 — which normalize_utc_timestamp reconciles, and which is why the column is
canonicalized rather than stored verbatim: two spellings in a fact column hash as two facts.
It is applied here rather than left to the model's validator because the validator raises, and this column is one cell of one row of an export with 30,000 of them from nineteen independent submitters. One unreadable date would otherwise cost the whole pass, with a bare pydantic traceback. A date that cannot be read is an unknown: withhold the cell, keep the assertion, and report the value once — exactly what an unrecognised classification wording already does.
Source code in enricher/src/just_dna_enricher/gene_validity.py
map_inheritance ¶
A submitter's own mode-of-inheritance wording → a VALID_INHERITANCE_MODE member, or None.
Source code in enricher/src/just_dna_enricher/gene_validity.py
parse_clingen_validity ¶
parse_clingen_validity(
text: str, *, unmapped: set[str] | None = None
) -> tuple[list[ValidityAssertion], str]
Parse ClinGen's gene-validity CSV → (assertions, release label).
The file's layout is the reason this is not two lines of csv.DictReader: four preamble rows
(title, FILE CREATED: …, the webpage), then a +++ separator, then the real header, then a
second +++ separator, then the data. Anchoring on the header's first cell rather than on a row
index means a fifth preamble line does not silently shift every column.
FILE CREATED: is the only version the file carries — ClinGen publishes no numbered release for
it — so it becomes the dataset label, exactly as clingen.parse_curation_list uses that file's
own date line.
Source code in enricher/src/just_dna_enricher/gene_validity.py
parse_gencc ¶
parse_gencc(
text: str, *, unmapped: set[str] | None = None
) -> tuple[list[ValidityAssertion], str]
Parse GenCC's submission export → (assertions, release label).
GenCC publishes no release identifier at all, so the label is derived from the latest
submitted_run_date the file carries — the same class of answer ClinGen's FILE CREATED: gives,
read from the data because the file states nothing else. unknown when even that is absent, which
is honest rather than a fabricated date.
Every row is one submitter's assertion and they are all kept. The submitted_as_* columns are the
submitter's own pre-harmonization wording; the unprefixed ones are GenCC's normalization onto
MONDO/HPO, which is what this table wants — the harmonization is GenCC's published work, not an
inference of ours.
Source code in enricher/src/just_dna_enricher/gene_validity.py
fetch_validity_export ¶
Download a submitter's whole export.
One request for the whole file rather than a per-gene query, because neither submitter offers one and both files are small enough to hold (1 MB and 28 MB). The timeout is generous for the same reason: GenCC's export is a single 28 MB response and a 60-second budget times out on a slow link while the request is perfectly healthy.
Source code in enricher/src/just_dna_enricher/gene_validity.py
enrich_gene_validity ¶
enrich_gene_validity(
spec_dir: Path,
*,
source: str = CLINGEN_SOURCE,
mode: str = "best_effort",
offline: bool = False,
write: bool = True,
export_text: str | None = None,
url: str | None = None,
) -> GeneValidityResult
Add curated gene–disease assertions to gene_validity.csv for the genes variants.csv names.
Existing rows are authoritative and merged, never clobbered — the standing rule for every pass,
and the standing consequence with it: to regenerate after a machinery change, delete the file
first. Two rows are the same row when _merge_key says so: the source's own assertion_id where
it publishes one, else (gene, disease_id, moi, submitter, dataset) — the source's own grain
rather than a convenience, because ClinGen carries 59 (gene, disease) pairs whose two rows differ
only by mode of inheritance, and GenCC carries the same pair from several submitters at different
strengths.
A gene the submitter has not curated gets no row, and is reported in missing. That follows
clingen.enrich_dosage_sensitivity exactly and for its reason: a curating body's silence means
nobody has assessed the gene yet, which is not a fact about the gene, and writing a not_found
row would state one. It is also why strict is a report, not a refusal to have looked — both
submitters curate a subset by design.
offline makes the pass a no-op with a warning (skipped_offline), never a failure: neither
submitter ships a snapshot. An injected export_text still wins, because handing over bytes you
already hold is not egress.
Source code in enricher/src/just_dna_enricher/gene_validity.py
398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 | |