just_dna_enricher.clinvar_draft¶
just_dna_enricher.clinvar_draft ¶
Draft a gene panel's variants.csv rows from the ClinVar snapshot (0.5.1, RM26).
The provider that RM4 has been waiting for, in the shape the charter allows: it drafts rows a human
then owns, with no compile-time reference materialization and no injected reference in the compile
path. The compiler stays inject-only; this is the network tier, so the snapshot it reads is found on
the usual cache ladder and provisioned from HuggingFace when absent (--offline forbids that, and an
explicit --snapshot is taken as given). It used to be a required argument, which meant the published
snapshot could not reach an author at all: they had to build 4.4M records from a 200 MB VCF first.
Why these rows are partial, and why that is the honest answer rather than a limitation.
VariantRow.genotype is required and ClinVar publishes alleles, not genotypes. Whether carrying a
pathogenic allele once is informative — a carrier, an affected proband, neither — is zygosity
interpretation, and it follows from the condition's inheritance mode, not from the allele. ClinVar
cannot supply it and this provider must not invent it: writing A/G because the alt is G would be
a clinical claim the source never made. reference_examples/pathogenic_clinvar/ is a human having
made exactly that call, by hand, per row.
So genotype is left carrying vocab.TEMPLATE_PLACEHOLDER, which no mode compiles: the file is
authored, in place, in the right order, and loudly incomplete until a human has decided. Rows are
placed into their gene's block (group_by=("gene",)), because a panel is read gene by gene and
stubs stranded at the end of a 500-row file are what made this unpleasant enough to look like a
blocker.
Identity is filled whole or not at all. With an rsID, only the rsID goes in; without one, the full
chrom/start/ref/alts. Never a subset — a lone alts on a position-only row makes
derive_variant_key mint a VRS ga4gh:VA.… id instead of chrom:start:ref, so a partial coordinate
silently changes which variant the row is.
What it will not fill, each a rule rather than an omission:
weight,direction,effect_size,effect_measure,effect_allele— ClinVar publishes no effect statistic, and a weight is the author's model of the finding.trait_efo_id— ClinVar'sconditionis free text and MedGen, not EFO. Mapping it is inference.acmg_sf— a different list this package deliberately does not hold (see the roadmap).curator,method— the spec'sdefaults:block owns those.
ClinVarDraftError ¶
Bases: RuntimeError
A gene-panel draft could not be completed.
ClinVarDraftResult
dataclass
¶
ClinVarDraftResult(
reports: list[DraftReport] = list(),
warnings: list[str] = list(),
skipped: bool = False,
)
What a panel draft did.
added
property
¶
Rows added across every table this run wrote — variants and their studies.
Ask added_for when you mean one of them: since the provider began drafting grounding
evidence too, a bare total no longer answers "how many variants did I get".
multi_allelic_rsids ¶
rsIDs that name more than one distinct allele event in this selection.
An rsID is a position/multi-allelic-level tag, not a per-allele one: ClinVar lists
rs773443949 in HFE as both G>A and G>T. An rsid-only row cannot say which, so drafting one
row per record would write two identical rows — and de-duplicating them would silently drop a real
allele. Found by drafting an actual panel; these records take the coordinate identity instead.
The event is (chrom, start, ref, alt), and keying the site on ref was the bug (S41).
The predicate used to group by (rsid, chrom, start, ref) and fire on >1 alt inside that group,
which reads as "more than one alt at one position" and is not: an ordinary ClinVar dup/del mirror
pair — A>AT beside ATT>A at one position, the same event written from either side — lands in
two groups of one alt each, so the rsID was never flagged, both records reduced to the same rsid
signature, and append_partial_rows dropped the second as already_present. A differing ref
breaks an rsid-only identity exactly as thoroughly as a differing alt; the docstring's own claim
was the correct rule and the code was narrower than it. Measured on the 2026-06-27 snapshot over
BRCA1/BRCA2/ATM/MLH1/MSH2: 942 rsIDs flagged before, 1,589 after, and the 647 newly flagged are
exactly the 647 identities that were collapsing — 725 records, of which 187 dropped a
better-reviewed record than the one kept, since select_by_gene orders by ref before
review_stars DESC and the survivor is therefore an artifact of allele spelling.
Distinctness is over the whole event rather than over records, deliberately: two rows of the same
allele (a re-submission under a second variation_id) are one claim written twice and collapsing
them loses nothing, while coordinate identity would not separate them anyway. On that measurement
the two readings coincide — every multi-record rsID is also multi-allele — but only this one is
true by construction, and the other would flag rsIDs whose collapse the fix cannot repair.
Source code in enricher/src/just_dna_enricher/clinvar_draft.py
sole_expressible_genotype ¶
The one genotype a caller can emit at this locus, or None where zygosity is a real decision.
The placeholder exists to protect a judgement, and on a non-diploid contig there is none to
protect (S6). Carrying a pathogenic allele is a carrier state or an affected one depending on
the condition's inheritance mode, which ClinVar does not state — that is why genotype is stubbed
at all. But the mitochondrial genome is haploid and chrY outside the pseudoautosomal regions is
hemizygous: exactly one genotype is expressible per allele there, so the decision the stub is
holding open does not exist, and every consumer of draft_gene_panel was left to rediscover
independently that its natural "write both zygosities" fill is wrong for those rows. One did, at
264 mitochondrial loci in a genome-wide panel and 260 in a cardiac one, each asserting a second
copy that is not there.
Y is decided per locus, three-valued, through the same predicate the compiler's ploidy check
uses: PAR1 and PAR2 recombine with X and are diploid in every karyotype, and XG and SPRY3
straddle a boundary, so a gene- or contig-wide verdict is wrong for half of either. True
(diploid) and None (no PAR table for this build) both keep the placeholder — an undecided
question is not an answer, and the author still has one to make.
The allele written is the ALT: a variant row is an annotation about carrying the finding, and on a haploid contig carrying it is spelled with one allele. The heteroplasmy axis is a different question with its own table kind, which is why the aggregated notice at the draft's end says so rather than leaving the reading implicit.
Source code in enricher/src/just_dna_enricher/clinvar_draft.py
draft_gene_panel ¶
draft_gene_panel(
spec_dir: Path,
genes: Sequence[str],
*,
snapshot: Path | None = None,
clin_sig: frozenset[str] = DEFAULT_CLIN_SIG,
min_review_stars: int = 2,
max_citations: int = DEFAULT_MAX_CITATIONS,
declared_use: str = "unstated",
offline: bool = False,
download: bool = True,
dry_run: bool = False,
) -> ClinVarDraftResult
Draft variants.csv rows for one or more genes, leaving genotype for the human.
Re-runnable and additive: run it per gene as a panel grows. A variant already in the file — stub or filled — is reported, never re-added, because a partial row keys on the identity columns rather than on a key that runs through the placeholder.
min_review_stars defaults to 2 (multiple submitters, no conflicts). A panel that silently mixes
a 0-star "no assertion criteria" submission with a 3-star expert-panel review is worse than one
that says which floor it drew from, and the floor belongs in the author's hands.
The snapshot is found, then provisioned — it used to be a required argument. snapshot=None
resolves the usual cache ladder and, unless offline, downloads the published one
(download.ensure_clinvar_snapshot), the same shape enrich() and the gene-metrics pass use. Before
this, the published snapshot could not reach an author: they had to build 4.4M records from a 200 MB
VCF themselves, or already know the cache path. That matters most for the citations, which are what
make a drafted panel compilable at all (studies.csv is mandatory and the VCF carries no PMIDs) and
which now travel with the published snapshot.
Source code in enricher/src/just_dna_enricher/clinvar_draft.py
619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 | |