just_dna_enricher.enrich¶
just_dna_enricher.enrich ¶
enrich — fill the source-independent resolution table for a module spec, then hand off to compile.
The resolver chain, first-hit-wins: an existing/human-authored row (authoritative, never clobbered) →
the local cache (a downloaded snapshot, offline) → live Ensembl (V2 GraphQL → V1 REST). It writes
resolution.csv beside the spec; the compiler then consumes it with no source knowledge and no
network. Two modes: best_effort fills what it can and records the rest as not_found; strict
fails unless every in-scope variant resolves to a position (the network analogue of the compiler's
strict=True). --offline clamps the chain to the cache alone (guaranteed zero egress).
Subject
dataclass
¶
Subject(
variant_key: str,
rsid: str | None,
chrom: str | None,
start: int | None,
ref: str | None,
alts: str | None,
constraint: str | None,
origin: str,
)
One row asking for a coordinate, normalized across the tables allowed to ask.
Resolution used to read variants.csv alone, so a PGx module — which by design carries no
variants.csv (one CSV = one concern) — enriched to an empty resolution.csv and shipped with no
coordinates at all. The chain itself was never variant-specific; only its input was. Normalizing
to a subject lets pharm_variants.csv and haplotypes.csv through the unchanged resolver,
caches, ordering and back-fill.
constraint is whatever the row knows about which alleles must be present at the locus, fed to
the shared hosting_verdict predicate so a one-to-many rsID drops the loci the row cannot be
about (and reports, rather than drops, the ones it cannot decide):
- a
VariantRow/PharmVariantRowsupplies its genotype (C/T); - a
HaplotypeRowsupplies its single defining allele (G) — the same membership question asked of one allele instead of two, which is why it reuses the same predicate rather than a parallel one; None(aPharmVariantRowwith no authored genotype) constrains nothing and keeps every locus.
EnrichmentError ¶
Bases: RuntimeError
Raised in strict mode when the chain cannot fully resolve the module.
SubjectDrift
dataclass
¶
One subject whose recorded facts and freshly derived facts disagree, under rederive=True.
The canary MODULE_LIFECYCLE describes, finally performed. Merge-not-clobber means an ordinary
re-run never re-asks about a recorded row, so a source that silently revised an answer moves no
fetched_at, no fact signature and no digest — and the only way to notice was to delete the table
and re-derive, which used to discard the curator's rows with it. With corrections living in the
overlay a full re-derivation costs nothing, and because the fresh table is staged beside the
current one both sides exist at the commit boundary, so the comparison is free.
before/after render the fact columns only (RESOLUTION_FACT_FIELDS), read off the
registry rather than restated here: the provenance columns move on every run by design, and a
report that called a new fetched_at a drift would cry wolf on every subject.
collect_subjects ¶
collect_subjects(
spec_dir: Path,
variants: list[VariantRow],
genome_build: str = "GRCh38",
) -> list[Subject]
Every row in the spec that needs a coordinate, variants.csv first, deduped by variant_key.
Order and precedence are load-bearing. variants.csv goes first so that when the same variant is
named by two tables, the SNP row wins — it is the only one carrying alts, which is a resolution
fact and therefore decides the compiled bytes. Letting a PGx row win would move an already
compiled module's artifact.digest, the same hazard the link ordering in the chain below exists
to avoid. Within that, first occurrence wins, so the emitted order is the authored order.
The PGx tables key without alts: a pharm annotation or a haplotype junction matches a
variant at chrom:start:ref regardless of allele. Mixing that up would mint a VRS allele id for a
row that never named an allele. heteroplasmy.csv is the exception and keys with alts,
because its own variant_key does — see the block below.
All three tables carry variant_key as a stamped field since 0.6 (RM43) — it was a property
on two of them and absent from HaplotypeRow entirely, which is why the block below derives one
inline. The stamped value is the same expression, frozen at load from the authored columns, so
nothing here changes; and PharmVariantRow.alts, added by the same item, is deliberately not read
here because it is compiler-filled data rather than an authored fact.
Source code in enricher/src/just_dna_enricher/enrich.py
192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 | |
select_par_representative ¶
select_par_representative(
loci: list[dict], *, build: str = "GRCh38"
) -> tuple[list[dict], list[dict]]
Split a one-to-many expansion into (kept, y_par_twins), preferring the X spelling.
A pseudoautosomal locus is one place on two contigs — PAR1 and PAR2 are the stretches X and Y share, so dbSNP maps one rsID to both and the expansion emits two rows for one finding. Probed 2026-08-04, every annotation source places PAR annotation on X and only the coordinate resolver disagrees: ClinVar has zero records in either PAR on Y (all 677 of its Y records lie outside the PARs), gnomAD v4 excludes the Y PAR from its callset outright (X PAR1 640000-641500 serves 880 variants, the same interval on Y serves none), and the ClinGen Allele Registry does mint a Y allele id but leaves the record a stub — a dbSNP cross-reference and nothing else, no ClinVar, no gnomAD, and a title that degrades to the bare genomic HGVS. Standard GRCh38 analysis sets then hard-mask the Y PAR, so the Y row cannot match a call either.
So keeping it records the sources' own convention, not the consumer's analysis set — which is
the distinction that makes this the enricher's decision to take (Principle 2: this is the only tier
permitted to hold a source convention). The caller keeps both with keep_par_twin.
This selects; it does not repair — the same contract as the allele-aware hosting_verdict
filter beside it. Every authored value is untouched, and the caller reports what was left out.
Two things it deliberately will not do:
- Fuse on geometry alone. A Y locus is dropped only when its
par_partnerX position is present and carries the sameref/alts. Partner coordinates say "same place"; they do not say "same variant", and a same-place different-allele pair is a real finding rather than a duplicate — so it is kept, and both rows survive. - Judge a gene. The verdict is per locus, because a real gene can straddle a PAR boundary: XG runs out of PAR1 (X:2,751,798-2,816,500 crosses the boundary at 2,781,479) and SPRY3 runs into PAR2 (X:155,612,298-155,782,459 crosses 155,701,383). Any module- or gene-scoped policy would be wrong for half of either one.
Source code in enricher/src/just_dna_enricher/enrich.py
spec_genome_build ¶
The build the module itself declares, which is the only build enrichment may resolve against.
enrich() took genome_build as a plain parameter defaulting to "GRCh38" and nothing ever
passed it — no CLI flag, no caller — so the GRCh38 gate on every link below was unreachable and a
genome_build: GRCh37 module was resolved against GRCh38 Ensembl, its GRCh38 coordinates written
into resolution.csv labelled GRCh38, and GRCh38 VRS ids minted for them. The compiler then
(correctly) refused to use any of it. The guard existed; the value never arrived — the same shape as
the VariantRow re-stamp bug, where a build parameter and its fall-through both existed and were
never reached.
A spec with no module_spec.yaml gets the format's own default. That is not a guess: enrich is
routinely pointed at a bare table directory in tests and by hand, and ModuleSpecConfig.genome_build
defaults to "GRCh38", so this returns what compiling that directory would assume. A spec whose
yaml is present but unreadable is a different case and raises — enrichment writes facts into that
directory, and picking a build for a module whose declaration cannot be read is exactly the
invention this function exists to remove.
Source code in enricher/src/just_dna_enricher/enrich.py
source_build_mismatch ¶
A warning when a drafting provider is about to write source_build coordinates into a module
that declares a different one — or None when the builds agree.
The gap this closes is that spec_genome_build had exactly one caller. It was written for the
bug where "the guard existed; the value never arrived", and then draft, draft-panel and
draft-clinpgx all shipped without asking it. Every source these providers read serves GRCh38 —
CPIC's allele_definitions, the ClinVar snapshot, the ClinPGx annotations — so drafting into a
genome_build: GRCh37 module writes 10,94942290 for rs1799853, whose GRCh37 position is
96702047, and nothing anywhere says a word. The compiler cannot catch it: a coordinate is legal
on either build, it is simply a different place. The whole point of reference_examples/grch37_build
is that this family of defect is silent by construction, and this is its ninth instance.
Reported, not repaired, and not refused. The provider still writes the row: refusing would make the three drafting commands unusable on a non-GRCh38 module rather than merely unhelpful, and stripping the coordinate to leave an rsid-only row is a different row than the author asked for — both are design decisions with real trade-offs, and the enricher's standing rule is that a pass which finds a disagreement says so and changes nothing. Which of the two the providers should eventually do is filed rather than decided here.
Returns None for the agreeing case, which is nearly every call — the house tri-state applies to
the answer, and "the builds agree" is not a finding.
Source code in enricher/src/just_dna_enricher/enrich.py
enrich ¶
enrich(
spec_dir: Path,
*,
mode: str = "best_effort",
offline: bool = False,
ensembl_cache: Path | None = None,
clinvar_cache: Path | None = None,
pubmind_cache: Path | None = None,
civic_cache: Path | None = None,
use_clinvar: bool = True,
use_gnomad: bool = True,
download: bool = True,
genome_build: str | None = None,
write: bool = True,
mint_vrs: bool = True,
verify_ref: bool = True,
verify_clinsig: bool = True,
verify_rsids: bool = True,
verify_datasets: bool = True,
keep_par_twin: bool = False,
rederive: bool = False,
keep_staging: bool = False,
progress: Callable[[int, int], None] | None = None,
resolver: EnsemblResolver | None = None,
gnomad_client: Optional[GnomadClient] = None,
grch37_client: Grch37Client | None = None,
release_probes: Mapping[str, ReleaseProbe]
| None = None,
) -> EnrichmentResult
Resolve a spec's variants into resolution.csv. See the module docstring for the chain/modes.
The chain is: existing rows → Ensembl cache → ClinVar cache (use_clinvar, stamps
source="clinvar") → live Ensembl → live gnomAD (use_gnomad, stamps source="gnomad"). Each
later link fills only what the earlier ones missed, so whichever link first knows a variant
decides its alts — and since alts is a fact column, that decides the compiled bytes. The
ordering is therefore chosen so no already-compiled module's artifact.digest can move when a new
link is added. --offline clamps the chain to the two local caches (zero egress).
mint_vrs stamps a ga4gh:VA.… allele id onto every resolved row (see vrs.mint_resolution_rows).
Substitutions mint offline with no dependency; indels need the sequence, so they mint only when the
run is online.
verify_ref checks each authored/resolved ref against the actual reference sequence and reports
disagreements — enrichment is partly validation of authored data, and this tier is the only one
that can perform it (see sequences.verify_reference_alleles). It never repairs, and severity
follows the mode, mirroring the compiler's VRS verify pass: strict treats a mismatch as fatal
(its contract is a reproducible artifact, and a wrong ref can silently mint a different allele
id), while best_effort warns and carries on. Needs sequence access, so it is skipped offline.
A mismatch is then asked one further question, over those rows only: does the coordinate read
as GRCh37 rather than as a wrong cell? That needs the live GRCh37 service, so it is skipped
offline, it is bounded (grch37.DEFAULT_DIAGNOSIS_LIMIT — a systematic wrong build answers the
same on every row), and grch37_client injects the client the way resolver/gnomad_client do.
It adds no flag of its own: verify_ref gates the family and --offline is the egress switch.
verify_clinsig compares each authored clin_sig against the ClinVar snapshot's own. It is
offline-capable (the snapshot is local) and is the one check whose severity does not follow the
mode: it warns in strict too, because failing a compile would make the format arbitrate a
clinical disagreement. See clinical.verify_clin_sig for the full argument.
verify_datasets compares every release this module records having been drafted from
(SourceRow.dataset) against the one that source publishes now, and reports the gap (RM85). It is
rederive's cheap neighbour: both ask has the world moved, one about the rows and one about the
release label, so this is the question you put first — it costs one request per source and
tells you whether the full re-derivation is worth running. It reads and never writes; repairing a
stale label is a re-draft, which is an author's decision and a different command. Severity follows
the mode. --offline makes it unchecked, never up to date, and an unreachable source is never
escalated by strict — nothing an author can edit clears a failed request. release_probes
injects the registry the way resolver/gnomad_client inject theirs.
keep_par_twin keeps both spellings of a pseudoautosomal locus. By default only the X one is
recorded, because every annotation source uses X and a standard GRCh38 analysis set hard-masks the
Y PAR — see select_par_representative for the probe behind that. Set it for a consumer whose
reference is unmasked. The switch belongs here and could not live on the compiler: resolution.csv
is injected data that travels with the module, so the choice is recorded and
compile → reverse → compile stays a fixed point either way, whereas a compiler flag would not
survive reverse_module rebuilding the spec from parquet alone (Principle 7).
The run is a transaction. What the sources answer is staged to disk as the run proceeds, in the
target's own directory (transaction.ResolutionJournal), and the table itself is written once — at
the gate, through a writer that renames into place. Two properties follow, and the second is the
one the item was really about. A kill at minute 29 leaves the staged answers, so the next run
resumes instead of paying the thirty minutes again. And a refused strict run commits nothing:
every refusal below raises before the write block, so the module is left exactly as it was — a
written promise rather than an accident of statement order. keep_staging leaves the staged
answers behind after a successful commit, for debugging; the default removes them.
The whole run is one read-modify-write window over resolution.csv, so it is held under an
advisory flock on the spec directory (transaction.spec_lock) — two concurrent runs were
last-writer-wins over a merge with neither able to see the other, and a shorter table is
indistinguishable from a module whose author resolved less. A second run refuses rather than
waiting; a platform or filesystem that will not take the lock degrades with a warning that says so.
write=False takes no lock and stages nothing: with nothing written there is no window to exclude,
which keeps the flag meaning one thing everywhere.
rederive re-asks the sources about every subject, including the ones already recorded, and
reports which of them changed value (EnrichmentResult.rederived). That is the drift canary: an
ordinary run gap-fills and never re-asks, so a source that silently revised an answer moves
nothing an author could notice. A recorded subject this run could not ask about keeps its recorded
rows — re-deriving must never be a way to shorten the table, which is the very incident this
transaction exists for.
progress is called (done, total) over subjects, total known before the first call. It
exists for a caller with an idle timeout: what such a caller needs is a keepalive with monotonic
progress, which a phase counter cannot give (a twenty-nine-minute phase emits nothing) and a link
counter cannot promise (its total is not known until resolution finds the links). Exceptions from
the callback are not swallowed; under the transaction that costs nothing, since the staged answers
survive an abort exactly as they survive a kill.
Source code in enricher/src/just_dna_enricher/enrich.py
748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 | |