Skip to content

just_dna_enricher.enrich

just_dna_enricher.enrich

enrich — fill the source-independent resolution table for a module spec, then hand off to compile.

The resolver chain, first-hit-wins: an existing/human-authored row (authoritative, never clobbered) → the local cache (a downloaded snapshot, offline) → live Ensembl (V2 GraphQL → V1 REST). It writes resolution.csv beside the spec; the compiler then consumes it with no source knowledge and no network. Two modes: best_effort fills what it can and records the rest as not_found; strict fails unless every in-scope variant resolves to a position (the network analogue of the compiler's strict=True). --offline clamps the chain to the cache alone (guaranteed zero egress).

Subject dataclass

Subject(
    variant_key: str,
    rsid: str | None,
    chrom: str | None,
    start: int | None,
    ref: str | None,
    alts: str | None,
    constraint: str | None,
    origin: str,
)

One row asking for a coordinate, normalized across the tables allowed to ask.

Resolution used to read variants.csv alone, so a PGx module — which by design carries no variants.csv (one CSV = one concern) — enriched to an empty resolution.csv and shipped with no coordinates at all. The chain itself was never variant-specific; only its input was. Normalizing to a subject lets pharm_variants.csv and haplotypes.csv through the unchanged resolver, caches, ordering and back-fill.

constraint is whatever the row knows about which alleles must be present at the locus, fed to the shared hosting_verdict predicate so a one-to-many rsID drops the loci the row cannot be about (and reports, rather than drops, the ones it cannot decide):

  • a VariantRow/PharmVariantRow supplies its genotype (C/T);
  • a HaplotypeRow supplies its single defining allele (G) — the same membership question asked of one allele instead of two, which is why it reuses the same predicate rather than a parallel one;
  • None (a PharmVariantRow with no authored genotype) constrains nothing and keeps every locus.

EnrichmentError

Bases: RuntimeError

Raised in strict mode when the chain cannot fully resolve the module.

SubjectDrift dataclass

SubjectDrift(variant_key: str, before: str, after: str)

One subject whose recorded facts and freshly derived facts disagree, under rederive=True.

The canary MODULE_LIFECYCLE describes, finally performed. Merge-not-clobber means an ordinary re-run never re-asks about a recorded row, so a source that silently revised an answer moves no fetched_at, no fact signature and no digest — and the only way to notice was to delete the table and re-derive, which used to discard the curator's rows with it. With corrections living in the overlay a full re-derivation costs nothing, and because the fresh table is staged beside the current one both sides exist at the commit boundary, so the comparison is free.

before/after render the fact columns only (RESOLUTION_FACT_FIELDS), read off the registry rather than restated here: the provenance columns move on every run by design, and a report that called a new fetched_at a drift would cry wolf on every subject.

collect_subjects

collect_subjects(
    spec_dir: Path,
    variants: list[VariantRow],
    genome_build: str = "GRCh38",
) -> list[Subject]

Every row in the spec that needs a coordinate, variants.csv first, deduped by variant_key.

Order and precedence are load-bearing. variants.csv goes first so that when the same variant is named by two tables, the SNP row wins — it is the only one carrying alts, which is a resolution fact and therefore decides the compiled bytes. Letting a PGx row win would move an already compiled module's artifact.digest, the same hazard the link ordering in the chain below exists to avoid. Within that, first occurrence wins, so the emitted order is the authored order.

The PGx tables key without alts: a pharm annotation or a haplotype junction matches a variant at chrom:start:ref regardless of allele. Mixing that up would mint a VRS allele id for a row that never named an allele. heteroplasmy.csv is the exception and keys with alts, because its own variant_key does — see the block below.

All three tables carry variant_key as a stamped field since 0.6 (RM43) — it was a property on two of them and absent from HaplotypeRow entirely, which is why the block below derives one inline. The stamped value is the same expression, frozen at load from the authored columns, so nothing here changes; and PharmVariantRow.alts, added by the same item, is deliberately not read here because it is compiler-filled data rather than an authored fact.

Source code in enricher/src/just_dna_enricher/enrich.py
def collect_subjects(
    spec_dir: Path, variants: list[VariantRow], genome_build: str = "GRCh38"
) -> list[Subject]:
    """Every row in the spec that needs a coordinate, `variants.csv` first, deduped by `variant_key`.

    Order and precedence are load-bearing. `variants.csv` goes first so that when the same variant is
    named by two tables, the SNP row wins — it is the only one carrying `alts`, which is a resolution
    *fact* and therefore decides the compiled bytes. Letting a PGx row win would move an already
    compiled module's `artifact.digest`, the same hazard the link ordering in the chain below exists
    to avoid. Within that, first occurrence wins, so the emitted order is the authored order.

    The PGx tables key **without** `alts`: a pharm annotation or a haplotype junction matches a
    variant at `chrom:start:ref` regardless of allele. Mixing that up would mint a VRS allele id for a
    row that never named an allele. `heteroplasmy.csv` is the exception and keys *with* `alts`,
    because its own `variant_key` does — see the block below.

    All three tables carry `variant_key` as a **stamped field** since 0.6 (RM43) — it was a property
    on two of them and absent from `HaplotypeRow` entirely, which is why the block below derives one
    inline. The stamped value is the same expression, frozen at load from the authored columns, so
    nothing here changes; and `PharmVariantRow.alts`, added by the same item, is deliberately not read
    here because it is compiler-filled data rather than an authored fact.
    """
    subjects: list[Subject] = [_subject_of_variant(v, genome_build) for v in variants]

    pharm_path = spec_dir / "pharm_variants.csv"
    if pharm_path.exists():
        rows, errors, _ = load_csv_rows(pharm_path, PharmVariantRow, "pharm_variants.csv")
        if errors:
            raise EnrichmentError(f"pharm_variants.csv is invalid: {errors[0]}")
        subjects.extend(
            Subject(r.variant_key, r.rsid, r.chrom, r.start, r.ref, None, r.genotype, "pharm_variants.csv")
            for r in rows
        )

    hap_path = spec_dir / "haplotypes.csv"
    if hap_path.exists():
        rows, errors, _ = load_csv_rows(hap_path, HaplotypeRow, "haplotypes.csv")
        if errors:
            raise EnrichmentError(f"haplotypes.csv is invalid: {errors[0]}")
        subjects.extend(
            Subject(
                derive_variant_key(r.rsid, r.chrom, r.start, r.ref),
                r.rsid,
                r.chrom,
                r.start,
                r.ref,
                None,
                r.allele,
                "haplotypes.csv",
            )
            for r in rows
        )

    # `heteroplasmy.csv` is the third table that can ask, and it was left out of the 0.5 round for no
    # reason anyone recorded: its coordinates are optional exactly like the PGx ones, so an
    # rsid-authored heteroplasmy module resolved to nothing and the compiler's positional-joinability
    # warning would have named a gap no tool could close.
    #
    # Two differences from the blocks above, both load-bearing. It **passes `alts`**, because
    # `HeteroplasmyRow.variant_key` mints one the same way `VariantRow` does (verified equal for both
    # the rsid and the coordinate shape), and a subject whose key carries an allele must carry the
    # allele. That makes the key **build-dependent**, so the load takes `genome_build` — the RM36
    # trap, and the reason the two blocks above rightly do not. Its `constraint` is `None`: a
    # heteroplasmy row is a measurement band over a locus, not a claim about a genotype, so it
    # constrains no locus out of a one-to-many expansion.
    het_path = spec_dir / "heteroplasmy.csv"
    if het_path.exists():
        rows, errors, _ = load_csv_rows(
            het_path, HeteroplasmyRow, "heteroplasmy.csv", genome_build=genome_build
        )
        if errors:
            raise EnrichmentError(f"heteroplasmy.csv is invalid: {errors[0]}")
        subjects.extend(
            Subject(
                r.variant_key
                or derive_variant_key(r.rsid, r.chrom, r.start, r.ref, r.alts, build=genome_build),
                r.rsid,
                r.chrom,
                r.start,
                r.ref,
                r.alts,
                None,
                "heteroplasmy.csv",
            )
            for r in rows
        )

    deduped: dict[str, Subject] = {}
    for s in subjects:
        deduped.setdefault(s.variant_key, s)
    return list(deduped.values())

select_par_representative

select_par_representative(
    loci: list[dict], *, build: str = "GRCh38"
) -> tuple[list[dict], list[dict]]

Split a one-to-many expansion into (kept, y_par_twins), preferring the X spelling.

A pseudoautosomal locus is one place on two contigs — PAR1 and PAR2 are the stretches X and Y share, so dbSNP maps one rsID to both and the expansion emits two rows for one finding. Probed 2026-08-04, every annotation source places PAR annotation on X and only the coordinate resolver disagrees: ClinVar has zero records in either PAR on Y (all 677 of its Y records lie outside the PARs), gnomAD v4 excludes the Y PAR from its callset outright (X PAR1 640000-641500 serves 880 variants, the same interval on Y serves none), and the ClinGen Allele Registry does mint a Y allele id but leaves the record a stub — a dbSNP cross-reference and nothing else, no ClinVar, no gnomAD, and a title that degrades to the bare genomic HGVS. Standard GRCh38 analysis sets then hard-mask the Y PAR, so the Y row cannot match a call either.

So keeping it records the sources' own convention, not the consumer's analysis set — which is the distinction that makes this the enricher's decision to take (Principle 2: this is the only tier permitted to hold a source convention). The caller keeps both with keep_par_twin.

This selects; it does not repair — the same contract as the allele-aware hosting_verdict filter beside it. Every authored value is untouched, and the caller reports what was left out.

Two things it deliberately will not do:

  • Fuse on geometry alone. A Y locus is dropped only when its par_partner X position is present and carries the same ref/alts. Partner coordinates say "same place"; they do not say "same variant", and a same-place different-allele pair is a real finding rather than a duplicate — so it is kept, and both rows survive.
  • Judge a gene. The verdict is per locus, because a real gene can straddle a PAR boundary: XG runs out of PAR1 (X:2,751,798-2,816,500 crosses the boundary at 2,781,479) and SPRY3 runs into PAR2 (X:155,612,298-155,782,459 crosses 155,701,383). Any module- or gene-scoped policy would be wrong for half of either one.
Source code in enricher/src/just_dna_enricher/enrich.py
def select_par_representative(loci: list[dict], *, build: str = "GRCh38") -> tuple[list[dict], list[dict]]:
    """Split a one-to-many expansion into `(kept, y_par_twins)`, preferring the X spelling.

    A pseudoautosomal locus is **one place on two contigs** — PAR1 and PAR2 are the stretches X and Y
    share, so dbSNP maps one rsID to both and the expansion emits two rows for one finding. Probed
    2026-08-04, every annotation source places PAR annotation on **X** and only the coordinate resolver
    disagrees: ClinVar has zero records in either PAR on Y (all 677 of its Y records lie outside the
    PARs), gnomAD v4 excludes the Y PAR from its callset outright (X PAR1 640000-641500 serves 880
    variants, the same interval on Y serves none), and the ClinGen Allele Registry does mint a Y allele
    id but leaves the record a stub — a dbSNP cross-reference and nothing else, no ClinVar, no gnomAD,
    and a title that degrades to the bare genomic HGVS. Standard GRCh38 analysis sets then hard-mask the
    Y PAR, so the Y row cannot match a call either.

    So keeping it records **the sources' own convention**, not the consumer's analysis set — which is
    the distinction that makes this the enricher's decision to take (Principle 2: this is the only tier
    permitted to hold a source convention). The caller keeps both with `keep_par_twin`.

    This **selects; it does not repair** — the same contract as the allele-aware `hosting_verdict`
    filter beside it. Every authored value is untouched, and the caller reports what was left out.

    Two things it deliberately will not do:

    * **Fuse on geometry alone.** A Y locus is dropped only when its `par_partner` X position is present
      *and* carries the same `ref`/`alts`. Partner coordinates say "same place"; they do not say "same
      variant", and a same-place different-allele pair is a real finding rather than a duplicate — so it
      is kept, and both rows survive.
    * **Judge a gene.** The verdict is per locus, because a real gene can straddle a PAR boundary: **XG**
      runs out of PAR1 (X:2,751,798-2,816,500 crosses the boundary at 2,781,479) and **SPRY3** runs into
      PAR2 (X:155,612,298-155,782,459 crosses 155,701,383). Any module- or gene-scoped policy would be
      wrong for half of either one.
    """
    by_place: dict[tuple[str, int, str, frozenset[str]], dict] = {}
    for locus in loci:
        chrom, start = normalize_chrom(locus.get("chrom")), locus.get("start")
        if chrom is None or start is None:
            continue
        ref, alts = _locus_alleles(locus)
        by_place[(chrom, int(start), ref, alts)] = locus

    kept: list[dict] = []
    twins: list[dict] = []
    for locus in loci:
        chrom, start = normalize_chrom(locus.get("chrom")), locus.get("start")
        partner = par_partner(chrom, int(start), build=build) if chrom == "Y" and start is not None else None
        ref, alts = _locus_alleles(locus)
        if partner is not None and (partner[0], partner[1], ref, alts) in by_place:
            twins.append(locus)
        else:
            kept.append(locus)
    return kept, twins

spec_genome_build

spec_genome_build(spec_dir: Path) -> str

The build the module itself declares, which is the only build enrichment may resolve against.

enrich() took genome_build as a plain parameter defaulting to "GRCh38" and nothing ever passed it — no CLI flag, no caller — so the GRCh38 gate on every link below was unreachable and a genome_build: GRCh37 module was resolved against GRCh38 Ensembl, its GRCh38 coordinates written into resolution.csv labelled GRCh38, and GRCh38 VRS ids minted for them. The compiler then (correctly) refused to use any of it. The guard existed; the value never arrived — the same shape as the VariantRow re-stamp bug, where a build parameter and its fall-through both existed and were never reached.

A spec with no module_spec.yaml gets the format's own default. That is not a guess: enrich is routinely pointed at a bare table directory in tests and by hand, and ModuleSpecConfig.genome_build defaults to "GRCh38", so this returns what compiling that directory would assume. A spec whose yaml is present but unreadable is a different case and raises — enrichment writes facts into that directory, and picking a build for a module whose declaration cannot be read is exactly the invention this function exists to remove.

Source code in enricher/src/just_dna_enricher/enrich.py
def spec_genome_build(spec_dir: Path) -> str:
    """The build the module itself declares, which is the only build enrichment may resolve against.

    `enrich()` took `genome_build` as a plain parameter defaulting to `"GRCh38"` and **nothing ever
    passed it** — no CLI flag, no caller — so the GRCh38 gate on every link below was unreachable and a
    `genome_build: GRCh37` module was resolved against GRCh38 Ensembl, its GRCh38 coordinates written
    into `resolution.csv` labelled `GRCh38`, and GRCh38 VRS ids minted for them. The compiler then
    (correctly) refused to use any of it. The guard existed; the value never arrived — the same shape as
    the `VariantRow` re-stamp bug, where a `build` parameter and its fall-through both existed and were
    never reached.

    A spec with no `module_spec.yaml` gets the format's own default. That is not a guess: `enrich` is
    routinely pointed at a bare table directory in tests and by hand, and `ModuleSpecConfig.genome_build`
    defaults to `"GRCh38"`, so this returns what compiling that directory would assume. A spec whose
    yaml is *present but unreadable* is a different case and raises — enrichment writes facts into that
    directory, and picking a build for a module whose declaration cannot be read is exactly the
    invention this function exists to remove.
    """
    path = spec_dir / "module_spec.yaml"
    if not path.exists():
        return "GRCh38"
    # Through the public loader since S74 — this call site is what made the case that one was needed,
    # since the workspace's own network tier was reaching into `_load_yaml`. It already raised on the
    # `None` half of the tuple, which is `load_spec`'s contract, so the translation is the only thing
    # left here: a pass owes its caller its own exception type.
    try:
        return load_spec(path).genome_build
    except SpecError as exc:
        # **Guard at the answerer (S103).** The question is the build, and a scaffold's unfilled
        # `title:` says nothing about it — yet `load_spec` validates the whole file, so `scaffold`
        # followed by `draft`, the reference README's own recipe, refused on `<<REPLACE>>` in three
        # fields the draft never reads. Only that one defect is looked past, and only outside the
        # build cell: a typo'd key, a wrong type, or a placeholder *in* `genome_build` still refuses,
        # because reading the default past those would reopen the `genome_bild:` hole the model's
        # `extra="forbid"` closed.
        declared, residual = _declared_build_behind_placeholders(path)
        if declared is not None:
            return declared
        # No `pass genome_build=` remedy in this sentence: that parameter belongs to `enrich()` alone,
        # and the drafters that share this reader have no such flag to offer. `residual` is the
        # diagnosis that remains once the stubs are looked past — the placeholder guard runs first
        # and would otherwise hide a `genome_bild:` behind "fill in the title".
        raise EnrichmentError(
            f"cannot read the module's genome_build: {residual or exc}. Enrichment resolves against "
            f"one assembly and records the answer under the module's declared build, so it will not "
            f"choose one for you — fix module_spec.yaml."
        ) from exc

source_build_mismatch

source_build_mismatch(
    spec_dir: Path,
    source: str,
    source_build: str = "GRCh38",
) -> str | None

A warning when a drafting provider is about to write source_build coordinates into a module that declares a different one — or None when the builds agree.

The gap this closes is that spec_genome_build had exactly one caller. It was written for the bug where "the guard existed; the value never arrived", and then draft, draft-panel and draft-clinpgx all shipped without asking it. Every source these providers read serves GRCh38 — CPIC's allele_definitions, the ClinVar snapshot, the ClinPGx annotations — so drafting into a genome_build: GRCh37 module writes 10,94942290 for rs1799853, whose GRCh37 position is 96702047, and nothing anywhere says a word. The compiler cannot catch it: a coordinate is legal on either build, it is simply a different place. The whole point of reference_examples/grch37_build is that this family of defect is silent by construction, and this is its ninth instance.

Reported, not repaired, and not refused. The provider still writes the row: refusing would make the three drafting commands unusable on a non-GRCh38 module rather than merely unhelpful, and stripping the coordinate to leave an rsid-only row is a different row than the author asked for — both are design decisions with real trade-offs, and the enricher's standing rule is that a pass which finds a disagreement says so and changes nothing. Which of the two the providers should eventually do is filed rather than decided here.

Returns None for the agreeing case, which is nearly every call — the house tri-state applies to the answer, and "the builds agree" is not a finding.

Source code in enricher/src/just_dna_enricher/enrich.py
def source_build_mismatch(spec_dir: Path, source: str, source_build: str = "GRCh38") -> str | None:
    """A warning when a drafting provider is about to write `source_build` coordinates into a module
    that declares a different one — or `None` when the builds agree.

    **The gap this closes is that `spec_genome_build` had exactly one caller.** It was written for the
    bug where "the guard existed; the value never arrived", and then `draft`, `draft-panel` and
    `draft-clinpgx` all shipped without asking it. Every source these providers read serves GRCh38 —
    CPIC's `allele_definitions`, the ClinVar snapshot, the ClinPGx annotations — so drafting into a
    `genome_build: GRCh37` module writes `10,94942290` for `rs1799853`, whose GRCh37 position is
    `96702047`, and nothing anywhere says a word. The compiler cannot catch it: a coordinate is legal
    on either build, it is simply a different place. The whole point of `reference_examples/grch37_build`
    is that this family of defect is *silent by construction*, and this is its ninth instance.

    **Reported, not repaired, and not refused.** The provider still writes the row: refusing would
    make the three drafting commands unusable on a non-GRCh38 module rather than merely
    unhelpful, and stripping the coordinate to leave an rsid-only row is a *different* row than the
    author asked for — both are design decisions with real trade-offs, and the enricher's standing
    rule is that a pass which finds a disagreement says so and changes nothing. Which of the two the
    providers should eventually do is filed rather than decided here.

    Returns `None` for the agreeing case, which is nearly every call — the house tri-state applies to
    the *answer*, and "the builds agree" is not a finding.
    """
    declared = spec_genome_build(spec_dir)
    if declared == source_build:
        return None
    return (
        f"{source} publishes {source_build} coordinates and this module declares "
        f"genome_build={declared!r}: every chrom/start drafted below names a {source_build} position, "
        f"recorded as though it were {declared}. Nothing downstream can detect this — a coordinate is "
        f"valid on either assembly, it is simply a different base — so either re-declare the module as "
        f"{source_build}, or delete the drafted coordinates and keep the rsIDs, which name a variant "
        f"without naming an assembly (`just-dna-enricher hint recover` converts in that direction)."
    )

enrich

enrich(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    offline: bool = False,
    ensembl_cache: Path | None = None,
    clinvar_cache: Path | None = None,
    pubmind_cache: Path | None = None,
    civic_cache: Path | None = None,
    use_clinvar: bool = True,
    use_gnomad: bool = True,
    download: bool = True,
    genome_build: str | None = None,
    write: bool = True,
    mint_vrs: bool = True,
    verify_ref: bool = True,
    verify_clinsig: bool = True,
    verify_rsids: bool = True,
    verify_datasets: bool = True,
    keep_par_twin: bool = False,
    rederive: bool = False,
    keep_staging: bool = False,
    progress: Callable[[int, int], None] | None = None,
    resolver: EnsemblResolver | None = None,
    gnomad_client: Optional[GnomadClient] = None,
    grch37_client: Grch37Client | None = None,
    release_probes: Mapping[str, ReleaseProbe]
    | None = None,
) -> EnrichmentResult

Resolve a spec's variants into resolution.csv. See the module docstring for the chain/modes.

The chain is: existing rows → Ensembl cache → ClinVar cache (use_clinvar, stamps source="clinvar") → live Ensembl → live gnomAD (use_gnomad, stamps source="gnomad"). Each later link fills only what the earlier ones missed, so whichever link first knows a variant decides its alts — and since alts is a fact column, that decides the compiled bytes. The ordering is therefore chosen so no already-compiled module's artifact.digest can move when a new link is added. --offline clamps the chain to the two local caches (zero egress).

mint_vrs stamps a ga4gh:VA.… allele id onto every resolved row (see vrs.mint_resolution_rows). Substitutions mint offline with no dependency; indels need the sequence, so they mint only when the run is online.

verify_ref checks each authored/resolved ref against the actual reference sequence and reports disagreements — enrichment is partly validation of authored data, and this tier is the only one that can perform it (see sequences.verify_reference_alleles). It never repairs, and severity follows the mode, mirroring the compiler's VRS verify pass: strict treats a mismatch as fatal (its contract is a reproducible artifact, and a wrong ref can silently mint a different allele id), while best_effort warns and carries on. Needs sequence access, so it is skipped offline.

A mismatch is then asked one further question, over those rows only: does the coordinate read as GRCh37 rather than as a wrong cell? That needs the live GRCh37 service, so it is skipped offline, it is bounded (grch37.DEFAULT_DIAGNOSIS_LIMIT — a systematic wrong build answers the same on every row), and grch37_client injects the client the way resolver/gnomad_client do. It adds no flag of its own: verify_ref gates the family and --offline is the egress switch.

verify_clinsig compares each authored clin_sig against the ClinVar snapshot's own. It is offline-capable (the snapshot is local) and is the one check whose severity does not follow the mode: it warns in strict too, because failing a compile would make the format arbitrate a clinical disagreement. See clinical.verify_clin_sig for the full argument.

verify_datasets compares every release this module records having been drafted from (SourceRow.dataset) against the one that source publishes now, and reports the gap (RM85). It is rederive's cheap neighbour: both ask has the world moved, one about the rows and one about the release label, so this is the question you put first — it costs one request per source and tells you whether the full re-derivation is worth running. It reads and never writes; repairing a stale label is a re-draft, which is an author's decision and a different command. Severity follows the mode. --offline makes it unchecked, never up to date, and an unreachable source is never escalated by strict — nothing an author can edit clears a failed request. release_probes injects the registry the way resolver/gnomad_client inject theirs.

keep_par_twin keeps both spellings of a pseudoautosomal locus. By default only the X one is recorded, because every annotation source uses X and a standard GRCh38 analysis set hard-masks the Y PAR — see select_par_representative for the probe behind that. Set it for a consumer whose reference is unmasked. The switch belongs here and could not live on the compiler: resolution.csv is injected data that travels with the module, so the choice is recorded and compile → reverse → compile stays a fixed point either way, whereas a compiler flag would not survive reverse_module rebuilding the spec from parquet alone (Principle 7).

The run is a transaction. What the sources answer is staged to disk as the run proceeds, in the target's own directory (transaction.ResolutionJournal), and the table itself is written once — at the gate, through a writer that renames into place. Two properties follow, and the second is the one the item was really about. A kill at minute 29 leaves the staged answers, so the next run resumes instead of paying the thirty minutes again. And a refused strict run commits nothing: every refusal below raises before the write block, so the module is left exactly as it was — a written promise rather than an accident of statement order. keep_staging leaves the staged answers behind after a successful commit, for debugging; the default removes them.

The whole run is one read-modify-write window over resolution.csv, so it is held under an advisory flock on the spec directory (transaction.spec_lock) — two concurrent runs were last-writer-wins over a merge with neither able to see the other, and a shorter table is indistinguishable from a module whose author resolved less. A second run refuses rather than waiting; a platform or filesystem that will not take the lock degrades with a warning that says so. write=False takes no lock and stages nothing: with nothing written there is no window to exclude, which keeps the flag meaning one thing everywhere.

rederive re-asks the sources about every subject, including the ones already recorded, and reports which of them changed value (EnrichmentResult.rederived). That is the drift canary: an ordinary run gap-fills and never re-asks, so a source that silently revised an answer moves nothing an author could notice. A recorded subject this run could not ask about keeps its recorded rows — re-deriving must never be a way to shorten the table, which is the very incident this transaction exists for.

progress is called (done, total) over subjects, total known before the first call. It exists for a caller with an idle timeout: what such a caller needs is a keepalive with monotonic progress, which a phase counter cannot give (a twenty-nine-minute phase emits nothing) and a link counter cannot promise (its total is not known until resolution finds the links). Exceptions from the callback are not swallowed; under the transaction that costs nothing, since the staged answers survive an abort exactly as they survive a kill.

Source code in enricher/src/just_dna_enricher/enrich.py
def enrich(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    offline: bool = False,
    ensembl_cache: Path | None = None,
    clinvar_cache: Path | None = None,
    pubmind_cache: Path | None = None,
    civic_cache: Path | None = None,
    use_clinvar: bool = True,
    use_gnomad: bool = True,
    download: bool = True,
    genome_build: str | None = None,
    write: bool = True,
    mint_vrs: bool = True,
    verify_ref: bool = True,
    verify_clinsig: bool = True,
    verify_rsids: bool = True,
    verify_datasets: bool = True,
    keep_par_twin: bool = False,
    rederive: bool = False,
    keep_staging: bool = False,
    progress: Callable[[int, int], None] | None = None,
    resolver: EnsemblResolver | None = None,
    gnomad_client: Optional["GnomadClient"] = None,
    grch37_client: Grch37Client | None = None,
    release_probes: Mapping[str, ReleaseProbe] | None = None,
) -> EnrichmentResult:
    """Resolve a spec's variants into `resolution.csv`. See the module docstring for the chain/modes.

    The chain is: existing rows → Ensembl cache → ClinVar cache (`use_clinvar`, stamps
    `source="clinvar"`) → live Ensembl → live gnomAD (`use_gnomad`, stamps `source="gnomad"`). Each
    later link fills only what the earlier ones missed, so whichever link *first* knows a variant
    decides its `alts` — and since `alts` is a fact column, that decides the compiled bytes. The
    ordering is therefore chosen so no already-compiled module's `artifact.digest` can move when a new
    link is added. `--offline` clamps the chain to the two local caches (zero egress).

    `mint_vrs` stamps a `ga4gh:VA.…` allele id onto every resolved row (see `vrs.mint_resolution_rows`).
    Substitutions mint offline with no dependency; indels need the sequence, so they mint only when the
    run is online.

    `verify_ref` checks each authored/resolved `ref` against the actual reference sequence and reports
    disagreements — enrichment is partly *validation* of authored data, and this tier is the only one
    that can perform it (see `sequences.verify_reference_alleles`). It never repairs, and severity
    follows the mode, mirroring the compiler's VRS verify pass: `strict` treats a mismatch as fatal
    (its contract is a reproducible artifact, and a wrong `ref` can silently mint a *different* allele
    id), while `best_effort` warns and carries on. Needs sequence access, so it is skipped offline.

    A mismatch is then asked one further question, over those rows **only**: does the coordinate read
    as GRCh37 rather than as a wrong cell? That needs the live GRCh37 service, so it is skipped
    offline, it is bounded (`grch37.DEFAULT_DIAGNOSIS_LIMIT` — a systematic wrong build answers the
    same on every row), and `grch37_client` injects the client the way `resolver`/`gnomad_client` do.
    It adds no flag of its own: `verify_ref` gates the family and `--offline` is the egress switch.

    `verify_clinsig` compares each authored `clin_sig` against the ClinVar snapshot's own. It is
    offline-capable (the snapshot is local) and is the **one check whose severity does not follow the
    mode**: it warns in `strict` too, because failing a compile would make the format arbitrate a
    clinical disagreement. See `clinical.verify_clin_sig` for the full argument.

    `verify_datasets` compares every release this module records having been drafted from
    (`SourceRow.dataset`) against the one that source publishes now, and reports the gap (RM85). It is
    `rederive`'s cheap neighbour: both ask *has the world moved*, one about the rows and one about the
    release **label**, so this is the question you put first — it costs one request per source and
    tells you whether the full re-derivation is worth running. It reads and never writes; repairing a
    stale label is a re-draft, which is an author's decision and a different command. Severity follows
    the mode. `--offline` makes it `unchecked`, never *up to date*, and an unreachable source is never
    escalated by `strict` — nothing an author can edit clears a failed request. `release_probes`
    injects the registry the way `resolver`/`gnomad_client` inject theirs.

    `keep_par_twin` keeps both spellings of a pseudoautosomal locus. By default only the X one is
    recorded, because every annotation source uses X and a standard GRCh38 analysis set hard-masks the
    Y PAR — see `select_par_representative` for the probe behind that. Set it for a consumer whose
    reference is unmasked. The switch belongs here and could not live on the compiler: `resolution.csv`
    is injected data that travels with the module, so the choice is *recorded* and
    `compile → reverse → compile` stays a fixed point either way, whereas a compiler flag would not
    survive `reverse_module` rebuilding the spec from parquet alone (Principle 7).

    **The run is a transaction.** What the sources answer is staged to disk as the run proceeds, in the
    target's own directory (`transaction.ResolutionJournal`), and the table itself is written once — at
    the gate, through a writer that renames into place. Two properties follow, and the second is the
    one the item was really about. A kill at minute 29 leaves the staged answers, so the next run
    resumes instead of paying the thirty minutes again. And **a refused `strict` run commits nothing**:
    every refusal below raises before the write block, so the module is left exactly as it was — a
    written promise rather than an accident of statement order. `keep_staging` leaves the staged
    answers behind after a successful commit, for debugging; the default removes them.

    The whole run is one read-modify-write window over `resolution.csv`, so it is held under an
    advisory `flock` on the spec directory (`transaction.spec_lock`) — two concurrent runs were
    last-writer-wins over a merge with neither able to see the other, and a shorter table is
    indistinguishable from a module whose author resolved less. A second run refuses rather than
    waiting; a platform or filesystem that will not take the lock degrades with a warning that says so.
    `write=False` takes no lock and stages nothing: with nothing written there is no window to exclude,
    which keeps the flag meaning one thing everywhere.

    `rederive` re-asks the sources about **every** subject, including the ones already recorded, and
    reports which of them changed value (`EnrichmentResult.rederived`). That is the drift canary: an
    ordinary run gap-fills and never re-asks, so a source that silently revised an answer moves
    nothing an author could notice. A recorded subject this run could not ask about keeps its recorded
    rows — re-deriving must never be a way to shorten the table, which is the very incident this
    transaction exists for.

    `progress` is called `(done, total)` over **subjects**, `total` known before the first call. It
    exists for a caller with an idle timeout: what such a caller needs is a keepalive with monotonic
    progress, which a phase counter cannot give (a twenty-nine-minute phase emits nothing) and a link
    counter cannot promise (its total is not known until resolution finds the links). Exceptions from
    the callback are not swallowed; under the transaction that costs nothing, since the staged answers
    survive an abort exactly as they survive a kill.
    """
    spec_dir = Path(spec_dir)
    # The lock spans the whole window — from the read of the recorded table to the commit — so it is
    # taken here and not inside the run. A `write=False` caller touches no file and takes none.
    with spec_lock(spec_dir, enabled=write, error=EnrichmentError):
        return _run_enrichment(
            spec_dir,
            mode=mode,
            offline=offline,
            ensembl_cache=ensembl_cache,
            clinvar_cache=clinvar_cache,
            pubmind_cache=pubmind_cache,
            civic_cache=civic_cache,
            use_clinvar=use_clinvar,
            use_gnomad=use_gnomad,
            download=download,
            genome_build=genome_build,
            write=write,
            mint_vrs=mint_vrs,
            verify_ref=verify_ref,
            verify_clinsig=verify_clinsig,
            verify_rsids=verify_rsids,
            verify_datasets=verify_datasets,
            keep_par_twin=keep_par_twin,
            rederive=rederive,
            keep_staging=keep_staging,
            progress=progress,
            resolver=resolver,
            gnomad_client=gnomad_client,
            grch37_client=grch37_client,
            release_probes=release_probes,
        )