Skip to content

just_dna_compiler.compiler

just_dna_compiler.compiler

Module spec compiler: validates a spec directory and compiles it to a composed multi-parquet artifact plus a manifest.json. A module composes from optional table kinds (RM2): the three-parquet SNP core (weights, annotations, studies) when it carries variants, plus one parquet per 0.4 table kind it includes (diplotypes, pharm_variants, pgs, the binning kinds, …). ARTIFACT_PARQUETS is the roster and len(ARTIFACT_PARQUETS) the count — stated that way rather than spelled, because this sentence carried a spelled-out figure for two releases while the tuple grew past it (RM218).

Public API

validate_spec(spec_dir) -> ValidationResult compile_module(spec_dir, output_dir, ...) -> CompilationResult (emits manifest.json) reverse_module(parquet_dir, output_dir, ...) -> Path

The DSL/manifest schema comes from just-dna-format; this package is the transform between them.

SpecError

Bases: ValueError

A module_spec.yaml that could not be loaded — see load_spec.

A ValueError subclass so a caller that already brackets a load with except (OSError, ValueError), the way read_verification's callers do, keeps working without knowing this type exists.

table_bindings

table_bindings() -> dict[str, tuple[str, ...]]

csv -> the parquet(s) it becomes, assembled from the registries and never named by hand.

Public because two surfaces outside this module need it and both would otherwise keep a copy: the docs site's generated per-table reference, and test_artifact_parquet_bindings.py, which asserts the union of the values is exactly ARTIFACT_PARQUETS. That equality is what makes the map total — a new table kind whose parquet is written by a literal at its own write site fails the test instead of quietly missing a reference page.

Four registries, each owning its own slice: SNP_CORE_PARQUETS (the core, where variants.csv fans out to two), _TABLE_KINDS (the optional authored kinds), _FACT_TABLES (the derived facts) and OVERRIDES_PARQUET (the overlay, whose CSV is authored but is not a table kind).

Source code in compiler/src/just_dna_compiler/compiler.py
def table_bindings() -> dict[str, tuple[str, ...]]:
    """`csv -> the parquet(s) it becomes`, assembled from the registries and never named by hand.

    Public because two surfaces outside this module need it and both would otherwise keep a copy: the
    docs site's generated per-table reference, and `test_artifact_parquet_bindings.py`, which asserts the
    union of the values is exactly `ARTIFACT_PARQUETS`. That equality is what makes the map **total** — a
    new table kind whose parquet is written by a literal at its own write site fails the test instead of
    quietly missing a reference page.

    Four registries, each owning its own slice: `SNP_CORE_PARQUETS` (the core, where `variants.csv`
    fans out to two), `_TABLE_KINDS` (the optional authored kinds), `_FACT_TABLES` (the derived facts)
    and `OVERRIDES_PARQUET` (the overlay, whose CSV is authored but is not a table kind).
    """
    bound: dict[str, tuple[str, ...]] = dict(SNP_CORE_PARQUETS)
    for csv_name, parquet, _model in _TABLE_KINDS:
        bound[csv_name] = (parquet,)
    for csv_name, parquet, _model in _FACT_TABLES:
        bound[csv_name] = (parquet,)
    bound[OVERRIDES_CSV] = (OVERRIDES_PARQUET,)
    return bound

authored_input_entries

authored_input_entries(spec_dir: Path) -> list[FileEntry]

The authored files a module is made of, hashed for the verification binding.

Public because two tiers must agree on it byte for byte (the RM41 lesson): the compiler recomputes the binding from this set when it decides whether to publish manifest.verification, and the enricher hashes the identical set into the attestation's module_hash so a later compile can tell whether the spec has been edited since the checks ran. A private symbol here would leave the enricher choosing between reaching into a private name and re-implementing the list, and a re-implementation is a place for the two to drift — at which point every attestation this workspace writes reads as stale to its own compiler.

Authored files only, which is the boundary just_dna_format.verification argues at length: the derived sidecars carry per-run noise (fetched_at) that would invalidate an attestation on a re-enrichment that changed nothing anyone claimed.

The bytes are newline-normalized, and manifest.inputs[] is deliberately not (RM82). This docstring said the compiler hashes these into manifest.inputs; it never did — that field is filled independently, by file_entries(spec_dir, _INPUT_FILES) over the raw bytes, and the two were only ever equal by coincidence of computing the same thing. Since 0.6 they are two different questions asked of one file set. manifest.inputs[] and artifact.digest answer are these the exact bytes, so they follow every byte, line endings included. The binding answers is this still the module those checks were put against, and an editor rewriting \r\n as \n changes no value, so it must not un-close a module. The asymmetry is the decision, not an inconsistency to tidy: see integrity.newline_normalized_file_entry for the transform and where it stops.

Source code in compiler/src/just_dna_compiler/compiler.py
def authored_input_entries(spec_dir: Path) -> list[FileEntry]:
    """The authored files a module is made of, hashed for the verification binding.

    Public because **two tiers must agree on it byte for byte** (the RM41 lesson): the compiler
    recomputes the binding from this set when it decides whether to publish `manifest.verification`,
    and the enricher hashes the identical set into the attestation's `module_hash` so a later compile
    can tell whether the spec has been edited since the checks ran. A private symbol here would leave
    the enricher choosing between reaching into a private name and re-implementing the list, and a
    re-implementation is a place for the two to drift — at which point every attestation this
    workspace writes reads as stale to its own compiler.

    **Authored files only**, which is the boundary `just_dna_format.verification` argues at length:
    the derived sidecars carry per-run noise (`fetched_at`) that would invalidate an attestation on a
    re-enrichment that changed nothing anyone claimed.

    **The bytes are newline-normalized, and `manifest.inputs[]` is deliberately not (RM82).** This
    docstring said the compiler hashes *these* into `manifest.inputs`; it never did — that field is
    filled independently, by `file_entries(spec_dir, _INPUT_FILES)` over the raw bytes, and the two
    were only ever equal by coincidence of computing the same thing. Since 0.6 they are two different
    questions asked of one file set. `manifest.inputs[]` and `artifact.digest` answer *are these the
    exact bytes*, so they follow every byte, line endings included. The binding answers *is this still
    the module those checks were put against*, and an editor rewriting `\\r\\n` as `\\n` changes no
    value, so it must not un-close a module. The asymmetry is the decision, not an inconsistency to
    tidy: see `integrity.newline_normalized_file_entry` for the transform and where it stops.
    """
    return newline_normalized_file_entries(Path(spec_dir), list(_INPUT_FILES))

load_spec

load_spec(
    path: Path,
    *,
    authority_keys: Iterable[str] | None = None,
) -> ModuleSpecConfig

Load and validate a module_spec.yaml, raising on anything wrong (S74).

The public route to a ModuleSpecConfig. The model has always been exported from just_dna_format.spec and the only thing that produced one was _load_yaml, underscored — so a consumer wanting weighting: or authorship: had to yaml.safe_load the file and read a raw dict, losing the authority-key handling and every diagnosis below, and carrying PyYAML for no reason except that ours was unreachable. Sibling of just_dna_format.read_manifest and read_verification: same shape, same contract, one per file a module carries.

It lives here rather than in the format tier because loading it needs pyyaml, and the format tier is pydantic + cryptography by charter. A consumer already depending on the compiler — which every caller of validate_spec is — can drop their own PyYAML with this.

authority_keys is inject-only and unchanged: pass just_dna_format.normalize.IDENTITY_AUTHORITY_KEYS so a registry-stamped namespace:/owner:/canonical_id: is stripped before validation rather than tripping extra="forbid". The format applies none by default, and a key outside the injected set still trips. Which keys were dropped is not reported here — that is validate_spec's .info, and a caller who needs it wants the validator rather than the loader.

Raises SpecError with every diagnosis joined, where _load_yaml returns them for accumulation. That difference is the whole reason both exist: validate_spec collects errors from a dozen sources and reports them together, which is right for a validator and wrong for a loader — a caller who just wants the object should not have to check a tuple's second element to find out it got None.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_spec(path: Path, *, authority_keys: Iterable[str] | None = None) -> ModuleSpecConfig:
    """Load and validate a `module_spec.yaml`, raising on anything wrong (S74).

    The public route to a `ModuleSpecConfig`. The model has always been exported from
    `just_dna_format.spec` and the only thing that produced one was `_load_yaml`, underscored — so a
    consumer wanting `weighting:` or `authorship:` had to `yaml.safe_load` the file and read a raw
    dict, losing the authority-key handling and every diagnosis below, and carrying **PyYAML** for no
    reason except that ours was unreachable. Sibling of `just_dna_format.read_manifest` and
    `read_verification`: same shape, same contract, one per file a module carries.

    It lives here rather than in the format tier because loading it needs `pyyaml`, and the format
    tier is `pydantic` + `cryptography` by charter. A consumer already depending on the compiler —
    which every caller of `validate_spec` is — can drop their own PyYAML with this.

    `authority_keys` is inject-only and unchanged: pass
    `just_dna_format.normalize.IDENTITY_AUTHORITY_KEYS` so a registry-stamped
    `namespace:`/`owner:`/`canonical_id:` is stripped before validation rather than tripping
    `extra="forbid"`. The format applies none by default, and a key outside the injected set still
    trips. **Which keys were dropped is not reported here** — that is `validate_spec`'s `.info`, and a
    caller who needs it wants the validator rather than the loader.

    Raises `SpecError` with every diagnosis joined, where `_load_yaml` returns them for accumulation.
    That difference is the whole reason both exist: `validate_spec` collects errors from a dozen
    sources and reports them together, which is right for a validator and wrong for a loader — a
    caller who just wants the object should not have to check a tuple's second element to find out
    it got `None`.
    """
    config, errors, _dropped = _load_yaml(Path(path), authority_keys)
    if config is None:
        raise SpecError("; ".join(errors) or f"module_spec.yaml could not be loaded from {path}")
    return config

load_csv_rows

load_csv_rows(
    path: Path,
    row_model: type,
    file_label: str,
    genome_build: str = DEFAULT_GENOME_BUILD,
) -> tuple[list[Any], list[str], list[str]]

Load a CSV and validate each row against a Pydantic model. Returns (rows, errors, warnings).

Public as of 0.5.1 (RM41), and it was public in practice long before. This is the only correct way to turn an authored CSV into row models, just-dna-enricher consumes it across a package boundary in a dozen places, and a downstream consumer wiring the pipeline server-side had the choice of reaching for a private symbol or re-implementing it. Re-implementing is a trap rather than a chore, because it is not csv.DictReader plus Model(**row) — it carries the two rules below, each of which this workspace has already had to fix once:

  • an empty cell becomes None, and the key is kept. MeasureBinRow.measure_kind has a default, so is_required() is False, but the model then receives None rather than its default and fails on type. A "" where this would have put None is a different failure again.
  • genome_build is told to each row (below), so a loader that omits it mints GRCh38 identities for a GRCh37 module.

_load_csv_rows remains as an alias so no caller breaks.

genome_build is told to each row, not read from it. A coordinate is not absolute, so a row deriving an identity from one needs the module's assembly — and a pydantic model built from a CSV dict has no module_spec.yaml in scope. Injecting it here keeps the build declared exactly once (the yaml) while reaching every row: it is a private attribute, so it is not a column, reaches no parquet, and moves no digest. See AuthoredModel._genome_build for why per-row or per-CSV declaration was rejected. Callers that load a build-independent table — the resolution and fact sidecars, which carry their own genome_build column — leave it at the default and are unaffected.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_csv_rows(
    path: Path, row_model: type, file_label: str, genome_build: str = DEFAULT_GENOME_BUILD
) -> tuple[list[Any], list[str], list[str]]:
    """Load a CSV and validate each row against a Pydantic model. Returns (rows, errors, warnings).

    **Public as of 0.5.1 (RM41), and it was public in practice long before.** This is the only correct
    way to turn an authored CSV into row models, `just-dna-enricher` consumes it across a package
    boundary in a dozen places, and a downstream consumer wiring the pipeline server-side had the
    choice of reaching for a private symbol or re-implementing it. Re-implementing is a trap rather
    than a chore, because it is not `csv.DictReader` plus `Model(**row)` — it carries the two rules
    below, each of which this workspace has already had to fix once:

    * **an empty cell becomes `None`, and the key is kept.** `MeasureBinRow.measure_kind` has a
      default, so `is_required()` is `False`, but the model then receives `None` rather than its
      default and fails on type. A `""` where this would have put `None` is a different failure again.
    * **`genome_build` is told to each row** (below), so a loader that omits it mints GRCh38
      identities for a GRCh37 module.

    `_load_csv_rows` remains as an alias so no caller breaks.

    `genome_build` is **told to each row**, not read from it. A coordinate is not absolute, so a row
    deriving an identity from one needs the module's assembly — and a pydantic model built from a CSV
    dict has no `module_spec.yaml` in scope. Injecting it here keeps the build declared exactly once
    (the yaml) while reaching every row: it is a private attribute, so it is not a column, reaches no
    parquet, and moves no digest. See `AuthoredModel._genome_build` for why per-row or per-CSV
    declaration was rejected. Callers that load a build-independent table — the resolution and fact
    sidecars, which carry their own `genome_build` column — leave it at the default and are unaffected.
    """
    errors: list[str] = []
    rows: list[Any] = []
    if not path.exists():
        return [], [f"{file_label} not found at {path}"], []

    with open(path, encoding="utf-8", newline="") as f:
        reader = csv.DictReader(f)
        if reader.fieldnames is None:
            return [], [f"{file_label} has no header row"], []
        for line_num, raw_row in enumerate(reader, start=2):
            # DictReader buckets any cells past the header under the `None` key (a list). Silently
            # dropping them would let a shifted/surplus column slip past `extra="forbid"` — a real
            # authoring error (a misaligned row) reported as valid. Flag a non-empty surplus instead.
            surplus = [s for s in (raw_row.get(None) or []) if isinstance(s, str) and s.strip()]
            if surplus:
                errors.append(
                    f"{file_label} line {line_num}: more values than header columns "
                    f"(surplus: {surplus}) — check for a shifted or extra column"
                )
                continue
            cleaned = {
                k.strip(): (v.strip() if isinstance(v, str) and v.strip() != "" else None)
                for k, v in raw_row.items()
                if k is not None
            }
            try:
                row = row_model.model_validate(cleaned)
                if isinstance(row, AuthoredModel):
                    row.with_genome_build(genome_build)
                rows.append(row)
            except ValidationError as exc:
                for err in exc.errors():
                    loc = " → ".join(str(x) for x in err["loc"])
                    errors.append(f"{file_label} line {line_num} [{loc}]: {err['msg']}")
    return rows, errors, []

load_spec_variants

load_spec_variants(
    spec_dir: Path,
) -> tuple[list[VariantRow], list[str], list[str]]

A spec directory's variants.csv, loaded and re-stamped for the build the module declares.

The other half of RM41. Two enricher checks take rows rather than a spec_dir — acmg.verify_acmg_sf and identifiers.check_identifiers — unlike every other pass, so a caller has to do this itself, and doing it right means three steps rather than one: read the declared build out of module_spec.yaml, inject it into every row, and then re-stamp the identities, since VariantRow._freeze_identity runs at construction where the yaml is not in scope.

Missing or unreadable yaml falls back to DEFAULT_GENOME_BUILD, matching what compiling that directory would assume — this is a read-only check helper, not the enrichment path, which refuses rather than choose a build for a module whose declaration cannot be read (it writes facts back).

Returns (variants, errors, warnings); an absent variants.csv is an empty list and one error, exactly as load_csv_rows reports it.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_spec_variants(spec_dir: Path) -> tuple[list[VariantRow], list[str], list[str]]:
    """A spec directory's `variants.csv`, loaded and re-stamped for the build the module declares.

    The other half of RM41. Two enricher checks take rows rather than a `spec_dir` —
    `acmg.verify_acmg_sf` and `identifiers.check_identifiers` — unlike every other pass, so a caller
    has to do this itself, and doing it *right* means three steps rather than one: read the declared
    build out of `module_spec.yaml`, inject it into every row, and then re-stamp the identities, since
    `VariantRow._freeze_identity` runs at construction where the yaml is not in scope.

    Missing or unreadable yaml falls back to `DEFAULT_GENOME_BUILD`, matching what compiling that
    directory would assume — this is a read-only check helper, not the enrichment path, which refuses
    rather than choose a build for a module whose declaration cannot be read (it writes facts back).

    Returns `(variants, errors, warnings)`; an absent `variants.csv` is an empty list and one error,
    exactly as `load_csv_rows` reports it.
    """
    spec_dir = Path(spec_dir)
    config = None
    if (spec_dir / "module_spec.yaml").exists():
        config, _, _ = _load_yaml(spec_dir / "module_spec.yaml")
    build = config.genome_build if config else DEFAULT_GENOME_BUILD
    variants, errors, warnings = load_csv_rows(
        spec_dir / "variants.csv", VariantRow, "variants.csv", genome_build=build
    )
    warnings.extend(_restamp_for_build(variants, build))
    return variants, errors, warnings

positional_placement

positional_placement(
    rows_by_csv: dict[str, list[Any]],
) -> tuple[int, int]

(rows, placed) across the three positional table kinds — the counts the manifest publishes.

The structured half of _check_positional_joinability, which reports the same facts as prose per table (S31). Public because the number is what a catalog wants: until 0.6 the only record of how much of a PGx table joins to a VCF was UNJOINABLE_PHRASE inside manifest.compilation.warnings, and a downstream registry substring-matched it. Anything a consumer can only learn from a warning string is an unversioned interface (RM44).

Call it after _apply_positional_resolution, so placed counts what the artifact actually carries rather than what the author typed — the fill is where an rsid-authored PGx module gets its coordinates, and before it the answer is the pre-RM43 one.

A module carrying no positional table returns (0, 0), which is a real answer and distinct from the None the manifest holds for a compile that never counted; see Compilation.positional_rows.

Source code in compiler/src/just_dna_compiler/compiler.py
def positional_placement(rows_by_csv: dict[str, list[Any]]) -> tuple[int, int]:
    """`(rows, placed)` across the three positional table kinds — the counts the manifest publishes.

    The structured half of `_check_positional_joinability`, which reports the same facts as prose per
    table (S31). Public because the number is what a catalog wants: until 0.6 the only record of how
    much of a PGx table joins to a VCF was `UNJOINABLE_PHRASE` inside `manifest.compilation.warnings`,
    and a downstream registry substring-matched it. Anything a consumer can only learn from a warning
    string is an unversioned interface (RM44).

    **Call it after `_apply_positional_resolution`**, so `placed` counts what the artifact actually
    carries rather than what the author typed — the fill is where an rsid-authored PGx module gets its
    coordinates, and before it the answer is the pre-RM43 one.

    A module carrying no positional table returns `(0, 0)`, which is a real answer and distinct from
    the `None` the manifest holds for a compile that never counted; see `Compilation.positional_rows`.
    """
    rows = [row for csv_name, _model in _POSITIONAL_TABLE_KINDS for row in rows_by_csv.get(csv_name) or []]
    placed = [row for row in rows if row.chrom is not None and row.start is not None]
    return len(rows), len(placed)

load_citing_rows

load_citing_rows(spec_dir: Path) -> dict[str, list[Any]]

Every citing table present beside a spec, keyed by CSV name — the annotation rows that ground their own claim with a pmid (MeasureBinRow.pmid RM47, PharmVariantRow.pmid RM132).

Public because a second tier needs it: the enricher's literature pass has to check these pointers alongside studies.csv, and its two alternatives were importing a private symbol or hand-keeping a parallel list of the citing kinds — the RM40/RM41 shape exactly, and the list would go stale on the next kind that declares the column.

Supersedes load_binning_rows, which stays and still means what it always did: it reads the binning kinds only, so a caller wanting the citations a module makes wants this one.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_citing_rows(spec_dir: Path) -> dict[str, list[Any]]:
    """Every **citing** table present beside a spec, keyed by CSV name — the annotation rows that
    ground their own claim with a `pmid` (`MeasureBinRow.pmid` RM47, `PharmVariantRow.pmid` RM132).

    Public because a second tier needs it: the enricher's literature pass has to check these pointers
    alongside `studies.csv`, and its two alternatives were importing a private symbol or hand-keeping a
    parallel list of the citing kinds — the RM40/RM41 shape exactly, and the list would go stale on the
    next kind that declares the column.

    Supersedes `load_binning_rows`, which stays and still means what it always did: it reads the
    binning kinds only, so a caller wanting *the citations a module makes* wants this one.
    """
    return _load_kind_rows(spec_dir, _CITING_TABLE_KINDS)

load_binning_rows

load_binning_rows(
    spec_dir: Path,
) -> dict[str, list[MeasureBinRow]]

Every binning table present beside a spec, keyed by CSV name (RM47).

Narrower than load_citing_rows since 0.7 and deliberately kept: a caller asking for the binning kinds is asking about thresholds, not about citations, and quietly widening what it returns would hand such a caller PharmVariantRows where it expects MeasureBinRows.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_binning_rows(spec_dir: Path) -> dict[str, list[MeasureBinRow]]:
    """Every **binning** table present beside a spec, keyed by CSV name (RM47).

    Narrower than `load_citing_rows` since 0.7 and deliberately kept: a caller asking for the binning
    kinds is asking about thresholds, not about citations, and quietly widening what it returns would
    hand such a caller `PharmVariantRow`s where it expects `MeasureBinRow`s.
    """
    return _load_kind_rows(spec_dir, _BINNING_TABLE_KINDS)

table_citations

table_citations(
    rows_by_csv: dict[str, list[Any]],
) -> list[str]

Digit-only PMIDs the module's annotation tables cite, de-duplicated — every citing kind (MeasureBinRow.pmid RM47, PharmVariantRow.pmid RM132).

Takes the whole table-kind map a caller already holds and reads only the citing kinds out of it, so a caller cannot accidentally hand over a haplotypes.csv (no pmid column) and get an attribute error. The kind set is derived from the models, never hand-listed, so a kind that gains the column is read here without an edit.

First-occurrence order rather than sorted, because it feeds emission order downstream (P7), and normalization goes through extract_pmids so a table pointer and studies.csv cannot drift into two spellings of one citation.

Source code in compiler/src/just_dna_compiler/compiler.py
def table_citations(rows_by_csv: dict[str, list[Any]]) -> list[str]:
    """Digit-only PMIDs the module's annotation tables cite, de-duplicated — **every** citing kind
    (`MeasureBinRow.pmid` RM47, `PharmVariantRow.pmid` RM132).

    Takes the whole table-kind map a caller already holds and reads only the citing kinds out of it, so
    a caller cannot accidentally hand over a `haplotypes.csv` (no `pmid` column) and get an attribute
    error. The kind set is derived from the models, never hand-listed, so a kind that gains the column
    is read here without an edit.

    First-occurrence order rather than sorted, because it feeds emission order downstream (P7), and
    normalization goes through `extract_pmids` so a table pointer and `studies.csv` cannot drift into
    two spellings of one citation.
    """
    return _citations_over(rows_by_csv, _CITING_TABLE_KINDS)

binning_citations

binning_citations(
    rows_by_csv: dict[str, list[Any]],
) -> list[str]

Digit-only PMIDs the binning tables cite (MeasureBinRow.pmid), de-duplicated.

Narrowed by table_citations since 0.7 the same way load_binning_rows is by load_citing_rows, and kept for the same reason. Nothing inside the compiler calls it any more: the literature cross-check must read every citation site, and this one answers a question about thresholds.

Source code in compiler/src/just_dna_compiler/compiler.py
def binning_citations(rows_by_csv: dict[str, list[Any]]) -> list[str]:
    """Digit-only PMIDs the **binning** tables cite (`MeasureBinRow.pmid`), de-duplicated.

    Narrowed by `table_citations` since 0.7 the same way `load_binning_rows` is by `load_citing_rows`,
    and kept for the same reason. Nothing inside the compiler calls it any more: the literature
    cross-check must read every citation site, and this one answers a question about thresholds.
    """
    return _citations_over(rows_by_csv, _BINNING_TABLE_KINDS)

load_overlay

load_overlay(
    spec_dir: Path,
) -> tuple[list[OverrideRow], list[str], list[str]]

Read overrides.csv and answer with (rows, errors, warnings) (RM124).

Public since RM136, because the enricher needs it: an author who corrects a derived cell through the overlay must not go on being told the same finding by the tier that writes the file. Private, it would have been reached into the way load_spec was before S74, or — worse — reimplemented in the enricher, which is the drift the overlay's own design refuses.

One loader for both public entry points, because both have to read it and a second copy is where validate and compile learn to disagree — the parity rule this module keeps re-learning (@validate-refuses-all). Everything it reports is structural: the rows parse or they do not, the key is duplicated or it is not, a key group carries one operation or several. Nothing here consults a derived table, which is what keeps every finding identical on both laps of a round trip.

A file that is present with no rows is an error, the same answer _TABLE_KINDS gives: an empty authored table is a header somebody meant to fill.

Source code in compiler/src/just_dna_compiler/compiler.py
def load_overlay(spec_dir: Path) -> tuple[list[OverrideRow], list[str], list[str]]:
    """Read `overrides.csv` and answer with `(rows, errors, warnings)` (RM124).

    **Public since RM136**, because the enricher needs it: an author who corrects a derived cell
    through the overlay must not go on being told the same finding by the tier that writes the file.
    Private, it would have been reached into the way `load_spec` was before S74, or — worse —
    reimplemented in the enricher, which is the drift the overlay's own design refuses.

    One loader for both public entry points, because both have to read it and a second copy is where
    `validate` and `compile` learn to disagree — the parity rule this module keeps re-learning
    (`@validate-refuses-all`). Everything it reports is structural: the rows parse or they do not, the
    key is duplicated or it is not, a key group carries one operation or several. Nothing here
    consults a derived table, which is what keeps every finding identical on both laps of a round
    trip.

    A file that is present with no rows is an **error**, the same answer `_TABLE_KINDS` gives: an
    empty authored table is a header somebody meant to fill.
    """
    path = spec_dir / OVERRIDES_CSV
    if not path.is_file():
        return [], [], []
    rows, errors, warnings = load_csv_rows(path, OverrideRow, OVERRIDES_CSV)
    if errors:
        return [], errors, warnings
    if not rows:
        return [], [f"{OVERRIDES_CSV} is present but has no rows."], warnings
    key_errors, key_warnings = _validate_table_kind(OVERRIDES_CSV, OverrideRow, rows)
    errors.extend(key_errors)
    warnings.extend(key_warnings)
    errors.extend(overlay_coherence_errors(rows))
    return rows, errors, warnings

validate_spec

validate_spec(
    spec_dir: Path,
    authority_keys: Iterable[str] | None = None,
    *,
    strict: bool = False,
    resolve_with_ensembl: bool = True,
) -> ValidationResult

Validate a module spec directory without producing output.

authority_keys (inject-only) is the set of consumer/registry-owned identity keys to strip from the authored module: block before validation — pass just_dna_format.normalize. IDENTITY_AUTHORITY_KEYS (or your own set) so a legacy spec carrying namespace:/owner:/ canonical_id: validates; the format applies none by default. Stripped keys are surfaced on .info. Everything else still trips extra="forbid".

strict mirrors compile_module's flag and exists for one reason: several checks are a mode ladder (warning in best_effort, error in strict), so without a mode here the pre-flight could not answer the question the author actually asked — the documented order is validate then compile --strict, and a modeless validate is a pre-flight for the other compile.

It mirrors compile_module's severities exactly, which is not the same as changing severity only — this said the latter, in both of the two docstrings carrying it, and it was false (RM218). Two findings are aggregates with no best_effort counterpart sentence: the unresolved-position refusal (strict compile: N variant(s) …) and build_disagreement_error. Their best_effort rung is a different sentence — the per-subject rsid_unresolved warning, which fires in both modes — so under strict the aggregate is genuinely added rather than promoted. The contract the two commands share is that validate(strict=x) and compile(strict=x) reach the same verdict, not that the two modes of validate differ by a severity column.

resolve_with_ensembl mirrors it for the same reason and is passed through by compile_module. The pre-flight applies the injected table to the positional 0.4 tables (RM43), and that decides whether their rows are reported as unjoinable — so a validate that ignored the master resolution switch would be more optimistic than the compile it precedes, which is the disagreement direction the parity rule exists to prevent.

Stats include genes/categories as lists (filtering None) plus variant_count, gene_count, study_count, and the ClinVar quality counts (clinvar_count/pathogenic_count/benign_count) — the fields the manifest needs. See ValidationResult.stats for the full key contract.

warnings is the complete list, unchanged; carried names the subset an author cannot clear and warnings_summary counts them by code (RM131). A caller that wants to keep building on this run's findings — as compile_module and close_module do — calls _validate_spec instead, for the reason given there.

Source code in compiler/src/just_dna_compiler/compiler.py
def validate_spec(
    spec_dir: Path,
    authority_keys: Iterable[str] | None = None,
    *,
    strict: bool = False,
    resolve_with_ensembl: bool = True,
) -> ValidationResult:
    """Validate a module spec directory without producing output.

    `authority_keys` (inject-only) is the set of consumer/registry-owned identity keys to strip from
    the authored `module:` block before validation — pass `just_dna_format.normalize.
    IDENTITY_AUTHORITY_KEYS` (or your own set) so a legacy spec carrying `namespace:`/`owner:`/
    `canonical_id:` validates; the format applies none by default. Stripped keys are surfaced on
    `.info`. Everything else still trips `extra="forbid"`.

    `strict` mirrors `compile_module`'s flag and exists for one reason: several checks are a **mode
    ladder** (warning in `best_effort`, error in `strict`), so without a mode here the pre-flight
    could not answer the question the author actually asked — the documented order is `validate` then
    `compile --strict`, and a modeless `validate` is a pre-flight for the *other* compile.

    **It mirrors `compile_module`'s severities exactly, which is not the same as changing severity
    only** — this said the latter, in both of the two docstrings carrying it, and it was false
    (RM218). Two findings are *aggregates* with no `best_effort` counterpart sentence: the
    unresolved-position refusal (`strict compile: N variant(s) …`) and `build_disagreement_error`.
    Their `best_effort` rung is a different sentence — the per-subject `rsid_unresolved` warning,
    which fires in **both** modes — so under `strict` the aggregate is genuinely *added* rather than
    promoted. The contract the two commands share is that **`validate(strict=x)` and
    `compile(strict=x)` reach the same verdict**, not that the two modes of `validate` differ by a
    severity column.

    `resolve_with_ensembl` mirrors it for the same reason and is passed through by `compile_module`.
    The pre-flight applies the injected table to the positional 0.4 tables (RM43), and that decides
    whether their rows are reported as unjoinable — so a `validate` that ignored the master resolution
    switch would be *more optimistic* than the compile it precedes, which is the disagreement
    direction the parity rule exists to prevent.

    Stats include `genes`/`categories` as lists (filtering None) plus `variant_count`,
    `gene_count`, `study_count`, and the ClinVar quality counts
    (`clinvar_count`/`pathogenic_count`/`benign_count`) — the fields the manifest needs. See
    `ValidationResult.stats` for the full key contract.

    `warnings` is the complete list, unchanged; `carried` names the subset an author cannot clear and
    `warnings_summary` counts them by code (RM131). A caller that wants to keep *building* on this
    run's findings — as `compile_module` and `close_module` do — calls `_validate_spec` instead, for
    the reason given there.
    """
    return _validate_spec(spec_dir, authority_keys, strict=strict, resolve_with_ensembl=resolve_with_ensembl)[
        0
    ]

variant_stats

variant_stats(variants: list[VariantRow]) -> dict[str, Any]

The variants.csv-derived facets of ValidationResult.stats / manifest.stats.

Its own function because it now has two callers, and the second is why: compile_module may discard a row for carrying an unusable symbolic allele (RM5), and the stats were computed by validate_spec before that happened. weights_rows counts the parquet and so is post-drop, so a published manifest claimed a variant_count one higher than the artifact contained — the RM44 class of defect exactly, a manifest number a catalog keys on and cannot check.

Source code in compiler/src/just_dna_compiler/compiler.py
def variant_stats(variants: list[VariantRow]) -> dict[str, Any]:
    """The `variants.csv`-derived facets of `ValidationResult.stats` / `manifest.stats`.

    Its own function because it now has **two** callers, and the second is why: `compile_module` may
    discard a row for carrying an unusable symbolic allele (RM5), and the stats were computed by
    `validate_spec` before that happened. `weights_rows` counts the parquet and so is post-drop, so a
    published manifest claimed a `variant_count` one higher than the artifact contained — the RM44
    class of defect exactly, a manifest number a catalog keys on and cannot check.
    """
    genes = sorted({v.gene for v in variants if v.gene})
    return {
        "variant_count": len({v.variant_key for v in variants}),
        "unique_rsids": len({v.rsid for v in variants if v.rsid is not None}),
        "gene_count": len(genes),
        "genes": genes,
        "categories": sorted({v.category for v in variants if v.category}),
        # ClinVar/quality flag counts over variant rows (ROADMAP item 5).
        "clinvar_count": sum(1 for v in variants if v.clinvar),
        "pathogenic_count": sum(1 for v in variants if v.pathogenic),
        "benign_count": sum(1 for v in variants if v.benign),
    }

module_stats

module_stats(
    variants: list[VariantRow],
    kind_rows: dict[str, list[Any]] | None = None,
) -> dict[str, Any]

variant_stats plus the gene facets taken over every authored table, not just variants.

PUBLIC, and it exists rather than a second parameter on variant_stats because that function's name is a promise about which table it reads and renaming it would be a major (S14's rule). What the two return differs in exactly two keys.

stats describes the module, and Stats has always said so — "card/detail stats derived from the spec", not from one table of it. variant_stats nevertheless derived genes from variants.csv alone, so a module whose lead table is diplotypes.csv, allele_function.csv, copynumbers.csv or any other gene-bearing kind published gene_count: 0, genes: [] however many of its rows named a gene — and a registry's gene index is fed from that field, so the module was unreachable by a gene search (S57). Measured on reference_examples/cyp2c19_star_alleles/: 1,332 rows carrying gene=CYP2C19 across three tables, and genes: [].

The honest workaround an author was left with was prose in the README, and the dishonest one — inventing an empty variants.csv to be discoverable — is what makes this ours to fix rather than a documentation note.

Only authored kinds count. _GENE_BEARING_TABLE_KINDS derives from _TABLE_KINDS, which deliberately excludes the derived fact sidecars, so a gene reaching gene_metrics.csv because a pass looked it up never becomes a gene the module claims to be about.

Source code in compiler/src/just_dna_compiler/compiler.py
def module_stats(variants: list[VariantRow], kind_rows: dict[str, list[Any]] | None = None) -> dict[str, Any]:
    """`variant_stats` plus the gene facets taken over **every** authored table, not just variants.

    PUBLIC, and it exists rather than a second parameter on `variant_stats` because that function's
    name is a promise about which table it reads and renaming it would be a major (S14's rule). What
    the two return differs in exactly two keys.

    **`stats` describes the module, and `Stats` has always said so** — *"card/detail stats derived from
    the spec"*, not from one table of it. `variant_stats` nevertheless derived `genes` from
    `variants.csv` alone, so a module whose lead table is `diplotypes.csv`, `allele_function.csv`,
    `copynumbers.csv` or any other gene-bearing kind published `gene_count: 0, genes: []` however many
    of its rows named a gene — and a registry's gene index is fed from that field, so the module was
    unreachable by a gene search (S57). Measured on `reference_examples/cyp2c19_star_alleles/`: 1,332
    rows carrying `gene=CYP2C19` across three tables, and `genes: []`.

    The honest workaround an author was left with was prose in the README, and the *dishonest* one —
    inventing an empty `variants.csv` to be discoverable — is what makes this ours to fix rather than a
    documentation note.

    Only authored kinds count. `_GENE_BEARING_TABLE_KINDS` derives from `_TABLE_KINDS`, which
    deliberately excludes the derived fact sidecars, so a gene reaching `gene_metrics.csv` because a
    pass looked it up never becomes a gene the module claims to be about.
    """
    stats = variant_stats(variants)
    genes = set(stats["genes"])
    for csv_name, _model in _GENE_BEARING_TABLE_KINDS:
        for row in (kind_rows or {}).get(csv_name) or []:
            gene = getattr(row, "gene", None)
            if gene:
                genes.add(gene)
    ordered = sorted(genes)
    stats["gene_count"] = len(ordered)
    stats["genes"] = ordered
    return stats

spec_tables

spec_tables(
    spec_dir: Path,
) -> tuple[dict[str, list[Any]], str]

The parsed, defaults-folded authored rows content_signature hashes, and the declared build.

PUBLIC, and the reason is that everything finer than a whole-module hash needs these rows and had no way to get them (S53). content_signature returned only the digest, so a tool answering what moved between two versions of this module — per table, per row — had to rebuild the mapping outside, and rebuilding it meant restating two private things: the table roster (_TABLE_KINDS) and the defaults: fold (_resolve_spec_defaults, _DEFAULTED_VARIANT_FIELDS).

The fold is the part that silently produces a wrong answer, which is why this returns the finished mapping rather than exporting the pieces. A caller hashing load_csv_rows output directly gets a different digest from content_signature for the same module: measured on reference_examples/hfe_hemochromatosis, writing one curator value on every variant row in one copy and the identical value under defaults: in another, content_signature agrees across the pair (correct — RM37) while the raw-rows build disagrees, so a per-table comparison built the obvious way reports twelve changed rows where there are none. Exporting _TABLE_KINDS and _resolve_spec_defaults separately would hand out three pieces that must be assembled in one order — load with the declared build injected, fold, then hash — and the order is the easy half to get wrong. One function that returns the finished mapping cannot be assembled wrongly.

The roster is authored tables only, so the licensing table is outside it: sources.csv / licensing.csv is hashed by integrity.source_signature instead, and neither renaming it nor editing a cell in it moves content_signature. Both verified on the same example. That is correct — the licence layer is its own identity — and it is stated here because it is the one authored, hand-editable table a licence audit sends an author looking for.

Raises ValueError if a present data CSV fails validation, exactly as content_signature does: the contract carries over unchanged, because that function is now this one plus the hash.

Source code in compiler/src/just_dna_compiler/compiler.py
def spec_tables(spec_dir: Path) -> tuple[dict[str, list[Any]], str]:
    """The parsed, defaults-folded authored rows `content_signature` hashes, and the declared build.

    PUBLIC, and the reason is that everything finer than a whole-module hash needs these rows and had
    no way to get them (S53). `content_signature` returned only the digest, so a tool answering *what
    moved between two versions of this module* — per table, per row — had to rebuild the mapping
    outside, and rebuilding it meant restating two private things: the table roster (`_TABLE_KINDS`)
    and the `defaults:` fold (`_resolve_spec_defaults`, `_DEFAULTED_VARIANT_FIELDS`).

    **The fold is the part that silently produces a wrong answer**, which is why this returns the
    finished mapping rather than exporting the pieces. A caller hashing `load_csv_rows` output directly
    gets a different digest from `content_signature` for the same module: measured on
    `reference_examples/hfe_hemochromatosis`, writing one `curator` value on every variant row in one
    copy and the identical value under `defaults:` in another, `content_signature` agrees across the
    pair (correct — RM37) while the raw-rows build disagrees, so a per-table comparison built the
    obvious way reports twelve changed rows where there are none. Exporting `_TABLE_KINDS` and
    `_resolve_spec_defaults` separately would hand out three pieces that must be assembled in one
    order — load with the declared build injected, fold, then hash — and the order is the easy half to
    get wrong. One function that returns the finished mapping cannot be assembled wrongly.

    **The roster is authored tables only**, so the licensing table is outside it: `sources.csv` /
    `licensing.csv` is hashed by `integrity.source_signature` instead, and neither renaming it nor
    editing a cell in it moves `content_signature`. Both verified on the same example. That is correct
    — the licence layer is its own identity — and it is stated here because it is the one authored,
    hand-editable table a licence audit sends an author looking for.

    Raises `ValueError` if a present data CSV fails validation, exactly as `content_signature` does:
    the contract carries over unchanged, because that function is now this one plus the hash.
    """
    spec_dir = Path(spec_dir)
    # The declared build is part of the content, not metadata about it: identical coordinate rows on
    # two assemblies describe two different loci. A spec whose yaml will not load falls back to the
    # format default — this function raises on an invalid *data* CSV, and an unreadable yaml is
    # `validate_spec`'s finding to report, not this one's.
    config, _, _ = _load_yaml(spec_dir / "module_spec.yaml")
    declared_build = config.genome_build if config else DEFAULT_GENOME_BUILD
    kinds: list[tuple[str, type[BaseModel]]] = [
        ("variants.csv", VariantRow),
        ("studies.csv", StudyRow),
        *((csv_name, model) for csv_name, _parquet, model in _TABLE_KINDS),
        # The overlay is authored input, so it is content (RM124) — by its value cells. Its three
        # provenance cells (`reason`/`decided_by`/`decided_at`) are `exclude=True` on the model and
        # so outside the hash (S87): rewording a reason is a patch, not a new identity. No published
        # module's signature moves: `content_signature` skips a table this loop finds no file for,
        # exactly as an unset optional column contributes nothing, and no module published to date
        # carries one.
        (OVERRIDES_CSV, OverrideRow),
    ]
    tables: dict[str, list[Any]] = {}
    for csv_name, model in kinds:
        path = spec_dir / csv_name
        if not path.is_file():
            continue
        rows, errors, _ = _load_csv_rows(path, model, csv_name, genome_build=declared_build)
        if errors:
            raise ValueError(f"cannot compute content_signature: {csv_name} is invalid: {errors[0]}")
        if model is VariantRow:
            # The only model carrying `Defaults`' fields. Safe to mutate: these rows were loaded here
            # and go nowhere else — the compile path loads its own copy.
            _resolve_spec_defaults(rows, config.defaults if config else Defaults())
        tables[csv_name] = rows
    return tables, declared_build

content_signature

content_signature(spec_dir: Path) -> str

Stable content identity over the raw authored data CSVs — name- and Ensembl-independent.

Reads variants.csv, studies.csv, and any present 0.4 table CSVs, validates each row, and hashes the normalized + deterministically-sorted rows via just_dna_format.integrity.content_signature. The data is read as authored (no Ensembl resolution, no parquet build), so this is cheap and reference-independent — a client can compute it without recompiling and dedup against a registry, surviving both metadata-strip and a recompile against a different reference. Raises ValueError if a present data CSV fails validation.

"As authored" means the rows, not the spelling: module_spec.yaml's defaults: block is folded into each variant row first (_resolve_spec_defaults, RM37), because a value written once under defaults: and the same value written on every row are the same content.

This is spec_tables plus the hash and nothing else, so a consumer wanting the rows behind the digest — per-table or per-row work — calls that instead of restating the roster and the fold (S53).

Source code in compiler/src/just_dna_compiler/compiler.py
def content_signature(spec_dir: Path) -> str:
    """Stable content identity over the raw authored data CSVs — name- and Ensembl-independent.

    Reads `variants.csv`, `studies.csv`, and any present 0.4 table CSVs, validates each row, and
    hashes the normalized + deterministically-sorted rows via
    `just_dna_format.integrity.content_signature`. The data is read **as authored** (no Ensembl
    resolution, no parquet build), so this is cheap and reference-independent — a client can compute
    it without recompiling and dedup against a registry, surviving both metadata-strip and a recompile
    against a different reference. Raises `ValueError` if a present data CSV fails validation.

    "As authored" means the *rows*, not the *spelling*: `module_spec.yaml`'s `defaults:` block is
    folded into each variant row first (`_resolve_spec_defaults`, RM37), because a value written once
    under `defaults:` and the same value written on every row are the same content.

    This is `spec_tables` plus the hash and nothing else, so a consumer wanting the rows behind the
    digest — per-table or per-row work — calls that instead of restating the roster and the fold (S53).
    """
    return _content_signature(*spec_tables(spec_dir))

compile_module

compile_module(
    spec_dir: Path,
    output_dir: Path,
    compression: str = "zstd",
    resolve_with_ensembl: bool = True,
    ensembl_cache: Path | None = None,
    compiled_by: str | None = None,
    ensembl_reference: str | None = None,
    log_files: list[Path] | None = None,
    provenance_file: Path | None = None,
    logo_file: Path | None = None,
    readme_file: Path | None = None,
    authority_keys: Iterable[str] | None = None,
    strict: bool = False,
    ba1_threshold: float = BA1_ALLELE_FREQUENCY_THRESHOLD,
) -> CompilationResult

Compile a module spec directory into parquet files plus a manifest.json.

Parameters:

Name Type Description Default
spec_dir Path

Path to the module spec directory.

required
output_dir Path

Directory for output parquet files + manifest.json.

required
compression str

Parquet compression codec.

'zstd'
resolve_with_ensembl bool

Master switch for resolution — of every kind, despite the name. With a resolution.csv present it drives the preferred, source-independent table path; ensembl_cache is the deprecated fallback. Setting it False with a resolution.csv present disables that table too, and warns, because the compile then succeeds with no resolved position on any row.

True
ensembl_cache Path | None

Deprecated (removed at 1.0). Path to a prebuilt Ensembl DuckDB or parquet cache dir. In-compiler DuckDB resolution has moved to just-dna-enricher; when given, this emits a DeprecationWarning and routes to the enricher (which must be installed). Prefer producing a resolution.csv (just-dna-enricher enrich) — the compiler then resolves with no reference and no network.

None
compiled_by str | None

Provenance tag for the manifest (the marketplace passes "marketplace-server"; a local compile leaves it None, so downloaders treat it as untrusted).

None
ensembl_reference str | None

Pinned reference id recorded in the manifest for reproducibility.

None
log_files list[Path] | None

Explicit run/provenance log files to record. If None, auto-discovers a top-level *.log plus per-role files under spec_dir/logs/. Logs are optional.

None
provenance_file Path | None

Explicit structured-provenance document. If None, auto-discovers spec_dir/provenance.json. Optional; summarized into manifest.provenance.

None
logo_file Path | None

Explicit module logo image. If None, auto-discovers spec_dir/logo.{png,jpg,jpeg}. Optional; hashed into manifest.logo, kept out of artifact.digest.

None
readme_file Path | None

Explicit module readme. If None, auto-discovers the first of manifest.README_CANDIDATES (README.md first) beside the spec. Optional; hashed into manifest.readme, kept out of artifact.digest and content_signature — prose about the module is not part of its identity, so fixing a caveat is a patch (S25).

None
authority_keys Iterable[str] | None

Inject-only set of consumer/registry-owned identity keys to strip from the authored module: block before validation (e.g. just_dna_format.normalize. IDENTITY_AUTHORITY_KEYS). None strips nothing.

None
strict bool

All-or-nothing compile. When True, fail (rather than emit a partial artifact) if any variant still lacks a resolved genomic position (chrom+start) after resolution — an unresolved position means the injected reference was incomplete/absent and the parquet bytes (hence artifact.digest) would not be reproducible. Default False keeps the best-effort behavior (positions left unset, surfaced as warnings).

False
ba1_threshold float

Allele frequency above which a pathogenic variant draws the ACMG BA1 warning (_check_ba1_lint). Defaults to ACMG's 5%. Raise it for a module curating a common recessive carrier allele, where the default fires on correct data. Warning-only in both modes, so this tunes noise, never whether the compile succeeds.

BA1_ALLELE_FREQUENCY_THRESHOLD
Source code in compiler/src/just_dna_compiler/compiler.py
4845
4846
4847
4848
4849
4850
4851
4852
4853
4854
4855
4856
4857
4858
4859
4860
4861
4862
4863
4864
4865
4866
4867
4868
4869
4870
4871
4872
4873
4874
4875
4876
4877
4878
4879
4880
4881
4882
4883
4884
4885
4886
4887
4888
4889
4890
4891
4892
4893
4894
4895
4896
4897
4898
4899
4900
4901
4902
4903
4904
4905
4906
4907
4908
4909
4910
4911
4912
4913
4914
4915
4916
4917
4918
4919
4920
4921
4922
4923
4924
4925
4926
4927
4928
4929
4930
4931
4932
4933
4934
4935
4936
4937
4938
4939
4940
4941
4942
4943
4944
4945
4946
4947
4948
4949
4950
4951
4952
4953
4954
4955
4956
4957
4958
4959
4960
4961
4962
4963
4964
4965
4966
4967
4968
4969
4970
4971
4972
4973
4974
4975
4976
4977
4978
4979
4980
4981
4982
4983
4984
4985
4986
4987
4988
4989
4990
4991
4992
4993
4994
4995
4996
4997
4998
4999
5000
5001
5002
5003
5004
5005
5006
5007
5008
5009
5010
5011
5012
5013
5014
5015
5016
5017
5018
5019
5020
5021
5022
5023
5024
5025
5026
5027
5028
5029
5030
5031
5032
5033
5034
5035
5036
5037
5038
5039
5040
5041
5042
5043
5044
5045
5046
5047
5048
5049
5050
5051
5052
5053
5054
5055
5056
5057
5058
5059
5060
5061
5062
5063
5064
5065
5066
5067
5068
5069
5070
5071
5072
5073
5074
5075
5076
5077
5078
5079
5080
5081
5082
5083
5084
5085
5086
5087
5088
5089
5090
5091
5092
5093
5094
5095
5096
5097
5098
5099
5100
5101
5102
5103
5104
5105
5106
5107
5108
5109
5110
5111
5112
5113
5114
5115
5116
5117
5118
5119
5120
5121
5122
5123
5124
5125
5126
5127
5128
5129
5130
5131
5132
5133
5134
5135
5136
5137
5138
5139
5140
5141
5142
5143
5144
5145
5146
5147
5148
5149
5150
5151
5152
5153
5154
5155
5156
5157
5158
5159
5160
5161
5162
5163
5164
5165
5166
5167
5168
5169
5170
5171
5172
5173
5174
5175
5176
5177
5178
5179
5180
5181
5182
5183
5184
5185
5186
5187
5188
5189
5190
5191
5192
5193
5194
5195
5196
5197
5198
5199
5200
5201
5202
5203
5204
5205
5206
5207
5208
5209
5210
5211
5212
5213
5214
5215
5216
5217
5218
5219
5220
5221
5222
5223
5224
5225
5226
5227
5228
5229
5230
5231
5232
5233
5234
5235
5236
5237
5238
5239
5240
5241
5242
5243
5244
5245
5246
5247
5248
5249
5250
5251
5252
5253
5254
5255
5256
5257
5258
5259
5260
5261
5262
5263
5264
5265
5266
5267
5268
5269
5270
5271
5272
5273
5274
5275
5276
5277
5278
5279
5280
5281
5282
5283
5284
5285
5286
5287
5288
5289
5290
5291
5292
5293
5294
5295
5296
5297
5298
5299
5300
5301
5302
5303
5304
5305
5306
5307
5308
5309
5310
5311
5312
5313
5314
5315
5316
5317
5318
5319
5320
5321
5322
5323
5324
5325
5326
5327
5328
5329
5330
5331
5332
5333
5334
5335
5336
5337
5338
5339
5340
5341
5342
5343
5344
5345
5346
5347
5348
5349
5350
5351
5352
5353
5354
5355
5356
5357
5358
5359
5360
5361
5362
5363
5364
5365
5366
5367
5368
5369
5370
5371
5372
5373
5374
5375
5376
5377
5378
5379
5380
5381
5382
5383
5384
5385
5386
5387
5388
5389
5390
5391
5392
5393
5394
5395
5396
5397
5398
5399
5400
5401
5402
5403
5404
5405
5406
5407
5408
5409
5410
5411
5412
5413
5414
5415
5416
5417
5418
5419
5420
5421
5422
5423
5424
5425
5426
5427
5428
5429
5430
5431
5432
5433
5434
5435
5436
5437
5438
5439
5440
5441
5442
5443
5444
5445
5446
5447
5448
5449
5450
5451
5452
5453
5454
5455
5456
5457
5458
5459
5460
5461
5462
5463
5464
5465
5466
5467
5468
5469
5470
5471
5472
5473
5474
5475
5476
5477
5478
5479
5480
5481
5482
5483
5484
5485
5486
5487
5488
5489
5490
5491
5492
5493
5494
5495
5496
5497
5498
5499
5500
5501
5502
5503
5504
5505
5506
5507
5508
5509
5510
5511
5512
5513
5514
5515
5516
5517
5518
5519
5520
5521
5522
5523
5524
5525
5526
5527
5528
5529
5530
5531
5532
5533
5534
5535
5536
5537
5538
5539
5540
5541
5542
5543
5544
5545
5546
5547
5548
5549
5550
5551
5552
5553
5554
5555
5556
5557
5558
5559
5560
5561
5562
5563
5564
5565
5566
5567
5568
5569
5570
5571
5572
5573
5574
5575
5576
5577
5578
5579
5580
5581
5582
5583
5584
5585
5586
5587
5588
5589
5590
5591
5592
5593
5594
5595
5596
5597
5598
5599
5600
5601
5602
5603
5604
5605
5606
5607
5608
5609
5610
5611
5612
5613
5614
5615
5616
5617
5618
5619
5620
5621
5622
5623
5624
5625
5626
5627
5628
5629
5630
5631
5632
5633
5634
5635
5636
5637
5638
5639
5640
5641
5642
5643
5644
5645
5646
5647
5648
5649
5650
5651
5652
5653
5654
5655
5656
5657
5658
5659
5660
5661
5662
5663
5664
5665
5666
5667
5668
5669
5670
5671
5672
5673
5674
5675
5676
5677
5678
5679
5680
5681
5682
5683
5684
5685
5686
5687
5688
5689
5690
5691
5692
5693
5694
5695
5696
5697
5698
5699
5700
5701
5702
5703
5704
5705
5706
5707
5708
5709
5710
5711
5712
5713
5714
5715
5716
5717
5718
5719
5720
5721
5722
5723
5724
5725
5726
5727
def compile_module(
    spec_dir: Path,
    output_dir: Path,
    compression: str = "zstd",
    resolve_with_ensembl: bool = True,
    ensembl_cache: Path | None = None,
    compiled_by: str | None = None,
    ensembl_reference: str | None = None,
    log_files: list[Path] | None = None,
    provenance_file: Path | None = None,
    logo_file: Path | None = None,
    readme_file: Path | None = None,
    authority_keys: Iterable[str] | None = None,
    strict: bool = False,
    ba1_threshold: float = BA1_ALLELE_FREQUENCY_THRESHOLD,
) -> CompilationResult:
    """Compile a module spec directory into parquet files plus a `manifest.json`.

    Args:
        spec_dir: Path to the module spec directory.
        output_dir: Directory for output parquet files + manifest.json.
        compression: Parquet compression codec.
        resolve_with_ensembl: Master switch for resolution — of **every** kind, despite the name.
            With a `resolution.csv` present it drives the preferred, source-independent table path;
            `ensembl_cache` is the deprecated fallback. Setting it False with a `resolution.csv`
            present disables that table too, and warns, because the compile then succeeds with no
            resolved position on any row.
        ensembl_cache: **Deprecated (removed at 1.0).** Path to a prebuilt Ensembl DuckDB or parquet
            cache dir. In-compiler DuckDB resolution has moved to `just-dna-enricher`; when given, this
            emits a `DeprecationWarning` and routes to the enricher (which must be installed). Prefer
            producing a `resolution.csv` (`just-dna-enricher enrich`) — the compiler then resolves with
            no reference and no network.
        compiled_by: Provenance tag for the manifest (the marketplace passes "marketplace-server";
            a local compile leaves it None, so downloaders treat it as untrusted).
        ensembl_reference: Pinned reference id recorded in the manifest for reproducibility.
        log_files: Explicit run/provenance log files to record. If None, auto-discovers a top-level
            `*.log` plus per-role files under `spec_dir/logs/`. Logs are optional.
        provenance_file: Explicit structured-provenance document. If None, auto-discovers
            `spec_dir/provenance.json`. Optional; summarized into `manifest.provenance`.
        logo_file: Explicit module logo image. If None, auto-discovers `spec_dir/logo.{png,jpg,jpeg}`.
            Optional; hashed into `manifest.logo`, kept out of `artifact.digest`.
        readme_file: Explicit module readme. If None, auto-discovers the first of
            `manifest.README_CANDIDATES` (`README.md` first) beside the spec. Optional; hashed into
            `manifest.readme`, kept out of `artifact.digest` and `content_signature` — prose about the
            module is not part of its identity, so fixing a caveat is a patch (S25).
        authority_keys: Inject-only set of consumer/registry-owned identity keys to strip from the
            authored `module:` block before validation (e.g. `just_dna_format.normalize.
            IDENTITY_AUTHORITY_KEYS`). None strips nothing.
        strict: All-or-nothing compile. When True, fail (rather than emit a partial artifact) if any
            variant still lacks a resolved genomic position (`chrom`+`start`) after resolution — an
            unresolved position means the injected reference was incomplete/absent and the parquet
            bytes (hence `artifact.digest`) would not be reproducible. Default False keeps the
            best-effort behavior (positions left unset, surfaced as warnings).
        ba1_threshold: Allele frequency above which a `pathogenic` variant draws the ACMG BA1 warning
            (`_check_ba1_lint`). Defaults to ACMG's 5%. Raise it for a module curating a common
            recessive carrier allele, where the default fires on correct data. Warning-only in both
            modes, so this tunes noise, never whether the compile succeeds.
    """
    spec_dir = Path(spec_dir)
    output_dir = Path(output_dir)

    # `strict` is deliberately NOT passed (the pre-flight runs in best_effort whatever this compile's
    # mode, which is why every mode-ladder check re-runs below); `resolve_with_ensembl` is, because it
    # is not a severity — it decides whether the injected table is consulted at all, and a pre-flight
    # answering that differently would seed `all_warnings` with a finding this compile contradicts.
    validation, validation_findings = _validate_spec(
        spec_dir, authority_keys, resolve_with_ensembl=resolve_with_ensembl
    )
    if not validation.valid:
        # The classified list, not `validation.warnings` — a failed compile publishes the same
        # readable channel a successful one does, and the model's copy has lost its codes.
        return CompilationResult(success=False, errors=validation.errors, warnings=validation_findings)

    config, _, _ = _load_yaml(spec_dir / "module_spec.yaml", authority_keys)
    assert config is not None
    module_name = config.module.name

    # A module composes from optional table kinds (RM2): load whatever is present.
    variants: list[VariantRow] = []
    if (spec_dir / "variants.csv").exists():
        variants, _, _ = _load_csv_rows(
            spec_dir / "variants.csv",
            VariantRow,
            "variants.csv",
            genome_build=config.genome_build,
        )
        # Same re-stamp as in `validate_spec`; this function re-loads its own rows, so the fix has to
        # happen on both copies or the artifact would carry the un-corrected keys.
        _restamp_for_build(variants, config.genome_build)
    studies: list[StudyRow] = []
    if (spec_dir / "studies.csv").exists():
        studies, _, _ = _load_csv_rows(
            spec_dir / "studies.csv", StudyRow, "studies.csv", genome_build=config.genome_build
        )
    # The 0.4 table kinds, loaded here rather than at materialization time. Two reasons, and the
    # second is the load-bearing one: the symbolic-allele check below has to reach them, and a check
    # that can refuse must do so **before** `output_dir.mkdir()` — the placement the licence gate
    # already has, for the same reason. (It also removes a second load of every table kind.)
    kind_rows: dict[str, list[Any]] = {}
    for csv_name, _parquet_name, model in _TABLE_KINDS:
        if (spec_dir / csv_name).exists():
            kind_rows[csv_name], _, _ = _load_csv_rows(
                spec_dir / csv_name, model, csv_name, genome_build=config.genome_build
            )

    # Seeded from the classified list rather than from `validation.warnings`, whose members pydantic
    # has already flattened to plain strings. The two lists carry identical text, so every `if w not
    # in all_warnings` de-duplication below still compares exactly what it compared before.
    all_warnings = list(validation_findings)

    # Symbolic/structural alleles the module cannot apply (RM5). Runs before anything is written, and
    # before resolution, because it decides which rows exist: under `best_effort` a row stating an
    # unusable rule is **dropped** (the warning says so — `reverse` will not re-emit it), under
    # `strict` the compile refuses. Re-run here rather than trusted from `validate_spec`, which ran in
    # `best_effort` whatever this compile's mode; the warnings it already produced are de-duplicated
    # on the message, the way ploidy's and allele-membership's are.
    symbolic_errors, symbolic_warnings, symbolic_drops = _check_symbolic_alleles(
        {"variants.csv": variants, **kind_rows}, strict=strict
    )
    all_warnings.extend(w for w in symbolic_warnings if w not in all_warnings)
    if symbolic_errors:
        return CompilationResult(success=False, errors=symbolic_errors, warnings=all_warnings)
    for table, drop_rows in symbolic_drops.items():
        if table == "variants.csv":
            variants = _apply_symbolic_drops(variants, drop_rows)
        else:
            kind_rows[table] = _apply_symbolic_drops(kind_rows[table], drop_rows)
    if symbolic_drops:
        # Re-derive the stats over what survived. `validate_spec` computed them from the full set, and
        # `weights_rows` counts the parquet, so leaving them would publish a `variant_count` higher
        # than the artifact holds (the RM44 class: a manifest number a catalog keys on and cannot
        # check). **After the whole loop, not inside its `variants.csv` branch** — since S57 put the
        # gene facets on every authored kind, a drop that empties the last row naming a gene in a
        # *kind* table has to move `genes` too, and the old placement could not see one.
        validation.stats.update(module_stats(variants, kind_rows))

    # The source-independent resolution table (0.5), if authored/produced beside the spec. When
    # present it is the *preferred* resolution path: the compiler consumes already-resolved facts and
    # owns no source convention (Ensembl/DuckDB/provisioning) — the strict inject-only end state
    # (Principle 2). An injected `ensembl_cache` (the DuckDB path) is the superseded fallback (P3).
    # The authored overlay (RM124), loaded before the first derived table it lies on. `validate_spec`
    # above already reported every finding it has — this pass re-loads its own copy, the way it
    # re-loads `variants.csv`, because the rows themselves are needed to build the artifact.
    overrides, overlay_errors, overlay_warnings = load_overlay(spec_dir)
    if overlay_errors:
        return CompilationResult(success=False, errors=overlay_errors, warnings=all_warnings)
    # De-duplicated on the message like every other check that runs on both sides: the pre-flight
    # already emitted these over the same bytes. Extended rather than discarded so a finding this
    # loader gains later cannot go missing from a compile that never runs `validate` separately.
    all_warnings.extend(w for w in overlay_warnings if w not in all_warnings)
    overlaid: set[str] = set()
    #: RM137, the compile's half of the same stash: unmatched `update` targets for the two tables a
    #: reverse rebuilds from something narrower, split below once `studies` and the citing tables are
    #: in scope. Same reason as in `validate_spec`, and the same function does the splitting.
    compile_deferred: dict[str, list[tuple[tuple[str, str], bool]]] = {}

    resolution_rows: list[ResolutionRow] = []
    resolution_table: dict[str, list[ResolutionRow]] = {}
    resolution_sources: list[str] = []
    resolution_sig: str | None = None
    # 0/0 for a module with no resolution table: no allele identities were attempted, which is a
    # different statement from "none were achieved" and is what the manifest should carry.
    vrs_alleles = vrs_identified = 0
    resolution_path, res_spelling_warnings, res_spelling_errors = _locate_sidecar(spec_dir, "resolution.csv")
    if res_spelling_errors:
        return CompilationResult(success=False, errors=res_spelling_errors, warnings=all_warnings)
    all_warnings.extend(w for w in res_spelling_warnings if w not in all_warnings)
    if resolution_path is not None:
        resolution_rows, res_errors, _ = _load_csv_rows(resolution_path, ResolutionRow, resolution_path.name)
        if res_errors:
            return CompilationResult(success=False, errors=res_errors, warnings=all_warnings)
        # The overlay first, so everything below — the membership table, the VRS pass, the coverage
        # count and `resolution_signature` — reads what the module asserts rather than what the last
        # enrichment happened to write (RM124).
        overlaid.add("resolution.csv")
        # Deferred and stashed on the PRE-overlay rows (RM137) — see `_classify_deferred_overlay_updates`.
        compile_deferred["resolution.csv"] = update_targets("resolution.csv", resolution_rows, overrides)
        resolution_rows, apply_errors, apply_warnings = apply_overrides(
            "resolution.csv", resolution_rows, overrides, defer_unmatched=True
        )
        if apply_errors:
            return CompilationResult(success=False, errors=apply_errors, warnings=all_warnings)
        all_warnings.extend(w for w in apply_warnings if w not in all_warnings)
        for row in resolution_rows:
            resolution_table.setdefault(row.variant_key, []).append(row)
        # Content-addressed identities are checkable against themselves — do it before anything is
        # written, so a tampered id never reaches an artifact. Dep-free (stdlib), see `_verify_vrs_ids`.
        # De-duplicated on the message, the same way ploidy and allele-membership are: `validate_spec`
        # ran this pass over the same injected rows and `all_warnings` was seeded from its result, so
        # every finding is already in the list once. Harmless while these were strict-mode *errors*
        # (which return early) and merely untidy for the coverage line (one per module); now that an
        # unverifiable allele warns in every mode it is one duplicated line per allele, and
        # `pathogenic_clinvar` alone would print 185 of them twice.
        vrs_errors, vrs_warnings = _verify_vrs_ids(resolution_rows)
        all_warnings.extend(w for w in vrs_warnings if w not in all_warnings)
        if vrs_errors:
            return CompilationResult(success=False, errors=vrs_errors, warnings=all_warnings)
        all_warnings.extend(w for w in _vrs_coverage_warnings(resolution_rows) if w not in all_warnings)
        vrs_alleles, vrs_identified, _gaps = _vrs_coverage(resolution_rows)
        # The table's identity is stamped **here**, where the table was read, rather than inside the
        # `variants`-gated resolution block below — which is where it used to sit, and which meant a
        # module with no `variants.csv` published `resolution_signature: null` while carrying a
        # perfectly good `resolution.csv` beside its spec. Four of the eleven reference examples are
        # exactly that shape, so a consumer holding a PGx manifest could not tell a module resolved
        # from one that never was, on the one field that answers it.
        #
        # It was gated that way for a real reason that RM43 removed: until 0.6 `reverse_module` rebuilt
        # this table from `weights.parquet` alone, so a table-only module round-tripped to a spec with
        # no `resolution.csv` and stamping the signature would have published an identity the next
        # compile could not reproduce (P7). Reverse now rebuilds it from the positional parquets too.
        #
        # **The residual, measured rather than assumed.** The signature reproduces exactly when the
        # artifact consumed the whole table — all eleven reference examples, and every clean row of
        # `test_resolution_matrix.py`. It does **not** reproduce when the table says more than the
        # module uses: a row about a variant the module does not carry, an unplaced one-to-many, a
        # coordinate the author overrode. Reverse rebuilds the table from the artifact, so a fact the
        # artifact never held has nowhere to come back from — the same structural reason
        # `rsid_alternates` is unrecoverable. That is **not new and not positional**: a plain SNP
        # module with one unused injected row has always behaved this way, and it is pinned in the
        # matrix now (`Case.table_says_more`) because stamping here is what first made it visible.
        # `artifact.digest` is reproducible throughout, which is why `strict` has nothing to refuse.
        #
        # Still gated on `resolve_with_ensembl`: under `--no-resolve` the table is deliberately not
        # consulted (the warning above says so), and a signature naming a table that shaped nothing
        # would claim the artifact was built from it.
        #
        # And gated on the table having ROWS, not merely existing. A header-only `resolution.csv`
        # hashes to the empty-set digest, which is a perfectly valid signature of nothing — publishing
        # it costs `resolution_signature is not None` its meaning ("this module was resolved"). The
        # deprecated `ensembl_cache` branch makes it worse than cosmetic: an empty `resolution_table`
        # is falsy there, so the DuckDB path runs and the manifest would name an injected table that
        # shaped none of the bytes — the same false claim the `--no-resolve` warning refuses to make.
        if resolve_with_ensembl and resolution_rows:
            resolution_sources = sorted({row.source for row in resolution_rows if row.source})
            resolution_sig = _resolution_signature(resolution_rows)

    # Do the alleles the module *states* exist at the loci it points at? Runs here, on the AUTHORED
    # rows, because resolution may expand one rsid into several loci that share this genotype — after
    # that expansion the check reports the siblings it was never about. See `_check_allele_membership`.
    allele_errors, allele_warnings = _check_allele_membership(variants, resolution_table, strict=strict)
    # De-duplicated on the message, the same way `_check_contig_ploidy` below is and for the same
    # reason: `compile_module` runs `validate_spec` first, in **best_effort** regardless of this
    # compile's mode, so a check living in both places emits its warning twice. Re-running it here is
    # not redundant — it is how a mode ladder reaches its real severity, since the inner pass cannot
    # know `strict`. What is redundant is printing the identical sentence a second time.
    all_warnings.extend(w for w in allele_warnings if w not in all_warnings)
    if allele_errors:
        return CompilationResult(success=False, errors=allele_errors, warnings=all_warnings)

    # RM91's study-side half, de-duplicated on the message for the same reason as the block above.
    study_allele_errors, study_allele_warnings = _check_study_effect_alleles(
        studies, resolution_table, strict=strict
    )
    all_warnings.extend(w for w in study_allele_warnings if w not in all_warnings)
    if study_allele_errors:
        return CompilationResult(success=False, errors=study_allele_errors, warnings=all_warnings)

    # De-duplicated on the message, like the two blocks above (RM94). This one was the exception, and
    # the reason it survived is worth keeping: `@no-rerun-with-counts` guards against a re-run whose
    # message embeds a count, because the two copies then disagree and the manifest publishes two
    # numbers. This message carries no count, so the copies agreed and the field was merely redundant
    # rather than self-contradicting -- an untidiness nobody was looking for. The wider rule the
    # neighbours already follow is the one to take from it: a check that runs on both sides dedupes on
    # the message, and re-running it is the normal case rather than the exception.
    #
    # The re-run itself earns its place and is not what was removed. `compile_module` runs
    # `validate_spec` in best_effort regardless of this compile's mode, so this second pass in the
    # caller's mode is the only thing that lets `--strict` escalate the warning into a refusal.
    p_value_errors, p_value_warnings = _check_p_value_num(studies, strict=strict)
    all_warnings.extend(w for w in p_value_warnings if w not in all_warnings)
    if p_value_errors:
        return CompilationResult(success=False, errors=p_value_errors, warnings=all_warnings)

    resolution_mode: str | None = None
    # `None` until the injected-table path runs and reports them (S33) — see the assignment below for
    # why no other branch can.
    expanded_keys: int | None = None
    expanded_rows: int | None = None
    # The flag reads as "do not use Ensembl", and since 0.5 made the compiler inject-only that is
    # exactly what a consumer migrating to `resolution.csv` expects it to mean. It is actually the
    # master switch for resolution *of any kind*, so turning it off with a complete, correct table
    # sitting beside the spec compiles **successfully** with `chrom=None` on every weight row — rows
    # that can never match a VCF. A silent success is the worst shape a mistake can take, so the
    # combination that cannot be intended says so. (Renaming it is a 1.0 conversation: the parameter
    # is part of a published signature.)
    if not resolve_with_ensembl and resolution_table:
        # The unread row count is in the message because a warning that quantifies over a table should
        # publish the size of what it skipped — the same reason `vrs_alleles` ships beside
        # `vrs_alleles_identified`. Rows, not keys: a one-to-many rsid contributes several.
        unread = sum(len(rows) for rows in resolution_table.values())
        all_warnings.append(
            CodedWarning(
                "resolution_disabled",
                f"--no-resolve (resolve_with_ensembl=False) switches off resolution entirely, including "
                f"the injected resolution.csv beside this spec ({unread} row(s), covering "
                f"{len(resolution_table)} variant key(s)), which was not read — every variant will compile "
                f"with no chrom/start and match no VCF. The flag names Ensembl but is the master switch; "
                f"drop it to use the injected table. There is no flag for 'do not reach the network' "
                f"because the compiler never does (CONSTITUTION P2) — omitting this one is that request.",
            )
        )
    if resolve_with_ensembl and variants:
        resolution_mode = "strict" if strict else "best_effort"
        resolve_warnings: list[str] = []
        resolve_strict_errors: list[str] = []
        if resolution_table:
            outcome = resolve_from_table(variants, resolution_table, genome_build=config.genome_build)
            variants = outcome.variants
            resolve_warnings = outcome.warnings
            resolve_strict_errors = outcome.strict_errors
            # S33. Only this branch can answer it: the deprecated `ensembl_cache` path returns a bare
            # (variants, warnings) pair and is removed at 1.0, so its modules keep `None` — "not
            # established", never "no expansion". Assigned here rather than derived from `variants`
            # further down, because after the expansion an expanded row is indistinguishable from an
            # ordinary coordinate-keyed one; the fact only exists while the loop is running.
            expanded_keys = outcome.expanded_keys
            expanded_rows = outcome.expanded_rows
            if outcome.errors:
                # Fatal in both modes — today only a curator-recorded `withdrawn` rsID. See
                # `ResolutionOutcome`.
                return CompilationResult(
                    success=False,
                    errors=[f"resolution: {e}" for e in outcome.errors],
                    warnings=all_warnings + resolve_warnings,
                )
            # `resolution_sources` / `resolution_sig` are stamped where the table is READ, several
            # blocks up — not here. This branch is about applying the table to `variants.csv`, and
            # tying the table's identity to that application is what left every table-only module
            # unstamped.
        elif ensembl_cache is not None:
            # DEPRECATED (removed at 1.0): the in-compiler DuckDB-reference path. Resolution now belongs
            # to the source-independent `resolution.csv` (produce it with `just-dna-enricher enrich`); the
            # compiler owns no source/DuckDB logic. This surface is kept working by routing to the
            # enricher — additive-within-a-major binds the wire/artifact *contract*, not this internal
            # call, so the legacy path can retire at the next major. Guarded optional import (CLAUDE.md).
            warnings.warn(
                "compile_module(ensembl_cache=...) / in-compiler DuckDB resolution is deprecated and "
                "will be removed at 1.0. Produce a resolution.csv (e.g. `just-dna-enricher enrich`) and "
                "the compiler consumes it with no reference and no network.",
                DeprecationWarning,
                stacklevel=2,
            )
            try:
                from just_dna_enricher.resolver import resolve_variants as _legacy_resolve
            except ImportError:
                return CompilationResult(
                    success=False,
                    errors=[
                        "ensembl_cache resolution now lives in just-dna-enricher (the network tier). "
                        "Install just-dna-enricher, or precompute resolution.csv and recompile without "
                        "ensembl_cache."
                    ],
                    warnings=all_warnings,
                )
            variants, resolve_warnings = _legacy_resolve(
                variants, ensembl_cache, genome_build=config.genome_build
            )
        elif any(v.chrom is None or v.start is None for v in variants):
            # Nothing injected: the compiler no longer auto-discovers or fetches a reference (P2,
            # tightened in 0.5). Variants lacking a position are left unresolved with a pointer.
            resolve_warnings = [
                CodedWarning(
                    "resolution_not_injected",
                    resolution_not_injected(
                        missing="No resolution.csv and no ensembl_cache injected",
                        remedy="Produce a resolution.csv with just-dna-enricher.",
                    ),
                )
            ]
        # De-duplicated on the message, the `_check_contig_ploidy` idiom: since S76 the pre-flight
        # emits the `rsid_unresolved` sentence for the same subjects, and `compile_module` runs that
        # pre-flight whatever its own mode, so appending blind published every such finding twice —
        # and `warnings_summary` counted 24 for 12 subjects. Safe here for the reason the rule
        # requires: no message resolution re-derives embeds a count, so two passes over one subject
        # produce the identical sentence rather than two that differ by a number.
        all_warnings.extend(w for w in resolve_warnings if w not in all_warnings)
        # The round-trip contract. `strict` promises a *reproducible* artifact, and these are the
        # conditions under which `compile → reverse → compile` cannot reproduce the resolution table
        # it started from (a dropped locus, an authored coordinate contradicting the table), plus the
        # one that is reproducible but rests on a guessed label (`ambiguous`). `best_effort` already
        # carries them as warnings above. See COMPILER.md § Resolution.
        if strict and resolve_strict_errors:
            return CompilationResult(
                success=False,
                errors=[f"strict resolution: {e}" for e in resolve_strict_errors],
                warnings=all_warnings,
            )
        # Resolution is an enrichment that can *change identity*: filling a coordinate or expanding a
        # one-to-many rsid into coord-keyed rows may collide with an already-authored row. validate_spec
        # ran on the pre-resolution set, so re-run the identity checks on the resolved set — a duplicate
        # (variant_key, genotype) or an inconsistent position must fail the compile, not silently land in
        # weights.parquet. Only errors are taken (warnings were already surfaced pre-resolution).
        post_errors, _ = _cross_validate_variants(variants)
        if post_errors:
            return CompilationResult(
                success=False,
                errors=[f"post-resolution: {e}" for e in post_errors],
                warnings=all_warnings,
            )

    # The non-diploid guardrail runs **here**, once, because this is the first point at which `chrom`
    # is final for every row — resolution has either run and filled it or was skipped. Inside
    # `_cross_validate_variants` it was emitted from the pre-resolution pass only (the post-resolution
    # call takes errors alone), so an rsID-authored MT or Y row — the shape every drafting provider
    # writes — was never checked at all.
    all_warnings.extend(
        w for w in _check_contig_ploidy(variants, config.genome_build) if w not in all_warnings
    )

    # Wrong-build coordinates, re-run **here** for the same reason ploidy is: this is the first point
    # at which `chrom`/`start` are final. `validate_spec` sees an rsid-only row with no coordinate at
    # all; resolution has since filled one, and the deprecated `ensembl_cache` path fills it from a
    # reference `validate_spec` never opened. Unlike ploidy this needs no de-duplication: it returns
    # errors, `compile_module` returns early on any validation error, so anything reaching here is a
    # coordinate the pre-flight could not have seen.
    build_errors = _check_build_coordinates(
        [_CoordinateTable("variants.csv", config.genome_build, True, variants)]
    )
    if build_errors:
        return CompilationResult(
            success=False,
            errors=[f"post-resolution: {e}" for e in build_errors],
            warnings=all_warnings,
        )

    # Outcome axis (orthogonal to the requested `resolution_mode` policy, Principle 5): did every
    # in-scope variant resolve to a genomic position? Vacuously true for a table-kind-only module —
    # which is why the denominator travels beside it into the manifest (RM44). Both come from the same
    # list, so the flag can never be published without the count that says what it quantified over.
    fully_resolved = all(v.chrom is not None and v.start is not None for v in variants)
    resolution_subjects = len(variants)

    # Strict (all-or-nothing): refuse to write a partial artifact. A variant still missing its
    # genomic position after resolution means the injected reference was incomplete or absent, so the
    # coordinate-anchored parquet bytes (and `artifact.digest`) would not be reproducible — the
    # failure mode behind "local hash differs from published". Best-effort (strict=False) leaves such
    # rows unset with a warning instead. Scope is the SNP-core VariantRow; the 0.4 table kinds carry
    # no positions.
    if strict and variants:
        unresolved = sorted(v.rsid or v.variant_key for v in variants if v.chrom is None or v.start is None)
        if unresolved:
            return CompilationResult(
                success=False,
                errors=[
                    f"strict compile: {len(unresolved)} variant(s) have unresolved genomic "
                    f"positions after resolution: {unresolved}. A partial artifact would not be "
                    f"byte-reproducible; inject a complete Ensembl reference (ensembl_cache=) or "
                    f"compile without strict."
                ],
                warnings=all_warnings,
            )

    # The wrong-build gate (S78, RM143), here for the same placement reason as the licence gate below
    # and one line ahead of it: a refusal must leave nothing written. It reads the attestation the
    # enricher wrote and acts on one recorded finding — see `build_disagreement_error` for why that
    # one and no other, and why this does not move the strict line. The block is re-read rather than
    # carried down from the pre-flight because a stale attestation is dropped by the reader, and the
    # gate must see what the reader saw.
    if strict:
        build_error = build_disagreement_error(_verification_block(spec_dir)[0])
        if build_error is not None:
            return CompilationResult(success=False, errors=[build_error], warnings=all_warnings)

    # Licensing gate. Loaded here rather than with the other fact tables because those are read
    # *after* `output_dir.mkdir()`, and a refusal must leave nothing written — this is the last point
    # at which that is still true. Purely computation over injected data: the compiler holds no
    # source→licence map (Principle 2 — it owns no source convention) and only reads what the
    # enricher recorded.
    sources_path, gate_spelling_warnings, gate_spelling_errors = _locate_sidecar(spec_dir, SOURCES_CSV)
    if gate_spelling_errors:
        return CompilationResult(success=False, errors=gate_spelling_errors, warnings=all_warnings)
    all_warnings.extend(w for w in gate_spelling_warnings if w not in all_warnings)
    if sources_path is not None:
        gate_rows, gate_load_errors, _ = _load_csv_rows(sources_path, SourceRow, sources_path.name)
        if gate_load_errors:
            return CompilationResult(success=False, errors=gate_load_errors, warnings=all_warnings)
        gate_errors = _check_license_gate(gate_rows)
        if gate_errors:
            return CompilationResult(success=False, errors=gate_errors, warnings=all_warnings)

    output_dir.mkdir(parents=True, exist_ok=True)

    # SNP core: weights/annotations only when the module actually has variants.
    weights_df = _build_weights(variants, config) if variants else None
    annotations_df = _build_annotations(variants, module_name) if variants else None
    studies_df = _build_studies(studies, module_name) if studies else None
    # Names taken from `SNP_CORE_PARQUETS` rather than spelled again: this is the site the binding was
    # extracted from, so a literal here would be the copy the constant exists to remove.
    _weights_name, _annotations_name = SNP_CORE_PARQUETS["variants.csv"]
    (_studies_name,) = SNP_CORE_PARQUETS["studies.csv"]
    if weights_df is not None:
        weights_df.write_parquet(output_dir / _weights_name, compression=compression)
    if annotations_df is not None:
        annotations_df.write_parquet(output_dir / _annotations_name, compression=compression)
    if studies_df is not None:
        studies_df.write_parquet(output_dir / _studies_name, compression=compression)

    # The positional fill (RM43), on the rows that survive. It runs **after** the symbolic-allele
    # ladder (RM5) and before `_build_table`, and that order is the only correct one: the drop decides
    # which rows exist, so filling first would resolve a coordinate onto a row about to be discarded
    # and — worse — leave the joinability report below counting rows the artifact does not contain.
    # The two never disagree because the drop happens up at load time, long before the resolution
    # table is even read; this comment is here so the distance between them does not read as accident.
    # Gated on `resolve_with_ensembl` like every other resolution, so `--no-resolve` consults nothing.
    fill_warnings, fill_applied = _apply_positional_resolution(
        kind_rows, resolution_table, config.genome_build, resolve=resolve_with_ensembl
    )
    all_warnings.extend(w for w in fill_warnings if w not in all_warnings)

    # 0.4 table kinds (RM1): materialize each present CSV via the generic materializer. The rows were
    # loaded above, before `mkdir`, so the symbolic-allele check could refuse without leaving a
    # half-written directory behind — and so a row it dropped is absent from the parquet here.
    table_rows: dict[str, int] = {}
    for csv_name, parquet_name, model in _TABLE_KINDS:
        rows = kind_rows.get(csv_name)
        if rows is None:
            continue
        table_df = _build_table(rows, model, module_name)
        table_df.write_parquet(output_dir / parquet_name, compression=compression)
        table_rows[parquet_name] = table_df.height

    # Re-run here so the finding reaches a caller who compiles without validating first, and
    # de-duplicated on the message the way `_check_contig_ploidy` is: `compile_module` runs
    # `validate_spec` itself, so a check living in both places otherwise prints its sentence twice.
    all_warnings.extend(
        w
        for w in _check_positional_joinability(
            kind_rows, resolution_table, config.genome_build, fill_applied=fill_applied
        )
        if w not in all_warnings
    )
    # The same facts as counts rather than as a sentence, for the catalog that was reading the
    # sentence (S31). Computed here, beside the check, so the two cannot describe different row sets.
    positional_rows, positional_rows_placed = positional_placement(kind_rows)
    all_warnings.extend(w for w in _check_binning_grounding(kind_rows, studies) if w not in all_warnings)
    # `kind_rows` is freshly loaded and never resolved, so re-running the binning check here produces
    # the identical sentence and the message-dedup above does its job.
    all_warnings.extend(w for w in _check_measure_shape(kind_rows) if w not in all_warnings)
    all_warnings.extend(w for w in _check_binning_deprecations(kind_rows) if w not in all_warnings)

    # **THREE checks of this round are deliberately NOT re-run here, and that is the fix rather than an
    # omission.** `_check_missing_allele_marker`, `_check_quality_inversion` and `_check_vcf_pointers`
    # all read authored cells (`alts`, `requires_callable` + `quality_from`, the pointer columns) that
    # resolution never fills, so `validate_spec`'s pass — which runs before any expansion — already has
    # the right answer, and it reaches this list through `all_warnings = list(validation.warnings)`
    # above. Running any of them again on `variants` would count the *expanded* rows: by this point
    # `variants` is `outcome.variants`, one row per resolved locus, so a one-to-many rsid becomes N rows
    # carrying one authored genotype and the same finding is reported with a different count and
    # different example keys. The de-duplication above keys on the **message**, so two sentences
    # differing only in their count can never collapse, and both are published into
    # `manifest.compilation.warnings` — a surface RM44 established that consumers parse. Measured twice,
    # independently, on the way in: an rsid-only `requires_callable` row over a two-locus
    # `resolution.csv` emitted "1 row(s) …" beside "2 row(s) …", and the pointer check emitted 328
    # beside 337 on `pathogenic_clinvar`.
    #
    # Nothing is lost by staying behind resolution: all three are warning-only in both modes, so there
    # is no severity for a re-run to recover — which is the whole reason the *mode-ladder* checks re-run
    # at all. This is the mirror of the `_check_contig_ploidy` lesson rather than a contradiction of it:
    # that warning had to **move** here because resolution fills its input; these three must stay behind
    # it because resolution fills nothing they read. The rule that covers both: re-run a check after
    # resolution exactly when resolution changes its input, and never when the message embeds a count.

    # 0.5 derived-fact sidecars: materialize each present CSV, and cross-check it against what the
    # module actually contains. A row describing something the module never mentions is a warning, not
    # an error — an over-broad sidecar is harmless (a stale gene left in after a variant was removed),
    # while failing the compile over it would punish the author for the enricher's generosity.
    # One branch per sidecar, keyed by model. A two-way `if/else` was fine for two tables and stops
    # being readable at three, so each entry states its own checks and its own builder; the loop below
    # stays generic. Errors are fatal, warnings accumulate — the same contract the SNP core uses.
    def _frequency_checks(rows: list) -> tuple[list[str], list[str]]:
        # De-duplicated on the message, the `_literature_checks` idiom eleven lines below and the
        # `_check_contig_ploidy` one before it: `compile_module` runs `validate_spec`, which has run
        # `_check_frequency_arithmetic` over these same injected rows since RM93 added it for parity,
        # so an unfiltered re-run publishes the finding **twice** in `manifest.compilation.warnings`
        # (RM44) and every consumer counting warnings overstates what is wrong with the module.
        # Re-running is the normal case and is not the bug; not filtering is (`@no-rerun-with-counts`).
        # Only the warnings need it — validate's errors abort the compile before this closure runs.
        errors, warns = _check_frequency_arithmetic(rows)
        warns = [w for w in warns if w not in all_warnings]
        warns.extend(_cross_check_frequencies(rows, variants))
        warns.extend(_check_ba1_lint(rows, variants, threshold=ba1_threshold))
        return errors, warns

    def _gene_metrics_checks(rows: list) -> tuple[list[str], list[str]]:
        # Both have run in the pre-flight since RM211, so both arrive twice. The **extend site**
        # below de-duplicates every fact-table check on the message — `@first-fact-check-on-both-
        # sides`, dedupe where the results are collected rather than in each closure — which is why
        # this one carries no filter of its own and `_gene_validity_checks` below carries none
        # either. Re-running is the normal case; what would break it is a message embedding a count,
        # and neither of these does.
        return [], [*_check_gene_metrics_arithmetic(rows), *_cross_check_gene_metrics(rows, variants)]

    def _literature_checks(rows: list) -> tuple[list[str], list[str]]:
        # De-duplicated on the message: `compile_module` runs `validate_spec`, which runs this same
        # check, so a finding living in both places would otherwise print twice (the
        # `_check_contig_ploidy` idiom).
        return [], [w for w in _cross_check_literature(rows, studies, kind_rows) if w not in all_warnings]

    def _gene_validity_checks(rows: list) -> tuple[list[str], list[str]]:
        # Both warn in either mode (see `_check_gene_validity_currency`), so nothing here reads
        # `strict` — the errors list stays empty by construction rather than by a branch.
        return [], [
            *_cross_check_gene_validity(rows, variants),
            *_check_gene_validity_currency(rows),
        ]

    def _clinical_assertion_checks(rows: list) -> tuple[list[str], list[str]]:
        return [], list(_cross_check_clinical_assertions(rows, variants))

    def _gwas_effect_checks(rows: list) -> tuple[list[str], list[str]]:
        return [], list(_cross_check_gwas_effects(rows, variants))

    def _expression_effect_checks(rows: list) -> tuple[list[str], list[str]]:
        """No checks, and the absence is a decision rather than a gap (RM194/RM200).

        **`_cross_check_gwas_effects`'s orphan warning deliberately does not transfer**, though the
        two tables look alike enough that copying it would be the obvious move. A `GwasEffectRow` is
        module-scoped by construction: the pass queries the Catalog *with the module's own rsIDs*, so
        a row naming an identity the module does not carry really is the residue of a narrowed
        variant list, which is what that warning is about.

        `expression_effects.csv` is **locus-wide by construction instead**. The pass queries a
        genomic interval and AlphaGenome answers for every scored variant in it, most of which the
        module does not author — and finding those is the entire point of the item, because slicing
        by gene position silently drops the promoters and enhancers that act on a gene without
        sitting in it. Running the identical check here would fire on nearly every row of every
        module, which is a warning that means "this table is working".

        A gene-scoped variant of it — warn when `gene` names no gene the module annotates — was
        considered and left unbuilt. A warning code is a permanent key (`@warning-code-names-the-
        finding`), nobody has asked for this one, and minting one speculatively costs more than the
        check would return. It is additive if a caller ever wants it.

        Returning `([], [])` rather than reporting a zero is `@tautology-zero`: this is a table with
        no check, not a check that always passes.
        """
        return [], []

    def _concordance_checks(rows: list) -> tuple[list[str], list[str]]:
        # De-duplicated on the message, the `_literature_checks` idiom: `compile_module` runs
        # `validate_spec`, which emits the identical sentences over the identical post-overlay rows,
        # and an unfiltered re-run would publish the count twice. Safe to run twice at all because
        # no compile step between the two passes touches this table or the overlay above it — which
        # is the standing test `@no-rerun-with-counts` actually applies.
        warns = [w for w in _concordance_warnings(rows) if w not in all_warnings]
        warns.extend(
            w
            for w in _cross_check_clin_sig_concordance(rows, variants, table="clin_sig_concordance.csv")
            if w not in all_warnings
        )
        return [], warns

    def _authority_call_checks(rows: list) -> tuple[list[str], list[str]]:
        # No check that the subject also appears in the parent record, deliberately: an author who
        # answers a contested subject suppresses the parent row through the overlay, and the calls
        # behind it stay — the evidence outliving the question is correct, and warning about it would
        # make answering a finding produce a finding.
        return [], [
            w
            for w in _cross_check_clin_sig_concordance(rows, variants, table="clin_sig_authority_calls.csv")
            if w not in all_warnings
        ]

    def _sources_checks(rows: list) -> tuple[list[str], list[str]]:
        # `sources.csv` is last in `_FACT_TABLES`, so the other sidecars are already parsed into
        # `fact_rows` and their `source` values can be cross-checked here. Warnings only — the gate
        # that can actually refuse already ran, before anything was written.
        #
        # `SourceRow` is excluded from the "used" set: the loop stores each model's rows into
        # `fact_rows` *before* calling its check, so including it would let sources.csv vouch for
        # itself and no orphan could ever be reported.
        # Resolution contributes its **authority**, not its `source`: that column names which *link*
        # answered (`ensembl-rest`, `cache`) while every other table's names a licensed source, so
        # comparing them by string made every enriched module warn that `ensembl-rest` has no terms
        # recorded (RM33). A row with no authority contributes nothing — `authored`/`reversed` have no
        # external source to declare, and an older `resolution.csv` written before the column existed
        # simply says nothing rather than saying the wrong thing.
        used = {r.authority for r in resolution_rows if r.authority}
        for model, parsed in fact_rows.items():
            if model in (SourceRow, LiteratureRow):
                continue
            used |= {getattr(r, "source", None) for r in parsed if getattr(r, "source", None)}
        warns = _source_checks(rows, {s for s in used if s})
        warns.extend(_check_declared_license_agrees(rows, config.license if config else None))
        return [], warns

    _FACT_HANDLERS: dict[type, tuple[Callable, Callable]] = {
        FrequencyRow: (_frequency_checks, lambda rows: _build_frequencies(rows, module_name)),
        GeneMetricsRow: (_gene_metrics_checks, lambda rows: _build_table(rows, GeneMetricsRow, module_name)),
        LiteratureRow: (_literature_checks, lambda rows: _build_table(rows, LiteratureRow, module_name)),
        GeneValidityRow: (
            _gene_validity_checks,
            lambda rows: _build_table(rows, GeneValidityRow, module_name),
        ),
        ClinicalAssertionRow: (
            _clinical_assertion_checks,
            lambda rows: _build_table(rows, ClinicalAssertionRow, module_name),
        ),
        GwasEffectRow: (
            _gwas_effect_checks,
            lambda rows: _build_table(rows, GwasEffectRow, module_name),
        ),
        ExpressionEffectRow: (
            _expression_effect_checks,
            lambda rows: _build_table(rows, ExpressionEffectRow, module_name),
        ),
        ClinSigConcordanceRow: (
            _concordance_checks,
            lambda rows: _build_table(rows, ClinSigConcordanceRow, module_name),
        ),
        ClinSigAuthorityCallRow: (
            _authority_call_checks,
            lambda rows: _build_table(rows, ClinSigAuthorityCallRow, module_name),
        ),
        SourceRow: (_sources_checks, lambda rows: _build_table(rows, SourceRow, module_name)),
    }

    fact_rows: dict[type, list] = {}
    for csv_name, parquet_name, model in _FACT_TABLES:
        # Spelling collisions and the deprecation notice were both already surfaced by the licence
        # gate above (for `sources.csv`) or are surfaced here for the first time; either way the
        # notices de-duplicate, the way ploidy's and the VRS pass's already do.
        fact_path, fact_spelling_warnings, fact_spelling_errors = _locate_sidecar(spec_dir, csv_name)
        if fact_spelling_errors:
            return CompilationResult(success=False, errors=fact_spelling_errors, warnings=all_warnings)
        all_warnings.extend(w for w in fact_spelling_warnings if w not in all_warnings)
        if fact_path is None:
            continue
        rows, fact_errors, _ = _load_csv_rows(fact_path, model, fact_path.name)
        if fact_errors:
            return CompilationResult(success=False, errors=fact_errors, warnings=all_warnings)
        # The overlay, before the table's own check runs (RM124) — the check reports on what the
        # module asserts, and the parquet, the fact signature and the manifest block are all built
        # from the same post-overlay rows.
        if csv_name in OVERRIDABLE_TABLES:
            overlaid.add(csv_name)
            lossy = csv_name in LOSSY_OVERLAY_TABLES or csv_name in VINDICATING_OVERLAY_TABLES
            if lossy:
                compile_deferred[csv_name] = update_targets(csv_name, rows, overrides)
            rows, apply_errors, apply_warnings = apply_overrides(
                csv_name, rows, overrides, defer_unmatched=lossy
            )
            if apply_errors:
                return CompilationResult(success=False, errors=apply_errors, warnings=all_warnings)
            all_warnings.extend(w for w in apply_warnings if w not in all_warnings)
        check, build = _FACT_HANDLERS[model]
        # **The check sees every row; everything after it sees the kept ones** (RM79). A literature
        # row for a citation no study and no bin names joins to nothing, so carrying it into the
        # parquet and the manifest is dead weight — but reporting it needs the full list, which is why
        # the split happens here rather than at load. `literature.csv` itself is untouched: it is
        # merge-not-clobber on purpose, and that pin is what makes a re-run cheap.
        check_errors, check_warnings = check(rows)
        if model is LiteratureRow:
            rows, _dropped = split_cited_literature(rows, studies, kind_rows)
        fact_rows[model] = rows
        if check_errors:
            return CompilationResult(success=False, errors=check_errors, warnings=all_warnings)
        # De-duplicated on the message, like every other check that runs on both sides (RM94): a
        # fact-table check the pre-flight also performs reaches the identical sentence twice, and a
        # doubled line in `manifest.compilation.warnings` doubles `warnings_summary`'s count with it.
        # Both passes read the same POST-OVERLAY rows — `validate_spec` applies the overlay in its own
        # loop — so a message embedding a count says the same number on both sides and
        # `@no-rerun-with-counts` is satisfied rather than dodged. The dedup was previously
        # unnecessary here because no fact check ran in the pre-flight; RM108's currency check is the
        # first, and stating the rule once is cheaper than remembering it for the second.
        all_warnings.extend(w for w in check_warnings if w not in all_warnings)
        fact_df = build(rows)
        fact_df.write_parquet(output_dir / parquet_name, compression=compression)
        table_rows[parquet_name] = fact_df.height

    # De-duplicated on the message like every other check that runs on both sides: `validate_spec`
    # ran this over the same overlay and `all_warnings` was seeded from its result.
    all_warnings.extend(w for w in _overlay_targets_missing(overrides, overlaid) if w not in all_warnings)
    # And RM137's split, deferred from both overlay sites to here — `studies` and the citing tables are
    # in scope now. De-duplicated on the message for the reason every both-sides check is: the
    # pre-flight computed the identical sentence from the identical inputs.
    all_warnings.extend(
        w
        for w in _classify_deferred_overlay_updates(
            compile_deferred, studies, kind_rows, variants, resolution_rows
        )
        if w not in all_warnings
    )

    # The overlay itself, materialized verbatim and in authored order — it is authored input, so it
    # goes to parquet the way `variants.csv` does rather than being re-derived from what it changed.
    # This is what `reverse_module` reads the corrections back out of (RM124).
    if overrides:
        overlay_df = _build_table(overrides, OverrideRow, module_name)
        overlay_df.write_parquet(output_dir / OVERRIDES_PARQUET, compression=compression)
        table_rows[OVERRIDES_PARQUET] = overlay_df.height

    frequency_rows: list[FrequencyRow] = fact_rows.get(FrequencyRow, [])
    gene_metrics_rows: list[GeneMetricsRow] = fact_rows.get(GeneMetricsRow, [])
    literature_rows: list[LiteratureRow] = fact_rows.get(LiteratureRow, [])
    gene_validity_rows: list[GeneValidityRow] = fact_rows.get(GeneValidityRow, [])
    clinical_assertion_rows: list[ClinicalAssertionRow] = fact_rows.get(ClinicalAssertionRow, [])
    gwas_effect_rows: list[GwasEffectRow] = fact_rows.get(GwasEffectRow, [])
    expression_effect_rows: list[ExpressionEffectRow] = fact_rows.get(ExpressionEffectRow, [])
    concordance_rows: list[ClinSigConcordanceRow] = fact_rows.get(ClinSigConcordanceRow, [])
    authority_call_rows: list[ClinSigAuthorityCallRow] = fact_rows.get(ClinSigAuthorityCallRow, [])
    source_rows: list[SourceRow] = fact_rows.get(SourceRow, [])

    logs = _collect_logs(spec_dir, output_dir, log_files)
    # Authored side-car assets are validated here (validate_spec does not read them). Surface a
    # malformed one as a compile error instead of letting the exception escape mid-compile.
    try:
        provenance = _collect_provenance(spec_dir, output_dir, provenance_file)
    except ValidationError as exc:
        return CompilationResult(
            success=False, errors=[f"provenance.json is invalid: {exc}"], warnings=all_warnings
        )
    try:
        logo = _collect_logo(spec_dir, output_dir, logo_file)
    except ValueError as exc:
        return CompilationResult(success=False, errors=[str(exc)], warnings=all_warnings)
    try:
        readme = _collect_readme(spec_dir, output_dir, readme_file)
    except ValueError as exc:
        return CompilationResult(success=False, errors=[str(exc)], warnings=all_warnings)
    # The attestation, re-read here because the block is what gets stamped (`validate_spec` ran the
    # same call for its warning and threw the block away). De-duplicated on the message for the
    # standard reason: the pre-flight already emitted the identical sentence.
    verification, verification_warnings = _verification_block(spec_dir)
    all_warnings.extend(w for w in verification_warnings if w not in all_warnings)

    # Content identity over the RAW authored data (re-read from disk, so pre-resolution and
    # reference-independent — the in-scope `variants` here are already resolved). Out of
    # `artifact.digest`; lets a registry dedup across recompile/metadata-strip.
    manifest = _build_manifest(
        config=config,
        spec_dir=spec_dir,
        output_dir=output_dir,
        validation=validation,
        weights_rows=weights_df.height if weights_df is not None else 0,
        warnings=all_warnings,
        dropped_rows={t: len(r) for t, r in sorted(symbolic_drops.items())},
        compiled_by=compiled_by,
        ensembl_reference=ensembl_reference,
        logs=logs,
        provenance=provenance,
        logo=logo,
        readme=readme,
        content_sig=content_signature(spec_dir),
        resolution_mode=resolution_mode,
        fully_resolved=fully_resolved,
        resolution_subjects=resolution_subjects,
        expanded_keys=expanded_keys,
        expanded_rows=expanded_rows,
        positional_rows=positional_rows,
        positional_rows_placed=positional_rows_placed,
        vrs_alleles=vrs_alleles,
        vrs_alleles_identified=vrs_identified,
        resolution_sig=resolution_sig,
        resolution_sources=resolution_sources,
        frequency=_frequency_block(frequency_rows),
        gene_metrics=_gene_metrics_block(gene_metrics_rows),
        gene_validity=_gene_validity_block(gene_validity_rows),
        clinical_assertions=_clinical_assertions_block(clinical_assertion_rows),
        gwas_effects=_gwas_effects_block(gwas_effect_rows),
        expression_effects=_expression_effects_block(expression_effect_rows),
        clin_sig_concordance=_clin_sig_concordance_block(concordance_rows, authority_call_rows),
        literature=_literature_block(literature_rows),
        sources=_sources_block(source_rows),
        verification=verification,
    )
    write_manifest(manifest, output_dir / "manifest.json")

    stats: dict[str, Any] = {
        "module_name": module_name,
        "weights_rows": weights_df.height if weights_df is not None else 0,
        "annotations_rows": annotations_df.height if annotations_df is not None else 0,
        "studies_rows": studies_df.height if studies_df is not None else 0,
        "table_rows": table_rows,
    }
    return CompilationResult(
        success=True,
        output_dir=output_dir,
        errors=[],
        warnings=all_warnings,
        stats=stats,
        manifest=manifest,
    )

close_module

close_module(
    spec_dir: Path,
    *,
    closed_by: str | None = None,
    private_key_pem: bytes | None = None,
    now: str | None = None,
    difficulty: int | None = None,
) -> ClosureResult

Declare a module's authoring phase finished, binding the statement to its authored bytes (RM73).

A flat CSV row records nothing about how it came to be, so authoring had no end and every check that needed one guessed. This is the end: a closure block inside the module's verification.json naming the hash of the authored files as they stand. A later edit moves that hash, the compiler recomputes it, and the closure is dropped along with the rest of the attestation — which is why there is no second file and no second binding here to keep in step.

Deliberate, never a side effect. validate_spec stays read-only and nothing stamps this on a passing run: a record written by whatever happened to execute says only someone ran a tool, which is the exact defect RM73 levels at an attestation produced as a by-product. So this is its own function behind its own command, and --private-key makes the act attributable rather than merely evident.

It refuses on an invalid spec and not on a warning. Declaring a set finished that the compiler will not accept is a contradiction; declaring one finished that carries an unresolvable rsID or an ungrounded threshold is ordinary, and refusing there would make closure unreachable for every module whose findings no authored edit can clear (P5, the not_covered class).

Existing check records survive only while they describe these bytes, and then the whole document is kept verbatim rather than rebuilt — producer names who put the checks, so stamping this tier's label over it would have the compiler claim an enricher's cross-checks. Records that no longer hold are dropped and named in dropped_checks: carrying them across would re-bind a claim to rows the check never saw, which is the failure module_hash exists to catch, committed by the tool instead of by an edit.

Source code in compiler/src/just_dna_compiler/compiler.py
def close_module(
    spec_dir: Path,
    *,
    closed_by: str | None = None,
    private_key_pem: bytes | None = None,
    now: str | None = None,
    difficulty: int | None = None,
) -> ClosureResult:
    """Declare a module's authoring phase finished, binding the statement to its authored bytes (RM73).

    A flat CSV row records nothing about how it came to be, so authoring had no end and every check
    that needed one guessed. This is the end: a `closure` block inside the module's `verification.json`
    naming the hash of the authored files as they stand. A later edit moves that hash, the compiler
    recomputes it, and the closure is dropped along with the rest of the attestation — which is why
    there is no second file and no second binding here to keep in step.

    **Deliberate, never a side effect.** `validate_spec` stays read-only and nothing stamps this on a
    passing run: a record written by whatever happened to execute says only *someone ran a tool*,
    which is the exact defect RM73 levels at an attestation produced as a by-product. So this is its
    own function behind its own command, and `--private-key` makes the act attributable rather than
    merely evident.

    **It refuses on an invalid spec and not on a warning.** Declaring a set finished that the compiler
    will not accept is a contradiction; declaring one finished that carries an unresolvable rsID or an
    ungrounded threshold is ordinary, and refusing there would make closure unreachable for every
    module whose findings no authored edit can clear (P5, the `not_covered` class).

    Existing check records survive **only while they describe these bytes**, and then the whole
    document is kept verbatim rather than rebuilt — `producer` names who put the *checks*, so stamping
    this tier's label over it would have the compiler claim an enricher's cross-checks. Records that no
    longer hold are dropped and named in `dropped_checks`: carrying them across would re-bind a claim
    to rows the check never saw, which is the failure `module_hash` exists to catch, committed by the
    tool instead of by an edit.
    """
    spec_dir = Path(spec_dir)
    validation, validation_findings = _validate_spec(spec_dir)
    if not validation.valid:
        return ClosureResult(
            closed=False,
            errors=[
                "This spec does not validate, so its authoring set cannot be declared finished. "
                "Fix the errors below and close it afterwards.",
                *validation.errors,
            ],
            # Everything the pre-flight found, except its reminder to run *this command* — which is
            # what an author was being told, as the first line of output, while running it. The
            # pre-flight is right to say it and this caller is the one context where it is answered
            # by definition. Filtered on the phrase rather than by re-deciding, so the two cannot
            # drift apart.
            warnings=[w for w in validation_findings if UNCLOSED_PHRASE not in w],
        )

    try:
        path = sidecar_write_path(spec_dir, VERIFICATION_JSON)
    except SidecarCollision as exc:
        return ClosureResult(closed=False, errors=[str(exc)])

    binding = _module_binding(spec_dir)
    stamp = now or now_utc_iso()
    statement = close(binding, closed_at=stamp, closed_by=closed_by, private_key_pem=private_key_pem)
    warnings: list[str] = []
    previous: VerificationDoc | None = None
    if path.is_file():
        try:
            previous = read_verification(path)
        except (OSError, ValueError) as exc:
            warnings.append(
                CodedWarning(
                    "closure_discarded_unreadable_record",
                    f"The existing {path.name} could not be read ({exc}); this closure replaces it, so "
                    f"any checks it recorded are gone. Re-run the checks (just-dna-enricher).",
                )
            )

    held = previous is not None and attestation_failure(previous, binding) is None
    if held:
        # Everything the document says still holds, so the closure is the only new claim in it and the
        # rest is kept **verbatim** — including `producer`, which names who put the *checks*. Writing
        # this tier's own label there would say the compiler ran an enricher's cross-checks, which is
        # a false claim manufactured by an unrelated act. Reusing the document also keeps the nonce
        # already mined over an unchanged payload, so closing costs no work rather than the same work
        # twice.
        doc = previous.model_copy(update={"closure": statement})
    else:
        # Either there was no document, or it no longer describes these bytes. Its records are dropped
        # rather than re-attested: re-binding them would claim a check was put against rows it never
        # saw, which is the failure `module_hash` exists to catch, committed by the tool instead of by
        # an edit.
        #
        # `producer` and `produced_at` both stay unset, as a pair: they describe the run that put the
        # checks, and this document has none. Stamping the time here left a closure-only manifest
        # reading `producer: null, produced_at: <now>, checks: []` — a timestamp for a run that did not
        # happen, beside the closure's own `closed_at` saying the same thing about the act that did.
        doc = attest([], binding, difficulty=difficulty, closure=statement)
    dropped = sorted(r.check for r in previous.records) if previous is not None and not held else []
    write_verification(doc, path)
    return ClosureResult(
        closed=True,
        path=path,
        module_hash=binding,
        signed=doc.closure is not None and doc.closure.signature is not None,
        dropped_checks=dropped,
        warnings=warnings,
    )

build_disagreement_error

build_disagreement_error(
    block: Verification | None,
) -> str | None

The one recorded finding strict refuses on, or None (S78, RM143).

This does not move the strict line, and the distinction is the whole item. strict means reproducible, never right — the compiler has no reference, so it cannot check a coordinate, and a whole file shifted by one base still passes. genome_build_agreement is the exception on internal-consistency grounds rather than correctness ones: a recorded finding there says the module's rows are on a different assembly than the genome_build it declares, which is one authored file contradicting another. Every other recorded finding is a disagreement between the module and an outside archive, where the archive is the stale side often enough that failing a build would have the format arbitrate someone else's dispute — that reasoning is unchanged and covers clinical_significance, reference_allele and the rest.

A fact the toolchain already established, not a check re-run here. The judgement is the enricher's, made against the GRCh37 service the compiler may never call (Principle 2); what changed is that it stops being discarded at the boundary. So the gate keys on a record the enricher wrote — findings > 0 on that one check — and the compiler adds no reference, no network and no opinion of its own.

Silent when no attestation exists, deliberately, and that is not a hole this leaves open: an unverified module is the ordinary case, _read_verification_block says nothing about it on purpose, and refusing there would fail every module that has never been enriched. What this closes is the case where the answer was obtained and thrown away.

A stale attestation is dropped before this sees it, which is the correct order: bytes that moved since the check ran are bytes the check did not judge.

Source code in compiler/src/just_dna_compiler/compiler.py
def build_disagreement_error(block: Verification | None) -> str | None:
    """The one recorded finding `strict` refuses on, or `None` (S78, RM143).

    **This does not move the strict line, and the distinction is the whole item.** `strict` means
    *reproducible*, never *right* — the compiler has no reference, so it cannot check a coordinate, and
    a whole file shifted by one base still passes. `genome_build_agreement` is the exception on
    internal-consistency grounds rather than correctness ones: a recorded finding there says the
    module's rows are **on a different assembly than the `genome_build` it declares**, which is one
    authored file contradicting another. Every other recorded finding is a disagreement between the
    module and an outside archive, where the archive is the stale side often enough that failing a
    build would have the format arbitrate someone else's dispute — that reasoning is unchanged and
    covers `clinical_significance`, `reference_allele` and the rest.

    **A fact the toolchain already established, not a check re-run here.** The judgement is the
    enricher's, made against the GRCh37 service the compiler may never call (Principle 2); what changed
    is that it stops being discarded at the boundary. So the gate keys on a *record* the enricher
    wrote — `findings > 0` on that one check — and the compiler adds no reference, no network and no
    opinion of its own.

    **Silent when no attestation exists, deliberately**, and that is not a hole this leaves open: an
    unverified module is the ordinary case, `_read_verification_block` says nothing about it on purpose,
    and refusing there would fail every module that has never been enriched. What this closes is the
    case where the answer *was* obtained and thrown away.

    A stale attestation is dropped before this sees it, which is the correct order: bytes that moved
    since the check ran are bytes the check did not judge.
    """
    if block is None:
        return None
    found = [r for r in block.checks if r.check == BUILD_AGREEMENT_CHECK and r.findings]
    if not found:
        return None
    total = sum(r.findings for r in found)
    subjects = sum(r.subjects for r in found)
    return (
        f"strict compile: verification.json records {total} row(s) of {subjects} whose coordinates "
        f"the enricher diagnosed as another assembly's ({BUILD_AGREEMENT_CHECK}). The module declares "
        f"a genome_build its own rows contradict, so the artifact would be internally consistent and "
        f"about the wrong locus. Read the record's `detail` for which rows and the rs-numbers to "
        f"author instead, fix the coordinates and re-run the checks — or compile without strict, "
        f"which builds it and says so."
    )

split_cited_literature

split_cited_literature(
    rows: list[LiteratureRow],
    studies: list[StudyRow],
    kind_rows: dict[str, list[Any]] | None = None,
) -> tuple[list[LiteratureRow], list[LiteratureRow]]

(kept, dropped) — the literature rows this module actually cites, and the rest (RM79).

The compiler discards the rest; literature.csv keeps them. A row describing a citation no study and no citing table row names is dead weight in the artifact: nothing joins to it, and it is only there because literature.csv is merge-not-clobber, so a citation the author has since deleted from studies.csv leaves its row behind. Keeping the row in the CSV is the point of that rule — it is the pin that makes a re-run cheap — and carrying it into the parquet and the manifest is a separate decision that nobody had taken deliberately.

What this settles. manifest.literature.missing_count counted exists is False over every row in the table while the citation_existence verification record counted over the module's current citations, so the two disagreed in a published manifest with nothing wrong in the module. Both were honest about their own subject, which is what made it a decision rather than a bug. Filtering here makes them the same subject by construction, rather than documenting a discrepancy a reader would have to reconcile.

cited empty means discard nothing, deliberately, and it is not the degenerate case it looks like: a module that cites nothing at all cannot distinguish "the sidecar is stale" from "the citations are not authored yet", and emptying its whole table on that reading would delete an enrichment pass's entire output. The if cited guard the orphan check already had is kept for the same reason it existed.

On the round trip. reverse_module rebuilds literature.csv from the parquet, so a reversed copy carries the kept rows only. That is a deterministic narrowing rather than a P7 breach — literature.csv is a machine-written derived sidecar, not an authored value (the RM69 reading of Principle 7's letter) — and it converges: everything in the parquet is cited by construction, so lap two discards nothing and the signatures are a fixed point. The rows are recoverable the way every derived sidecar's are, by re-running the pass.

Source code in compiler/src/just_dna_compiler/compiler.py
def split_cited_literature(
    rows: list[LiteratureRow],
    studies: list[StudyRow],
    kind_rows: dict[str, list[Any]] | None = None,
) -> tuple[list[LiteratureRow], list[LiteratureRow]]:
    """`(kept, dropped)` — the literature rows this module actually cites, and the rest (RM79).

    **The compiler discards the rest; `literature.csv` keeps them.** A row describing a citation no
    study and no citing table row names is dead weight in the artifact: nothing joins to it, and it
    is only there
    because `literature.csv` is merge-not-clobber, so a citation the author has since deleted from
    `studies.csv` leaves its row behind. Keeping the row in the CSV is the point of that rule — it is
    the pin that makes a re-run cheap — and carrying it into the parquet and the manifest is a
    separate decision that nobody had taken deliberately.

    **What this settles.** `manifest.literature.missing_count` counted `exists is False` over *every*
    row in the table while the `citation_existence` verification record counted over the module's
    *current* citations, so the two disagreed in a published manifest with nothing wrong in the
    module. Both were honest about their own subject, which is what made it a decision rather than a
    bug. Filtering here makes them the same subject **by construction**, rather than documenting a
    discrepancy a reader would have to reconcile.

    **`cited` empty means discard nothing**, deliberately, and it is not the degenerate case it looks
    like: a module that cites nothing at all cannot distinguish "the sidecar is stale" from "the
    citations are not authored yet", and emptying its whole table on that reading would delete an
    enrichment pass's entire output. The `if cited` guard the orphan check already had is kept for the
    same reason it existed.

    **On the round trip.** `reverse_module` rebuilds `literature.csv` from the parquet, so a reversed
    copy carries the kept rows only. That is a deterministic narrowing rather than a P7 breach —
    `literature.csv` is a machine-written derived sidecar, not an authored value (the RM69 reading of
    Principle 7's letter) — and it **converges**: everything in the parquet is cited by construction,
    so lap two discards nothing and the signatures are a fixed point. The rows are recoverable the way
    every derived sidecar's are, by re-running the pass.
    """
    cited = cited_pmids(studies, kind_rows)
    if not cited:
        return list(rows), []
    kept = [r for r in rows if r.pmid in cited]
    return kept, [r for r in rows if r.pmid not in cited]

cited_pmids

cited_pmids(
    studies: list[StudyRow],
    kind_rows: dict[str, list[Any]] | None = None,
) -> set[str]

Every PMID this module cites, from both citation sites, through the one normalizer.

Extracted so RM137's reachability predicate asks the same question split_cited_literature answers, rather than a second statement of it. The two would drift silently and in the worst direction: the predicate would call a row unreachable that the drop had kept, so a healthy overlay would report a finding forever.

The empty case is the caller's to interpret, and it is not "nothing is cited". A module citing nothing cannot distinguish a stale sidecar from citations not yet authored, so split_cited_literature discards nothing there — and the predicate must mirror that or every literature update on such a module reads as unreachable. literature_target_survives builds the mirror; nothing should test this set for emptiness on its own.

Source code in compiler/src/just_dna_compiler/compiler.py
def cited_pmids(studies: list[StudyRow], kind_rows: dict[str, list[Any]] | None = None) -> set[str]:
    """Every PMID this module cites, from both citation sites, through the one normalizer.

    Extracted so RM137's reachability predicate asks the **same** question `split_cited_literature`
    answers, rather than a second statement of it. The two would drift silently and in the worst
    direction: the predicate would call a row unreachable that the drop had kept, so a healthy overlay
    would report a finding forever.

    **The empty case is the caller's to interpret, and it is not "nothing is cited".** A module citing
    nothing cannot distinguish a stale sidecar from citations not yet authored, so `split_cited_literature`
    discards nothing there — and the predicate must mirror that or every literature `update` on such a
    module reads as unreachable. `literature_target_survives` builds the mirror; nothing should test
    this set for emptiness on its own.
    """
    cited: set[str] = set()
    for study in studies:
        cited.update(extract_pmids(study.pmid))
    cited.update(table_citations(kind_rows or {}))
    return cited

literature_target_survives

literature_target_survives(
    studies: list[StudyRow],
    kind_rows: dict[str, list[Any]] | None = None,
) -> Callable[[str], bool]

Can an artifact of this module carry a literature.csv row for this PMID? (RM137)

True when the PMID is cited, because split_cited_literature keeps exactly the cited rows — and True for everything when the module cites nothing at all, which is that function's own guard reproduced rather than restated. Without the guard a module with no citations would mark every literature correction unreachable: a stable false positive, which is worse than the unstable true one RM137 is about.

Source code in compiler/src/just_dna_compiler/compiler.py
def literature_target_survives(
    studies: list[StudyRow], kind_rows: dict[str, list[Any]] | None = None
) -> Callable[[str], bool]:
    """Can an artifact of this module carry a `literature.csv` row for this PMID? (RM137)

    True when the PMID is cited, because `split_cited_literature` keeps exactly the cited rows — and
    **True for everything when the module cites nothing at all**, which is that function's own guard
    reproduced rather than restated. Without the guard a module with no citations would mark every
    literature correction unreachable: a stable false positive, which is worse than the unstable true
    one RM137 is about.
    """
    cited = cited_pmids(studies, kind_rows)
    if not cited:
        return lambda pmid: True
    return lambda pmid: pmid in cited

resolution_target_survives

resolution_target_survives(
    variants: list[VariantRow],
    resolution_rows: list[ResolutionRow],
) -> Callable[[str], bool]

Can an artifact of this module carry a resolution.csv row for this variant_key? (RM137)

resolution.csv has no parquet, so reverse_module rebuilds it from the SNP core and _write_resolution_csv skips a row with no resolved position — "rows without a resolved position carry no fact and are skipped". So the surviving set is the subjects this module can place.

Computed from the authored coordinates and the injected table together, which is what makes it answer the same on both laps: on lap 1 the unpositioned row is present and its own cells say it is unpositioned; on lap 2 the row is gone and the authored side still says the same thing. Neither reading depends on the row being there to be matched.

Source code in compiler/src/just_dna_compiler/compiler.py
def resolution_target_survives(
    variants: list[VariantRow], resolution_rows: list[ResolutionRow]
) -> Callable[[str], bool]:
    """Can an artifact of this module carry a `resolution.csv` row for this `variant_key`? (RM137)

    `resolution.csv` has no parquet, so `reverse_module` rebuilds it from the SNP core and
    `_write_resolution_csv` skips a row with no resolved position — *"rows without a resolved position
    carry no fact and are skipped"*. So the surviving set is the subjects this module can **place**.

    Computed from the authored coordinates and the injected table together, which is what makes it
    answer the same on both laps: on lap 1 the unpositioned row is present and its own cells say it is
    unpositioned; on lap 2 the row is gone and the authored side still says the same thing. Neither
    reading depends on the row being there to be matched.
    """
    placed: set[str] = {row.variant_key for row in resolution_rows if row.chrom and row.start is not None}
    placed.update(v.variant_key for v in variants if v.chrom and v.start is not None)
    return lambda subject: subject in placed

reverse_module

reverse_module(
    parquet_dir: Path,
    output_dir: Path,
    module_name: str | None = None,
    title: str | None = None,
    description: str | None = None,
    report_title: str | None = None,
    icon: str = "database",
    color: str = "#6435c9",
    version: str | None = None,
    write_resolution: bool = True,
    genome_build: str | None = None,
) -> Path

Reverse-engineer a parquet module back into the spec DSL (yaml + csv). Returns output_dir.

version (like title/description) is authored module: metadata, out of artifact.digest and so not materialized into any parquet. Since RM103 it is recovered from the artifact's own manifest.json when the caller supplies none, rather than dropped: an explicit argument still wins, and a bare parquet directory with no manifest still leaves the key out of the block. What is recovered is identity.version_coerced_from where the compile recorded one, falling back to identity.version — the pre-coercion string, because re-emitting the coerced one gives the next compile nothing to coerce and version_coerced_from then goes absent on lap 2, which is a module disagreeing with its own round trip on a published field.

genome_build is not in that class, even though it reaches the artifact the same way (the manifest, never a parquet column). A wrong title is cosmetic; a wrong build relocates every coordinate in the module. This used to be hardcoded "GRCh38", so compile → reverse → compile on a genome_build: GRCh37 module re-emitted it as GRCh38 and the recompile minted ga4gh:VA.… ids — GRCh38 allele identities for GRCh37 positions — moving artifact.digest and asserting a variant at a base the module never named. Resolution order is therefore: this argument, else the artifact's own manifest.json, else "GRCh38" for a bare parquet directory that records nothing.

write_resolution (default True) also emits resolution.csv — the resolved facts recovered from the artifact — so reverse → compile reproduces the identical artifact.digest with no network and no Ensembl reference (Principle 7 hardened from reference-dependent to self-contained). A coord-keyed row's resolved rsid, dropped from variants.csv, is carried here and restored on recompile via resolution.resolve_from_table.

The authored tables are written under their one legal name at the root. The machine-written sidecars go through layout.sidecar_write_path, so a fresh tree gets the preferred spelling (licensing.csv, not the deprecated sources.csv _FACT_TABLES still names for its parquet) and an output directory that already carries a copy has that copy overwritten rather than joined by a second one. Reversing into a directory that already holds two copies of one sidecar raises layout.SidecarCollision: which of two hand-editable claims to overwrite is not something this function may decide silently.

Source code in compiler/src/just_dna_compiler/compiler.py
def reverse_module(
    parquet_dir: Path,
    output_dir: Path,
    module_name: str | None = None,
    title: str | None = None,
    description: str | None = None,
    report_title: str | None = None,
    icon: str = "database",
    color: str = "#6435c9",
    version: str | None = None,
    write_resolution: bool = True,
    genome_build: str | None = None,
) -> Path:
    """Reverse-engineer a parquet module back into the spec DSL (yaml + csv). Returns output_dir.

    `version` (like `title`/`description`) is authored `module:` metadata, out of `artifact.digest`
    and so not materialized into any parquet. **Since RM103 it is recovered from the artifact's own
    `manifest.json` when the caller supplies none**, rather than dropped: an explicit argument still
    wins, and a bare parquet directory with no manifest still leaves the key out of the block. What is
    recovered is `identity.version_coerced_from` where the compile recorded one, falling back to
    `identity.version` — the pre-coercion string, because re-emitting the coerced one gives the next
    compile nothing to coerce and `version_coerced_from` then goes absent on lap 2, which is a module
    disagreeing with its own round trip on a published field.

    `genome_build` is **not** in that class, even though it reaches the artifact the same way (the
    manifest, never a parquet column). A wrong title is cosmetic; a wrong build relocates every
    coordinate in the module. This used to be hardcoded `"GRCh38"`, so
    `compile → reverse → compile` on a `genome_build: GRCh37` module re-emitted it as GRCh38 and the
    recompile minted `ga4gh:VA.…` ids — GRCh38 allele identities for GRCh37 positions — moving
    `artifact.digest` and asserting a variant at a base the module never named. Resolution order is
    therefore: this argument, else the artifact's own `manifest.json`, else `"GRCh38"` for a bare
    parquet directory that records nothing.

    `write_resolution` (default True) also emits `resolution.csv` — the resolved facts recovered from
    the artifact — so `reverse → compile` reproduces the identical `artifact.digest` with **no network
    and no Ensembl reference** (Principle 7 hardened from reference-dependent to self-contained). A
    coord-keyed row's resolved rsid, dropped from `variants.csv`, is carried here and restored on
    recompile via `resolution.resolve_from_table`.

    The authored tables are written under their one legal name at the root. The machine-written
    sidecars go through `layout.sidecar_write_path`, so a fresh tree gets the **preferred** spelling
    (`licensing.csv`, not the deprecated `sources.csv` `_FACT_TABLES` still names for its parquet) and
    an output directory that already carries a copy has that copy overwritten rather than joined by a
    second one. Reversing into a directory that already holds two copies of one sidecar raises
    `layout.SidecarCollision`: which of two hand-editable claims to overwrite is not something this
    function may decide silently."""
    parquet_dir = Path(parquet_dir)
    output_dir = Path(output_dir)
    output_dir.mkdir(parents=True, exist_ok=True)

    # SNP core is optional (RM2): a module may have no weights.parquet.
    weights_path = parquet_dir / "weights.parquet"
    weights_df = pl.read_parquet(weights_path) if weights_path.is_file() else None

    # Every sidecar destination is resolved BEFORE the first write, and only for the tables this
    # artifact will actually produce. `sidecar_write_path` raises on an output directory that already
    # holds two copies of one table, and resolving late would raise it *after* `module_spec.yaml` and
    # the authored CSVs had been rewritten — a refusal that leaves a half-rebuilt spec behind. The
    # collision is refused with nothing touched instead, which is what the rest of this layout does.
    sidecar_paths: dict[str, Path] = {
        csv_name: sidecar_write_path(output_dir, csv_name)
        for csv_name, parquet_name, _ in _FACT_TABLES
        if (parquet_dir / parquet_name).is_file()
    }
    # Not `and weights_df is not None`: since RM43 the lookup table is rebuilt from the positional
    # parquets too, so a table-only module emits one and needs its path resolved the same way. The
    # writer itself returns before creating a file when there is nothing to write, so resolving the
    # path unconditionally cannot leave an empty `resolution.csv` behind.
    if write_resolution:
        sidecar_paths["resolution.csv"] = sidecar_write_path(output_dir, "resolution.csv")

    if module_name is None:
        module_name = _module_name_from_parquets(parquet_dir) or parquet_dir.name
    if genome_build is None:
        genome_build = _genome_build_from_artifact(parquet_dir) or "GRCh38"
    if version is None:
        version = _authored_version_from_artifact(parquet_dir)

    # **The attestation cannot be carried, and the silence about that was the defect (RM45).**
    # `verification.json` records checks the *enricher* put against sources this tier cannot reach,
    # and it is bound to the authored bytes by a hash — so reverse has nothing to rebuild it from and
    # must not invent one. What it can do is say so: without this line a module round-trips into a
    # spec that recompiles to a manifest with no `verification` block at all, and
    # `manifest.compilation.warnings` — a surface consumers parse (RM44) — differs between a module
    # and its own round trip with nothing edited. Losing a record of what was checked is acceptable;
    # losing it invisibly is the S16 silent-success shape.
    dropped_verification = _artifact_verification(parquet_dir)
    if dropped_verification is not None:
        logger.warning(_verification_loss_notice(dropped_verification), output_dir)

    default_curator = "unknown"
    default_method = "unknown"
    # `priority` is intentionally NOT defaulted. It is Optional with no `Defaults.priority` fallback
    # ('ai-module-creator'/'literature-review' back curator/method, but priority defaults to None),
    # so a null priority is *authored-absent*. Inferring a default from the mode would fabricate a
    # value for rows that never set one — turning weights `['high', None]` into `['high', 'high']`
    # on recompile (a Principle 7 idempotency break). It is written verbatim, per row, instead.
    default_priority: str | None = None
    if weights_df is not None:
        default_curator = _most_common(weights_df, "curator") or "unknown"
        default_method = _most_common(weights_df, "method") or "unknown"

    defaults_dict: dict[str, Any] = {"curator": default_curator, "method": default_method}

    module_block: dict[str, Any] = {
        "name": module_name,
        "title": title or module_name.replace("_", " ").title(),
        "description": description or f"Annotation module: {module_name}",
        "report_title": report_title or module_name.replace("_", " ").title(),
        "icon": icon,
        "color": color,
    }
    if version is not None:
        module_block["version"] = version
    spec = {
        "schema_version": "1.0",
        "module": module_block,
        "defaults": defaults_dict,
        "genome_build": genome_build,
    }
    (output_dir / "module_spec.yaml").write_text(
        yaml.dump(spec, default_flow_style=False, sort_keys=False), encoding="utf-8"
    )

    # variants.csv + studies.csv only when the module has them.
    if weights_df is not None:
        ann_lookup: dict[tuple, dict[str, str]] = {}
        ann_key_columns: tuple[str, ...] = ()
        ann_path = parquet_dir / "annotations.parquet"
        if ann_path.exists():
            ann_df = pl.read_parquet(ann_path)
            # Which columns follow `variant_key` in this artifact's annotation key, read off the
            # artifact rather than assumed — three generations of the table are in the wild and each
            # keyed differently: 0.6 on (variant_key, genotype, conclusion, negatives) (RM80), 0.5 on
            # the variant-effect pair, and the oldest on variant_key alone. Detected once and used by
            # both the lookup and the weights-side probe, so the two cannot key differently.
            # Order is fixed here, not by column order in the parquet, or the two sides could agree
            # on the members and disagree on the tuple.
            ann_key_columns = tuple(
                name for name in ("genotype", "conclusion", "negatives") if name in ann_df.columns
            )
            for row in ann_df.iter_rows(named=True):
                # variant_key so position-only variants (rsid null) match; fall back to rsid for an
                # older artifact compiled before the variant_key column existed.
                base = row.get("variant_key") or row.get("rsid")
                if base is None:
                    continue
                key = (base, *(row.get(name) for name in ann_key_columns))
                ann_lookup[key] = {
                    "gene": row.get("gene", ""),
                    "phenotype": row.get("phenotype", ""),
                    "category": row.get("category", ""),
                }
        _write_variants_csv(
            weights_df,
            ann_lookup,
            ann_key_columns,
            default_curator,
            default_method,
            default_priority,
            output_dir / "variants.csv",
            genome_build=genome_build,
        )
    studies_path = parquet_dir / "studies.parquet"
    if studies_path.exists():
        _write_studies_csv(pl.read_parquet(studies_path), output_dir / "studies.csv")

    # 0.4 table kinds (RM1): each present parquet → its authored CSV.
    positional_frames: list[tuple[str, pl.DataFrame]] = []
    for csv_name, parquet_name, model in _TABLE_KINDS:
        kind_path = parquet_dir / parquet_name
        if kind_path.is_file():
            kind_df = pl.read_parquet(kind_path)
            _write_table_csv(kind_df, model, output_dir / csv_name)
            if csv_name in {name for name, _model in _POSITIONAL_TABLE_KINDS}:
                positional_frames.append((csv_name, kind_df))

    # `resolution.csv` last, because it is rebuilt from everything above. It is written from the SNP
    # core **and** the positional tables since 0.6 (RM43): once the compiler fills a resolved
    # coordinate into `pharm_variants`/`haplotypes`/`heteroplasmy`, a reverse that dropped the lookup
    # table would emit a spec whose recompile leaves those parquets unfilled — so `compile → reverse →
    # compile` would stop reproducing the artifact, which is Principle 7.
    # Through `sidecar_paths`, never `output_dir / "resolution.csv"` — the literal join RM51 abolishes.
    # Worth spelling out because the two halves of this arrived in different lanes and nearly cancelled:
    # RM43 *moved* this call out of the `weights_df is not None` block so a table-only module emits one,
    # while RM51's repair was applied to the call site at its old address. Taking either side of that
    # merge alone loses the other.
    if write_resolution:
        _write_resolution_csv(
            weights_df,
            positional_frames,
            sidecar_paths["resolution.csv"],
            genome_build=genome_build,
        )

    # 0.5 derived-fact sidecars: same round-trip, minus the columns that are recomputed rather than
    # stored. `_write_table_csv` drops any parquet column the model does not declare, so
    # `allele_frequency` (derived on write, absent from `FrequencyRow`'s fields) falls away by
    # construction rather than by a special case — re-deriving it on the next compile reproduces the
    # identical parquet.
    #
    # The filename comes from `sidecar_paths` (resolved above), not `output_dir / csv_name`: `_FACT_TABLES`
    # names the licence table by its *deprecated* spelling (the parquet and the manifest key keep it,
    # since only a major may rename those), so joining that name on by hand emitted `sources.csv` and
    # made `compile → reverse → compile` deprecation-warn on a module whose own compile is silent —
    # and `manifest.compilation.warnings` is a published field (RM44), so the module and its own round
    # trip disagreed on it. This is the same "write to the file you read" rule every other writer
    # follows rather than an exception to it: reverse builds a spec tree from an artifact, so on a
    # fresh directory there is nothing to follow and the rule yields the preferred spelling, while
    # reversing over a tree that already carries the old name (or a `derived/` split) overwrites that
    # copy instead of leaving a second one behind — which is the collision the rule exists to prevent.
    for csv_name, parquet_name, model in _FACT_TABLES:
        fact_path = parquet_dir / parquet_name
        if fact_path.is_file():
            _write_table_csv(pl.read_parquet(fact_path), model, sidecar_paths[csv_name])

    # The authored overlay (RM124), at the spec **root** and under its one legal name — it is authored
    # like `variants.csv`, not a machine-written sidecar, so it does not go through `sidecar_paths`.
    #
    # **This emits the post-overlay derived tables above AND the overlay, so the overlay applies
    # twice, and that is the design rather than an oversight.** The alternative — emitting the
    # *pre*-overlay tables so the apply happens exactly once — needs the overlay to record the value
    # it replaced, which is a derived cell inside an authored table and rots the moment the source
    # moves. All three operations are idempotent set operations instead (an update to a value already
    # present, an insert of a row already keyed, a suppress of a row already absent are each a
    # no-op), so the second lap is a fixed point and buys the round trip at no schema cost. It is
    # checked by test, never assumed — Principle 7 requires that of every derivation.
    overlay_parquet = parquet_dir / OVERRIDES_PARQUET
    if overlay_parquet.is_file():
        _write_table_csv(pl.read_parquet(overlay_parquet), OverrideRow, output_dir / OVERRIDES_CSV)

    return output_dir