Skip to content

just_dna_enricher.constraint_build

just_dna_enricher.constraint_build

gnomAD gene-constraint snapshot builder — the [dev] half of the gene-metrics pass.

Turns gnomAD's v4.1 constraint metrics TSV into a small gene-level parquet the enricher reads offline (gnomad_constraint/data/*.parquet), in the same on-disk layout as the Ensembl and ClinVar snapshots. Modelled directly on clinvar_build, down to the streaming download with the .part rename and the release.json provenance record.

Why this one can ship offline and frequency cannot is pure size. The source TSV is 95.5 MB — per transcript, 55 columns — and reduces to one row per gene with a curated column subset, landing in single-digit MB. The v4.1 frequency data is 58 GB of exomes plus 742 GB of genomes, so there is no equivalent slice to build.

The row pick is the hard part, and it is load-bearing rather than cosmetic. The TSV is per-transcript and mixes RefSeq with Ensembl rows for the same gene, and — verified against the real file, not assumed — both carry mane_select=true::

A1BG   1                 NM_130786.4       canonical=true  mane_select=true    <- RefSeq
A1BG   ENSG00000121410   ENST00000263100   canonical=true  mane_select=true    <- Ensembl
A1BG   ENSG00000121410   ENST00000600966   canonical=false mane_select=false

(BRCA1 and MYH7 show the same shape, so this is the file's general structure, not an A1BG quirk.) A naive "first mane_select row wins" therefore returns whichever the file happens to list first, and for A1BG that is the RefSeq row whose gene_id is the bare NCBI id 1 — useless as a stable identity. The rule here is mane_select AND an ENSG-shaped gene_id, falling back to canonical on ENSG, and recording the gene as unresolved rather than guessing. _pick_row is fed the rows in both orders by a test, because same-input-different-order is exactly what the naive implementation gets wrong.

Builder-only (just-dna-enricher[dev]): polars is a guarded import, so the runtime gene-metrics path (which reads the built parquet through DuckDB) never needs it and the core install stays polars-free (Goal 2).

ConstraintBuildResult dataclass

ConstraintBuildResult(
    out_dir: Path,
    parquet_file: Path,
    gene_count: int,
    source_sha256: str,
    source_rows: int = 0,
    unresolved_genes: list[str] = list(),
)

Outcome of a snapshot build (paths + provenance + what was dropped and why).

ConstraintBuildError

Bases: RuntimeError

This lane had no error type at all until RM187, which is why its download leaked.

Every other builder here owns one. This module raised whatever its inputs raised, so caches._rebuild_constraint caught (FileNotFoundError, ImportError, OSError) and a transport failure went straight past it. A lane without its own type cannot be caught as that lane, and the downloader has nothing to translate into (@client-exception-contract).

ConstraintUnavailable

Bases: ConstraintBuildError

The gnomAD download did not answer — the fetch failed, not the file.

download_constraint_tsv

download_constraint_tsv(
    dest: Path, url: str = DEFAULT_CONSTRAINT_URL
) -> Path

Stream the gnomAD constraint TSV to dest (atomic .part rename), returning the path.

Provisioning-only — the gene-metrics pass never calls this; it reads a prebuilt cache. Shares net.stream_to_file with every other bulk download here, so the retry and the translation are one rule rather than eleven copies of one (RM187).

Source code in enricher/src/just_dna_enricher/constraint_build.py
def download_constraint_tsv(dest: Path, url: str = DEFAULT_CONSTRAINT_URL) -> Path:
    """Stream the gnomAD constraint TSV to ``dest`` (atomic ``.part`` rename), returning the path.

    Provisioning-only — the gene-metrics pass never calls this; it reads a prebuilt cache. Shares
    `net.stream_to_file` with every other bulk download here, so the retry and the translation are
    one rule rather than eleven copies of one (RM187).
    """
    return stream_to_file(
        dest,
        url,
        error_cls=ConstraintUnavailable,
        what="the gnomAD constraint TSV",
        remedy="Pass --source constraint=<tsv> to build from a copy you already hold.",
    ).path

build_snapshot

build_snapshot(
    tsv: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_CONSTRAINT_URL,
) -> ConstraintBuildResult

Convert the constraint TSV into data/gnomad_constraint.parquet + release.json.

One row per gene, chosen by _pick_row, sorted by gene so a rebuild is byte-identical (Principle 7); release.json's built_at is the only per-run-varying byte and lives outside the parquet, exactly as in the ClinVar builder.

Source code in enricher/src/just_dna_enricher/constraint_build.py
def build_snapshot(
    tsv: Path, out_dir: Path, *, source_url: str = DEFAULT_CONSTRAINT_URL
) -> ConstraintBuildResult:
    """Convert the constraint TSV into `data/gnomad_constraint.parquet` + `release.json`.

    One row per gene, chosen by `_pick_row`, sorted by `gene` so a rebuild is byte-identical
    (Principle 7); `release.json`'s `built_at` is the only per-run-varying byte and lives outside the
    parquet, exactly as in the ClinVar builder.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the gnomAD constraint snapshot; install the publisher/dev "
            "surface with `pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    tsv = Path(tsv)
    if not tsv.exists():
        raise FileNotFoundError(f"gnomAD constraint TSV not found at {tsv}")
    out_dir = Path(out_dir)
    data_dir = out_dir / "data"
    data_dir.mkdir(parents=True, exist_ok=True)

    by_gene: dict[str, list[dict]] = {}
    source_rows = 0
    for row in _iter_rows(tsv):
        gene = (row.get("gene") or "").strip()
        if not gene:
            continue
        source_rows += 1
        by_gene.setdefault(gene, []).append(row)

    records: list[dict] = []
    unresolved: list[str] = []
    for gene in sorted(by_gene):
        picked = _pick_row(by_gene[gene])
        if picked is None:
            unresolved.append(gene)
            continue
        records.append(_gene_record(gene, picked))

    schema = _empty_schema()
    frame = pl.DataFrame(records, schema=schema).sort("gene")
    parquet_path = data_dir / "gnomad_constraint.parquet"
    frame.write_parquet(parquet_path)

    source_sha256 = _sha256_file(tsv)
    _write_release_json(
        out_dir,
        source_url=source_url,
        source_sha256=source_sha256,
        gene_count=frame.height,
        source_rows=source_rows,
        unresolved_count=len(unresolved),
    )
    logger.info(
        "Built gnomAD constraint snapshot: %d genes from %d transcript rows → %s "
        "(%d gene(s) had no MANE/canonical Ensembl row and were dropped)",
        frame.height,
        source_rows,
        parquet_path,
        len(unresolved),
    )
    return ConstraintBuildResult(
        out_dir=out_dir,
        parquet_file=parquet_path,
        gene_count=frame.height,
        source_sha256=source_sha256,
        source_rows=source_rows,
        unresolved_genes=unresolved,
    )