just_dna_enricher.constraint_build¶
just_dna_enricher.constraint_build ¶
gnomAD gene-constraint snapshot builder — the [dev] half of the gene-metrics pass.
Turns gnomAD's v4.1 constraint metrics TSV into a small gene-level parquet the enricher reads offline
(gnomad_constraint/data/*.parquet), in the same on-disk layout as the Ensembl and ClinVar snapshots.
Modelled directly on clinvar_build, down to the streaming download with the .part rename and the
release.json provenance record.
Why this one can ship offline and frequency cannot is pure size. The source TSV is 95.5 MB — per transcript, 55 columns — and reduces to one row per gene with a curated column subset, landing in single-digit MB. The v4.1 frequency data is 58 GB of exomes plus 742 GB of genomes, so there is no equivalent slice to build.
The row pick is the hard part, and it is load-bearing rather than cosmetic. The TSV is
per-transcript and mixes RefSeq with Ensembl rows for the same gene, and — verified against the real
file, not assumed — both carry mane_select=true::
A1BG 1 NM_130786.4 canonical=true mane_select=true <- RefSeq
A1BG ENSG00000121410 ENST00000263100 canonical=true mane_select=true <- Ensembl
A1BG ENSG00000121410 ENST00000600966 canonical=false mane_select=false
(BRCA1 and MYH7 show the same shape, so this is the file's general structure, not an A1BG quirk.) A
naive "first mane_select row wins" therefore returns whichever the file happens to list first, and
for A1BG that is the RefSeq row whose gene_id is the bare NCBI id 1 — useless as a stable
identity. The rule here is mane_select AND an ENSG-shaped gene_id, falling back to
canonical on ENSG, and recording the gene as unresolved rather than guessing. _pick_row is fed the
rows in both orders by a test, because same-input-different-order is exactly what the naive
implementation gets wrong.
Builder-only (just-dna-enricher[dev]): polars is a guarded import, so the runtime gene-metrics
path (which reads the built parquet through DuckDB) never needs it and the core install stays
polars-free (Goal 2).
ConstraintBuildResult
dataclass
¶
ConstraintBuildResult(
out_dir: Path,
parquet_file: Path,
gene_count: int,
source_sha256: str,
source_rows: int = 0,
unresolved_genes: list[str] = list(),
)
Outcome of a snapshot build (paths + provenance + what was dropped and why).
ConstraintBuildError ¶
Bases: RuntimeError
This lane had no error type at all until RM187, which is why its download leaked.
Every other builder here owns one. This module raised whatever its inputs raised, so
caches._rebuild_constraint caught (FileNotFoundError, ImportError, OSError) and a transport
failure went straight past it. A lane without its own type cannot be caught as that lane, and
the downloader has nothing to translate into (@client-exception-contract).
ConstraintUnavailable ¶
Bases: ConstraintBuildError
The gnomAD download did not answer — the fetch failed, not the file.
download_constraint_tsv ¶
Stream the gnomAD constraint TSV to dest (atomic .part rename), returning the path.
Provisioning-only — the gene-metrics pass never calls this; it reads a prebuilt cache. Shares
net.stream_to_file with every other bulk download here, so the retry and the translation are
one rule rather than eleven copies of one (RM187).
Source code in enricher/src/just_dna_enricher/constraint_build.py
build_snapshot ¶
build_snapshot(
tsv: Path,
out_dir: Path,
*,
source_url: str = DEFAULT_CONSTRAINT_URL,
) -> ConstraintBuildResult
Convert the constraint TSV into data/gnomad_constraint.parquet + release.json.
One row per gene, chosen by _pick_row, sorted by gene so a rebuild is byte-identical
(Principle 7); release.json's built_at is the only per-run-varying byte and lives outside the
parquet, exactly as in the ClinVar builder.