Live gnomAD v4.1 queries — the third source in the network tier, and the first rate-limited one.
One GraphQL endpoint serves all three roles gnomAD plays here:
resolve_rsids — the resolver-chain link (rsID → loci), appended after live Ensembl;
fetch_frequencies — the frequency pass (coordinate → per-ancestry-group AC/AN);
fetch_gene_constraint — the gene pass's live fallback (symbol → pLI/LOEUF/Z scores).
Rate limiting is the design constraint, not an afterthought. gnomAD allows 10 requests per IP per
60 seconds, so one request per variant is unusable — a forty-variant module would take four minutes.
Everything here is therefore batched: GraphQL field aliasing puts many variant(...) lookups in one
POST. Probed live, batches of 20 and 25 both succeeded while 29 returned HTTP 400, so the batch size is
20 with a 6 second pacing gate — one request per 6s is exactly the 10/minute budget, and about
200 variants/minute.
Pacing comes before retry on purpose. tenacity retries transport errors, timeouts and 429s, but a
blind retry spends the same budget that caused the 429; the gate is what keeps us inside the limit, and
the retry is only for the genuinely transient case.
Partial failures are normal and must not sink a batch. GraphQL reports per-alias failures in
errors[] while still returning data for every alias that worked — a probed 20-alias batch came back
resolved=17, errors=3. Treating a non-empty errors[] as total failure would discard 17 good rows, so
every parser here reads both halves.
The query shapes and field names were checked against the live schema by introspection rather than
recalled: variant(variantId:|rsid:|vrsId:, dataset:), variant_search(query:, dataset:),
gene(gene_symbol:, reference_genome:), and a VariantPopulation that carries ac/an but
deliberately no af — the frequency is ours to compute.
GnomadError
Bases: RuntimeError
A gnomAD query failed outright (transport, HTTP status, or a malformed response body).
RateLimitedError
Bases: GnomadError
gnomAD answered HTTP 429. Retried with backoff; the pacing gate should normally prevent it.
GnomadClient
dataclass
GnomadClient(
settings: GnomadSettings = GnomadSettings(),
gate: PacingGate | None = None,
_client: Client | None = None,
)
Batched, paced gnomAD GraphQL client. Reused across a pass, so one gate covers every request.
resolve_rsids
resolve_rsids(rsids: list[str]) -> dict[str, list[dict]]
rsid -> [{chrom, start, ref, alts}, ...], shaped exactly like the cache links' output.
A multi-allelic rsID cannot be looked up by rsID: gnomAD answers rs334 with "Multiple
variants found, query using variant ID to select one." rather than a record. That is not an
error to swallow — rs334 is sickle-cell, a variant a module is very likely to carry — so those
rsIDs are collected and retried through variant_search, which does return every matching
variant id, and the alleles are aggregated into one locus. An rsID that returns nothing at all
is simply absent from the result (the caller falls through to "not found").
Source code in enricher/src/just_dna_enricher/gnomad.py
| def resolve_rsids(self, rsids: list[str]) -> dict[str, list[dict]]:
"""`rsid -> [{chrom, start, ref, alts}, ...]`, shaped exactly like the cache links' output.
A **multi-allelic rsID cannot be looked up by rsID**: gnomAD answers `rs334` with "Multiple
variants found, query using variant ID to select one." rather than a record. That is not an
error to swallow — `rs334` is sickle-cell, a variant a module is very likely to carry — so those
rsIDs are collected and retried through `variant_search`, which does return every matching
variant id, and the alleles are aggregated into one locus. An rsID that returns nothing at all
is simply absent from the result (the caller falls through to "not found").
"""
wanted = dedupe(r for r in rsids if r)
if not wanted:
return {}
out: dict[str, list[dict]] = {}
ambiguous: list[str] = []
for batch in batched(wanted, self.settings.batch_size):
fields = [
f"{_alias(i)}: variant(rsid: {_gql_string(rsid)}, dataset: {self.settings.dataset}) "
f"{{ variant_id }}"
for i, rsid in enumerate(batch)
]
data, errors = self._post_batch(fields, what="rsid resolution")
for i, rsid in enumerate(batch):
alias = _alias(i)
if _MULTIPLE_VARIANTS_RE.search(errors.get(alias, "")):
ambiguous.append(rsid)
continue
node = data.get(alias)
if node and node.get("variant_id"):
out[rsid] = _loci_from_variant_ids([node["variant_id"]])
for batch in batched(ambiguous, self.settings.batch_size):
fields = [
f"{_alias(i)}: variant_search(query: {_gql_string(rsid)}, "
f"dataset: {self.settings.dataset}) {{ variant_id }}"
for i, rsid in enumerate(batch)
]
data, _ = self._post_batch(fields, what="rsid search")
for i, rsid in enumerate(batch):
hits = data.get(_alias(i)) or []
ids = [h.get("variant_id") for h in hits if h.get("variant_id")]
if ids:
out[rsid] = _loci_from_variant_ids(ids)
return out
|
fetch_frequencies
fetch_frequencies(
variant_ids: list[str],
) -> dict[str, dict]
variantId -> {populations: [...], caid, vrs_id} for already-resolved coordinates.
The multi-allelic-rsID problem does not arise here: this pass keys on chrom-pos-ref-alt,
which names exactly one allele. Each population entry carries AC/AN — never a frequency, which
the API does not expose per group and which we compute from the counts instead.
Source code in enricher/src/just_dna_enricher/gnomad.py
| def fetch_frequencies(self, variant_ids: list[str]) -> dict[str, dict]:
"""`variantId -> {populations: [...], caid, vrs_id}` for already-resolved coordinates.
The multi-allelic-rsID problem does not arise here: this pass keys on `chrom-pos-ref-alt`,
which names exactly one allele. Each population entry carries AC/AN — never a frequency, which
the API does not expose per group and which we compute from the counts instead.
"""
wanted = dedupe(v for v in variant_ids if v)
if not wanted:
return {}
subset = self.settings.frequency_set
out: dict[str, dict] = {}
for batch in batched(wanted, self.settings.batch_size):
fields = [
f"{_alias(i)}: variant(variantId: {_gql_string(vid)}, dataset: {self.settings.dataset}) "
f"{{ variant_id caid rsids vrs {{ _id }} "
f"{subset} {{ ac an homozygote_count hemizygote_count "
f"faf95 {{ popmax popmax_population }} "
f"populations {{ id ac an homozygote_count hemizygote_count }} }} }}"
for i, vid in enumerate(batch)
]
data, _ = self._post_batch(fields, what="frequency")
for i, vid in enumerate(batch):
node = data.get(_alias(i))
if not node:
continue
joint = node.get(subset) or {}
rsids = node.get("rsids") or []
out[vid] = {
"populations": _populations_from_joint(joint),
"caid": node.get("caid"),
"vrs_id": (node.get("vrs") or {}).get("_id"),
# One rsID only when gnomAD reports exactly one: several means the id is
# position-level over multiple alleles, and picking one would mis-attribute it.
"rsid": rsids[0] if len(rsids) == 1 else None,
}
return out
|
fetch_gene_constraint
fetch_gene_constraint(
symbols: list[str],
) -> dict[str, dict]
gene symbol -> constraint metrics dict (empty entry omitted).
A module has tens of genes at most, so this is one or two batched requests — which is why the
gene pass can fall back to the live API without the rate limit mattering, while the frequency
pass cannot go the other way and ship an offline snapshot.
Source code in enricher/src/just_dna_enricher/gnomad.py
| def fetch_gene_constraint(self, symbols: list[str]) -> dict[str, dict]:
"""`gene symbol -> constraint metrics dict` (empty entry omitted).
A module has tens of genes at most, so this is one or two batched requests — which is why the
gene pass can fall back to the live API without the rate limit mattering, while the frequency
pass cannot go the other way and ship an offline snapshot.
"""
wanted = dedupe(s for s in symbols if s)
if not wanted:
return {}
out: dict[str, dict] = {}
for batch in batched(wanted, self.settings.batch_size):
fields = [
f"{_alias(i)}: gene(gene_symbol: {_gql_string(symbol)}, "
f"reference_genome: {self.settings.reference_genome}) "
f"{{ gene_id symbol mane_select_transcript {{ ensembl_id }} "
f"gnomad_constraint {{ pli oe_lof oe_lof_lower oe_lof_upper oe_mis "
f"lof_z mis_z syn_z obs_lof exp_lof flags }} }}"
for i, symbol in enumerate(batch)
]
data, _ = self._post_batch(fields, what="gene constraint")
for i, symbol in enumerate(batch):
node = data.get(_alias(i))
if not node:
continue
constraint = node.get("gnomad_constraint")
if not constraint:
continue
mane = node.get("mane_select_transcript") or {}
flags = constraint.get("flags") or []
out[symbol] = {
"gene": node.get("symbol") or symbol,
"gene_id": node.get("gene_id"),
"transcript": mane.get("ensembl_id"),
# The live route returns the gene's MANE Select transcript alongside its
# constraint, so when both are present the metrics are the MANE ones by
# construction — unlike the bulk TSV, where the row has to be picked.
"mane_select": bool(mane.get("ensembl_id")) or None,
"pli": constraint.get("pli"),
"loeuf": constraint.get("oe_lof_upper"),
"oe_lof": constraint.get("oe_lof"),
"oe_lof_lower": constraint.get("oe_lof_lower"),
"lof_z": constraint.get("lof_z"),
"mis_z": constraint.get("mis_z"),
"syn_z": constraint.get("syn_z"),
"oe_mis": constraint.get("oe_mis"),
"obs_lof": constraint.get("obs_lof"),
"exp_lof": constraint.get("exp_lof"),
"constraint_flags": normalize_constraint_flags(flags),
}
return out
|
covers_locus
covers_locus(
chrom: str | None,
start: int | None,
*,
build: str = "GRCh38",
) -> bool | None
Is this locus inside gnomAD's callset at all? Three-valued, and None means "cannot say".
gnomAD excludes the Y pseudoautosomal region: like a standard GRCh38 analysis set it hard-masks
the Y PAR, because those bases are a duplicate of the X PAR and reads there cannot be placed. Probed
live 2026-08-04 — region(chrom:"X", 640000-641500) serves 880 variants and the identical
interval on Y serves none, while a single-variant lookup for Y-640851-C-T returns
Variant not found where X-640851-C-T resolves.
That matters because "the source has no such allele" and "the source does not look here" are
different answers, and only the first is a fact. Without this, the one-to-many expansion of a PAR
rsID handed the frequency pass a Y locus, gnomAD said nothing, and the pass recorded not_found —
an absence nobody established. Three-valued so an unknown build withholds rather than asserting
coverage it cannot vouch for (PAR intervals are per-assembly — RM15).
This is a source convention and it lives here, in the network tier. The PAR geometry is an
assembly constant and belongs to the format tier (vrs.in_pseudoautosomal_region); "gnomAD masks
it" is a fact about gnomAD, and Principle 2 makes the enricher the only tier permitted to hold one.
Source code in enricher/src/just_dna_enricher/gnomad.py
| def covers_locus(chrom: str | None, start: int | None, *, build: str = "GRCh38") -> bool | None:
"""Is this locus inside gnomAD's callset at all? **Three-valued**, and `None` means "cannot say".
gnomAD excludes the **Y pseudoautosomal region**: like a standard GRCh38 analysis set it hard-masks
the Y PAR, because those bases are a duplicate of the X PAR and reads there cannot be placed. Probed
live 2026-08-04 — `region(chrom:"X", 640000-641500)` serves **880** variants and the identical
interval on Y serves **none**, while a single-variant lookup for `Y-640851-C-T` returns
`Variant not found` where `X-640851-C-T` resolves.
That matters because "the source has no such allele" and "the source does not look here" are
different answers, and only the first is a fact. Without this, the one-to-many expansion of a PAR
rsID handed the frequency pass a Y locus, gnomAD said nothing, and the pass recorded `not_found` —
an absence nobody established. Three-valued so an unknown build withholds rather than asserting
coverage it cannot vouch for (PAR intervals are per-assembly — RM15).
**This is a source convention and it lives here, in the network tier.** The PAR *geometry* is an
assembly constant and belongs to the format tier (`vrs.in_pseudoautosomal_region`); "gnomAD masks
it" is a fact about gnomAD, and Principle 2 makes the enricher the only tier permitted to hold one.
"""
if build != "GRCh38":
return None
in_par = in_pseudoautosomal_region(chrom, start, build=build)
# Left uncollapsed on purpose: `return not (...)` would read as a two-valued answer, and the
# `None` above is the third. Keeping the branches apart keeps the three outcomes visible.
if normalize_chrom(chrom) == "Y" and in_par is True: # noqa: SIM103
return False
return True
|