Skip to content

just_dna_enricher.grch37

just_dna_enricher.grch37

The old assembly: rs-number recovery and a wrong-build diagnosis (RM48).

An author curating from older literature has hg19/GRCh37 coordinates and the module must be GRCh38. Nothing in these four packages converts, so the conversion happens off-tool and lands as an ordinary authored coordinate carrying no provenance at all. This module is the online half of the answer; the offline half (a position past the end of its contig, a contig only one build names) lives in the compiler, where it needs neither network nor sequence.

Recovery, never liftover — and the reporter argued their own request down. If the paper gives an rs-number, liftover is unnecessary and strictly worse: authoring the rs-number produces the independent second value resolution._verify cross-examines. So liftover is only reachable where there is no rs-number and only an old coordinate — and in exactly that case the lifted coordinate becomes the row's sole identity with nothing to check it against, which is a generator of unverifiable-by-construction identities: the hazard class behind the 3,038-variant off-by-one this tree has already paid for once. Recovering the rs-number converts an unverifiable coordinate into a verifiable one, using machinery the enricher already has.

No chain file, no provisioned asset, no new licence. The roadmap's stated blocker was that rs-number recovery needs "either an hg19-keyed dbSNP surface or a chain file … i.e. the whole snapshot apparatus for one authoring convenience". Probed 2026-08-13 and false: Ensembl runs a permanent GRCh37 REST service at grch37.rest.ensembl.org, same API shape, serving both dbSNP variants (/overlap/region) and reference bases (/sequence/region). The same request that answers "which rs-numbers sit here on GRCh37" also discriminates the builds outright — 7:140453135..140453137 is CAC on GRCh37 and GTT on GRCh38.

Live-only, skipped offline. This is an authoring-time convenience; nothing in a compile depends on it, and a second whole-genome variant snapshot in the old assembly is exactly the apparatus worth not building. --offline reports skipped_offline, which is a first-class answer and never a pass.

It reports; it writes nothing. Filling the recovered rs-number would make resolution verify a value against the very service that produced it (hints.REDUNDANCY_BEARING, and rsid is identity_bearing in lookup._REFUSAL_BY_COLUMN). The author types the rs-number; a later enrich resolves it through the ordinary chain and resolution.csv's source column records which link answered. Provenance goes there, never into an authored coordinate.

Grch37Client dataclass

Grch37Client(
    settings: Grch37Settings = Grch37Settings(),
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

Paced, retrying access to the GRCh37 REST service. Every read is three-valued.

Hold one per session, the way LookupClients says to: the PacingGate state and the connection pool are both thrown away by a client built per question, and the gate is what keeps a panel-sized run inside Ensembl's published budget.

variants_at

variants_at(chrom: str, start: int) -> list[dict] | None

dbSNP variants overlapping chrom:start on GRCh37, or None for could-not-ask.

[] is an answer — the service was reached and records nothing there — while None means the question never completed. A 4xx counts as an answer: this service 400s on a contig it does not have and on a position past the end of one, and both of those really are "GRCh37 has nothing for you here", stated by the service rather than guessed at by us.

Only dbSNP features with an rs id are returned. The endpoint also serves HGMD-PUBLIC records (CM112509 at 7:140453136), and an id that is not an rs-number is not what this recovers.

Source code in enricher/src/just_dna_enricher/grch37.py
def variants_at(self, chrom: str, start: int) -> list[dict] | None:
    """dbSNP variants overlapping `chrom:start` on GRCh37, or `None` for could-not-ask.

    `[]` is an answer — the service was reached and records nothing there — while `None` means the
    question never completed. A 4xx counts as an answer: this service 400s on a contig it does not
    have and on a position past the end of one, and both of those really are "GRCh37 has nothing
    for you here", stated by the service rather than guessed at by us.

    Only dbSNP features with an `rs` id are returned. The endpoint also serves HGMD-PUBLIC records
    (`CM112509` at 7:140453136), and an id that is not an rs-number is not what this recovers.
    """
    try:
        response = self._get(
            f"/overlap/region/human/{chrom}:{start}-{start}",
            "application/json",
            {"feature": "variation"},
        )
    except httpx.HTTPStatusError as exc:
        if exc.response.status_code >= 500:
            logger.warning("GRCh37 overlap for %s:%d failed (%s)", chrom, start, exc)
            return None
        logger.info("GRCh37 has no region %s:%d (%s)", chrom, start, exc)
        return []
    except (httpx.TransportError, httpx.TimeoutException, httpx.HTTPError) as exc:
        logger.warning("GRCh37 overlap for %s:%d could not be reached: %s", chrom, start, exc)
        return None
    try:
        features = response.json()
    except ValueError as exc:
        # A 200 that is not JSON — a maintenance page — is the service answering with something
        # that cannot be read: could-not-ask, never "nothing there". Left raw it was a bare
        # `JSONDecodeError` out of a method whose whole contract is three-valued.
        logger.warning("GRCh37 overlap for %s:%d did not answer JSON: %s", chrom, start, exc)
        return None
    if not isinstance(features, list):
        logger.warning(
            "GRCh37 overlap for %s:%d answered %s, not a list", chrom, start, type(features).__name__
        )
        return None
    return [
        feature
        for feature in features
        if feature.get("assembly_name") == GRCH37_BUILD
        and feature.get("source") == "dbSNP"
        and str(feature.get("id", "")).startswith("rs")
    ]

reference_bases

reference_bases(
    chrom: str, start: int, end: int
) -> str | None

The GRCh37 bases over the 1-based inclusive span chrom:start..end. Three-valued.

Bases are an answer; "" is also an answer — the service was reached and has no sequence there — and None means the question never completed. Collapsing the middle case into None is precisely the S20 inversion this module's docstring cites as its design rule, and it is reachable with ordinary data rather than only under an outage: 20:63500000 sits inside GRCh38's chromosome 20 and past the end of GRCh37's, so the wrong-build case itself 400s here. Reported as "could not be asked", it would tell an author to re-run for an answer the service had already given.

start/end are the 1-based VCF convention this whole codebase stores, passed through unconverted — the instinctive -1 is the off-by-one that cost 3,038 rows once already.

Source code in enricher/src/just_dna_enricher/grch37.py
def reference_bases(self, chrom: str, start: int, end: int) -> str | None:
    """The GRCh37 bases over the 1-based inclusive span `chrom:start..end`. **Three-valued.**

    Bases are an answer; `""` is *also* an answer — the service was reached and has no sequence
    there — and `None` means the question never completed. Collapsing the middle case into `None`
    is precisely the S20 inversion this module's docstring cites as its design rule, and it is
    reachable with ordinary data rather than only under an outage: `20:63500000` sits inside
    GRCh38's chromosome 20 and past the end of GRCh37's, so **the wrong-build case itself** 400s
    here. Reported as "could not be asked", it would tell an author to re-run for an answer the
    service had already given.

    `start`/`end` are the 1-based VCF convention this whole codebase stores, passed through
    unconverted — the instinctive `-1` is the off-by-one that cost 3,038 rows once already.
    """
    try:
        response = self._get(f"/sequence/region/human/{chrom}:{start}..{end}", "text/plain")
    except httpx.HTTPStatusError as exc:
        if exc.response.status_code >= 500:
            logger.warning("GRCh37 sequence for %s:%d-%d failed (%s)", chrom, start, end, exc)
            return None
        logger.info("GRCh37 has no sequence at %s:%d-%d (%s)", chrom, start, end, exc)
        return ""
    except (httpx.TransportError, httpx.TimeoutException, httpx.HTTPError) as exc:
        logger.warning("GRCh37 sequence for %s:%d-%d could not be reached: %s", chrom, start, end, exc)
        return None
    return response.text.strip().upper()

RsidRecovery dataclass

RsidRecovery(
    chrom: str,
    start: int,
    outcome: str,
    ref: str | None = None,
    alts: str | None = None,
    rsids: list[str] = list(),
    candidates: list[dict] = list(),
    note: str = "",
)

What the GRCh37 service said about one old coordinate. Advisory; nothing is written.

BuildDiagnosis dataclass

BuildDiagnosis(
    variant_key: str,
    chrom: str,
    start: int,
    claimed: str,
    reason: str,
    grch37_bases: str | None = None,
    rsids: list[str] = list(),
    shift_claimed: int | None = None,
)

One ref-mismatched row, and what GRCh37 says about the same coordinate.

Only produced for rows that already failed the reference-base check, which is what bounds the cost: verify_reference_alleles is the filter, and a module whose refs all agree with GRCh38 makes no request here at all.

BuildDiagnosisResult dataclass

BuildDiagnosisResult(
    diagnoses: list[BuildDiagnosis] = list(),
    not_checked: str | None = None,
    examined: int = 0,
    total: int = 0,
)

The wrong-build pass's answer, with a reason when it did not run and what it looked at.

An empty diagnoses list otherwise says two opposite things — "asked, and nothing pointed at another build" and "never asked" — which is the same defect clin_sig_not_checked exists to prevent. not_checked is None exactly when the pass ran.

examined/total are the denominator, published for the reason RM44 states: a finding that quantifies over a subset has to say what the subset was, or a reader takes it for the whole.

recover_rsid

recover_rsid(
    chrom: str,
    start: int,
    *,
    ref: str | None = None,
    alts: str | None = None,
    client: Grch37Client | None = None,
    offline: bool = False,
) -> RsidRecovery

Which rs-number GRCh37 dbSNP records at this old coordinate.

The match is anchored on the authored position: a candidate must start exactly at start, and must carry the authored ref (and every authored alt) when those are given. Anchoring is what keeps this honest — a variant that merely overlaps the position is a different event, and accepting one would hand back an rs-number for a variant the author never meant.

The consequence to know: an indel authored in VCF's padded spelling (POS on the base before the event) will not match Ensembl's unpadded record, and comes back none rather than wrong. That is RM31's one-indel-several-spellings problem, and this reports the shorter search rather than reconciling frames it was not given.

Source code in enricher/src/just_dna_enricher/grch37.py
def recover_rsid(
    chrom: str,
    start: int,
    *,
    ref: str | None = None,
    alts: str | None = None,
    client: Grch37Client | None = None,
    offline: bool = False,
) -> RsidRecovery:
    """Which rs-number GRCh37 dbSNP records at this old coordinate.

    The match is **anchored on the authored position**: a candidate must start exactly at `start`, and
    must carry the authored `ref` (and every authored alt) when those are given. Anchoring is what
    keeps this honest — a variant that merely *overlaps* the position is a different event, and
    accepting one would hand back an rs-number for a variant the author never meant.

    The consequence to know: an indel authored in VCF's padded spelling (POS on the base *before* the
    event) will not match Ensembl's unpadded record, and comes back `none` rather than wrong. That is
    RM31's one-indel-several-spellings problem, and this reports the shorter search rather than
    reconciling frames it was not given.
    """
    query = RsidRecovery(chrom=str(chrom), start=int(start), outcome="none", ref=ref, alts=alts)
    if offline:
        query.outcome = "skipped_offline"
        return query
    owned = client is None
    client = client or Grch37Client()
    try:
        features = client.variants_at(query.chrom, query.start)
    finally:
        if owned:
            client.close()
    if features is None:
        query.outcome = "unchecked"
        return query

    wanted_ref = ref.strip().upper() if ref else None
    wanted_alts = {a.strip().upper() for a in (alts or "").split(",") if a.strip()}
    matched: list[dict] = []
    for feature in features:
        alleles = [str(a).strip().upper() for a in feature.get("alleles") or []]
        if not alleles or feature.get("start") != query.start:
            continue
        if wanted_ref is not None and alleles[0] != wanted_ref:
            continue
        if wanted_alts and not wanted_alts <= set(alleles[1:]):
            continue
        matched.append(
            {
                "rsid": str(feature["id"]),
                "start": feature.get("start"),
                "end": feature.get("end"),
                "alleles": alleles,
            }
        )
    # Sorted, never first-met: the emitted order is what a caller reads back and what a test compares,
    # and Ensembl does not promise one (Principle 7's ordering rule).
    query.candidates = sorted(matched, key=lambda c: c["rsid"])
    query.rsids = [c["rsid"] for c in query.candidates]
    if len(query.rsids) == 1:
        query.outcome = "recovered"
    elif len(query.rsids) > 1:
        query.outcome = "ambiguous"
    else:
        query.outcome = "none"
        query.note = (
            "The search is anchored on the exact position, so an indel authored in VCF's padded "
            "spelling (POS on the base before the event) can miss a record that is really there."
            if not features
            else f"{len(features)} dbSNP record(s) overlap the position, none starting at it with "
            f"the alleles given."
        )
    return query

diagnose_wrong_build

diagnose_wrong_build(
    mismatches: Sequence[RefMismatch],
    *,
    client: Grch37Client | None = None,
    offline: bool = False,
    limit: int = DEFAULT_DIAGNOSIS_LIMIT,
) -> BuildDiagnosisResult

Do these ref-mismatched rows read as GRCh37 coordinates in a GRCh38 module?

Runs only on rows the reference-base check already rejected, which is the whole cost control: the expensive question is asked of a set someone else has already narrowed to the suspicious rows.

The asymmetry is worth stating rather than implying generality: verify_reference_alleles skips any build refget_accession has no table for, so a RefMismatch only ever exists for a GRCh38 module — which makes "the other assembly" always GRCh37 here, not a parameter.

Three tiers of evidence, and the message says which one it has. Neither the mismatch nor this diagnosis is repaired: an authored ref and an authored coordinate are the evidence that something upstream went wrong, and rewriting either destroys it.

Source code in enricher/src/just_dna_enricher/grch37.py
def diagnose_wrong_build(
    mismatches: Sequence[RefMismatch],
    *,
    client: Grch37Client | None = None,
    offline: bool = False,
    limit: int = DEFAULT_DIAGNOSIS_LIMIT,
) -> BuildDiagnosisResult:
    """Do these ref-mismatched rows read as GRCh37 coordinates in a GRCh38 module?

    **Runs only on rows the reference-base check already rejected**, which is the whole cost control:
    the expensive question is asked of a set someone else has already narrowed to the suspicious rows.

    The asymmetry is worth stating rather than implying generality: `verify_reference_alleles` skips
    any build `refget_accession` has no table for, so a `RefMismatch` only ever exists for a GRCh38
    module — which makes "the other assembly" always GRCh37 here, not a parameter.

    Three tiers of evidence, and the message says which one it has. Neither the mismatch nor this
    diagnosis is repaired: an authored `ref` and an authored coordinate are the evidence that
    something upstream went wrong, and rewriting either destroys it.
    """
    if offline:
        return BuildDiagnosisResult(not_checked="skipped_offline", total=len(mismatches))
    if not mismatches:
        return BuildDiagnosisResult(not_checked="no_ref_mismatches")
    owned = client is None
    client = client or Grch37Client()
    examined = mismatches[:limit]
    diagnoses: list[BuildDiagnosis] = []
    try:
        for mismatch in examined:
            width = len(mismatch.claimed)
            bases = client.reference_bases(mismatch.chrom, mismatch.start, mismatch.start + width - 1)
            if bases is None:
                diagnoses.append(
                    BuildDiagnosis(
                        variant_key=mismatch.variant_key,
                        chrom=mismatch.chrom,
                        start=mismatch.start,
                        claimed=mismatch.claimed,
                        reason="unchecked",
                    )
                )
                continue
            if bases != mismatch.claimed:
                # GRCh37 does not explain this row either — including `""`, where GRCh37 answered that
                # it has no such place at all (a contig it lacks, or a position past the end of one).
                # Withhold: the mismatch stands on its own, and inventing a build hypothesis with no
                # evidence would be a false accusation.
                continue
            recovery = recover_rsid(mismatch.chrom, mismatch.start, ref=mismatch.claimed, client=client)
            reason = (
                "dbsnp_corroborated"
                if recovery.rsids
                else ("multi_base_match" if width > 1 else "single_base_match")
            )
            diagnoses.append(
                BuildDiagnosis(
                    variant_key=mismatch.variant_key,
                    chrom=mismatch.chrom,
                    start=mismatch.start,
                    claimed=mismatch.claimed,
                    reason=reason,
                    grch37_bases=bases,
                    rsids=list(recovery.rsids),
                    shift_claimed=mismatch.shift,
                )
            )
    finally:
        if owned:
            client.close()
    return BuildDiagnosisResult(diagnoses=diagnoses, examined=len(examined), total=len(mismatches))

summarize_build_diagnoses

summarize_build_diagnoses(
    diagnoses: Sequence[BuildDiagnosis],
    *,
    examples: int = 3,
) -> list[str]

One line per reason class, with a count, a few named rows, and which reading wins.

Grouped by reason and never by row: a panel authored from pre-GRCh38 literature produces one of these per variant, and the shared cause is the only actionable thing in the pile.

Source code in enricher/src/just_dna_enricher/grch37.py
def summarize_build_diagnoses(diagnoses: Sequence[BuildDiagnosis], *, examples: int = 3) -> list[str]:
    """One line per reason class, with a count, a few named rows, and which reading wins.

    Grouped by *reason* and never by row: a panel authored from pre-GRCh38 literature produces one of
    these per variant, and the shared cause is the only actionable thing in the pile.
    """
    grouped: dict[str, list[BuildDiagnosis]] = {}
    for diagnosis in diagnoses:
        grouped.setdefault(diagnosis.reason, []).append(diagnosis)
    lines: list[str] = []
    for reason, found in grouped.items():
        named = ", ".join(
            f"{d.chrom}:{d.start}" + (f" → {d.rsids[0]}" if d.rsids else "") for d in found[:examples]
        )
        more = f", and {len(found) - examples} more" if len(found) > examples else ""
        shifted = [d for d in found if d.shift_claimed is not None]
        supersedes = (
            f" {len(shifted)} of these were also read as a ±1 coordinate shift, and this supersedes "
            f"that: a neighbouring base equal to the authored ref is a one-in-four coincidence, and "
            f"the true variant is hundreds of bases away rather than one."
            if shifted and reason in _SUPERSEDES_SHIFT
            else ""
        )
        lines.append(f"{len(found)} row(s) — {_DIAGNOSIS_NOTES[reason]} ({named}{more}).{supersedes}")
    return lines