Skip to content

just_dna_enricher.mane_build

just_dna_enricher.mane_build

Build the MANE snapshot ([dev]) — the numbering frame, as a cache instead of a sentence (RM168).

MANE (Matched Annotation from NCBI and EMBL-EBI) publishes one agreed transcript per protein-coding gene, matched base-for-base between a RefSeq and an Ensembl accession. This workspace already leaned on it: the CIViC identity protocol pins a numbering frame with MANE.GRCh38.v1.5.summary.txt.gz, "downloaded once and cited" — a procedure step in a probe document, where every other reference table here is a cache with a location and a recorded release. The 33 curated name→identity answers that shipped with that protocol were derived in a frame nothing in the code could read. This builder is that frame, cached and pinned.

MANE is the default, not the answer, and the bound is measured rather than asserted. The summary carries CDKN2A twice for one gene — NM_000077.5 MANE Select beside NM_058195.4 MANE Plus Clinical, with different CDS numbering — and only ~0.4 % of rows are MANE Plus Clinical, which is exactly the class a remembered accession hides and a table shows. RUNX1, by contrast, is a single row, so the 27-residue RUNX1c/RUNX1b offset the identity protocol derived by translating each isoform's CDS is not in MANE and cannot be. The table therefore makes the CDKN2A class of problem visible and is silent on the RUNX1 class, and a pass that treated it as an oracle would be wrong in a way the file itself cannot warn about. MANE_status is carried as a column and never collapsed; a builder keeping one row per gene would drop the Plus Clinical rows and reintroduce the blind spot this item closes.

Scope is a transcript-identity aid. Generating c./p. notation is a separately deferred feature with its own unanswered questions, and nothing here proposes it.

Three files, because the third of them is the currency check and the second is the negative roster. Shipping the summary without the thing that notices it going stale is the defect this item is about, and all three together are under 1.2 MB:

  • summary — one row per MANE transcript; the cross-map (NCBI GeneID, Ensembl gene, HGNC id, symbol, both nuc/prot accession pairs, GRCh38 coordinates) as well as the frame.
  • changed_select_accessions — every gene whose MANE Select moved, with the release it moved from and Update_Affects_CDS: the numbering-frame axis, stated by the source. A gene absent from this table has had a stable frame, which is a positive statement the cache can make and a memory cannot.
  • protein_coding_genes_not_in_mane — the genes MANE deliberately has no answer for, each with a reason. @unreachable-not-absent served by the source: "MANE has no answer for this gene" is distinguishable from "nobody asked", with the reason attached, and pending MANE review is a third state that is neither. The reason vocabulary is derived from the file and recorded in release.json; it is not restated here, because a seven-member list written down beside the file is a registry nothing iterates (@registry-completeness).

The versioned directory is pinned, never current/. release_1.5/ back to release_0.5/ all exist, so a pin is a URL rather than a hope. current/ is still the right place to discover the newest version — its README_versions.txt names it — and resolving that to a versioned URL before downloading anything is the honest shape.

release.json copies README_versions.txt rather than restating it. That file is 96 bytes and publishes MANE Version, NCBI RefSeq Annotation Release and Ensembl Release; two of the three appear in no filename, so parsing a filename would reconstruct less information than the source hands over (@probe-the-real-file). It is parsed generically, label by label, so a line MANE adds travels through instead of being dropped by a reader that knew three names.

Three parquets under data/, keyed by filename, on the CPIC precedent. CPIC's snapshot already holds five distinct schemas in one data/ directory; the glob-union hazard (@snapshot-layout-locations) belongs to readers that union data/*.parquet into a single view, and this snapshot ships no reader — anything reading it later reads by filename.

Terms: NCBI publishes a policy, not a licence (@no-named-licence, and see MANE_TERMS). MANE is a joint NCBI/EMBL-EBI product and only NCBI's side was read, so nothing here asserts anything about EMBL-EBI's terms for the same tables.

Builder-only: polars is a guarded [dev] import, exactly as in the sibling builders.

ManeTable dataclass

ManeTable(name: str, source_suffix: str, parquet: str)

One of the three files a MANE release publishes that this snapshot carries.

A registry rather than three sets of parallel constants: the downloader, the builder, the CLI and release.json all walk it, so adding a fourth file is one entry instead of four edits (@registry-completeness).

ManeBuildError

Bases: RuntimeError

A MANE snapshot could not be built from the files given.

ManeUnavailable

Bases: ManeBuildError

A MANE release file could not be reached.

A subclass, so except ManeBuildError keeps catching everything it did while a caller who wants to tell "NCBI is unreachable" from "the file you handed me is malformed" can. The distinction is carried by the type rather than by a message, because a reword must not be able to flip a caller's verdict (@client-exception-contract) — and the subclassing makes a caller's except order load-bearing: list this arm before its parent or it is dead code.

ManeDownload dataclass

ManeDownload(
    path: Path,
    sha256: str | None,
    url: str | None = None,
    etag: str | None = None,
    last_modified: str | None = None,
)

What a download established about one file's bytes — each half None where the server was silent. The ETag and Last-Modified are kept beside the sha256 because a versioned directory is supposed to be immutable, and recording all three is what would turn a revision of one into a finding rather than a silent change of answer.

ManeBuildResult dataclass

ManeBuildResult(
    out_dir: Path,
    parquet_files: dict[str, Path] = dict(),
    rows: dict[str, int] = dict(),
    release: str | None = None,
    versions: dict[str, str] = dict(),
    dataset: str | None = None,
    mane_status_counts: dict[str, int] = dict(),
    update_affects_cds_counts: dict[str, int] = dict(),
    excluded_reasons: dict[str, int] = dict(),
    unparsable_gene_id: dict[str, int] = dict(),
    unparsable_update_affects_cds: int = 0,
    unparsable_coordinate: int = 0,
    source_sha256: dict[str, str | None] = dict(),
)

Outcome of a build: the paths, what each table holds, and the residue nothing could parse.

mane_release_dir_url

mane_release_dir_url(release: str) -> str

The versioned directory for a release, e.g. …/MANE_human/release_1.5.

Source code in enricher/src/just_dna_enricher/mane_build.py
def mane_release_dir_url(release: str) -> str:
    """The versioned directory for a release, e.g. `…/MANE_human/release_1.5`."""
    return f"{MANE_FTP_BASE}/release_{release}"

mane_release_url

mane_release_url(release: str, source_suffix: str) -> str

The URL of one release file.

MANE repeats the version in every filename (release_1.5/MANE.GRCh38.v1.5.summary.txt.gz), which is a shape a caller should not have to know — the same reason civic_build.civic_release_url exists.

Source code in enricher/src/just_dna_enricher/mane_build.py
def mane_release_url(release: str, source_suffix: str) -> str:
    """The URL of one release file.

    MANE repeats the version in every filename (`release_1.5/MANE.GRCh38.v1.5.summary.txt.gz`), which
    is a shape a caller should not have to know — the same reason `civic_build.civic_release_url`
    exists.
    """
    return f"{mane_release_dir_url(release)}/MANE.GRCh38.v{release}.{source_suffix}"

mane_versions_url

mane_versions_url(release: str | None = None) -> str

README_versions.txt in a pinned release, or in current/ when release is None.

The None case is the only read of the mutable path, and it exists so a build can discover which version is newest before pinning it. Nothing is downloaded from current/.

Source code in enricher/src/just_dna_enricher/mane_build.py
def mane_versions_url(release: str | None = None) -> str:
    """`README_versions.txt` in a pinned release, or in `current/` when `release` is `None`.

    The `None` case is the **only** read of the mutable path, and it exists so a build can discover
    which version is newest before pinning it. Nothing is downloaded from `current/`.
    """
    directory = MANE_CURRENT_DIRNAME if release is None else f"release_{release}"
    return f"{MANE_FTP_BASE}/{directory}/{MANE_VERSIONS_FILENAME}"

parse_versions

parse_versions(text: str) -> dict[str, str]

README_versions.txt as label -> value, verbatim and in file order.

Parsed generically rather than into three named fields: the file is a tab-separated label/value list, and a reader that knew exactly MANE Version, NCBI RefSeq Annotation Release and Ensembl Release would silently drop a fourth line MANE added. Blank lines and lines with no tab are skipped rather than raising — this is provenance, and a provenance failure is not a data failure.

Source code in enricher/src/just_dna_enricher/mane_build.py
def parse_versions(text: str) -> dict[str, str]:
    """`README_versions.txt` as `label -> value`, verbatim and in file order.

    Parsed generically rather than into three named fields: the file is a tab-separated label/value
    list, and a reader that knew exactly `MANE Version`, `NCBI RefSeq Annotation Release` and
    `Ensembl Release` would silently drop a fourth line MANE added. Blank lines and lines with no tab
    are skipped rather than raising — this is provenance, and a provenance failure is not a data
    failure.
    """
    versions: dict[str, str] = {}
    for line in text.splitlines():
        label, tab, value = line.partition("\t")
        if not tab:
            continue
        label, value = label.strip(), value.strip()
        if label:
            versions[label] = value
    return versions

parse_ncbi_gene_id

parse_ncbi_gene_id(raw: str | None) -> int | None

The numeric NCBI GeneID from either spelling MANE uses, or None.

The source spells its own key two ways, and the whole point of the three tables is that they join: summary writes GeneID:1029 while changed_select_accessions and protein_coding_genes_not_in_mane write a bare 1029. One normalizer, tested on both raw spellings (@one-normalizer-two-spellings), so a cross-table question — has this gene's frame ever moved? — is a join on one integer rather than a string comparison that silently never matches. None for anything else: unreadable, never guessed.

Source code in enricher/src/just_dna_enricher/mane_build.py
def parse_ncbi_gene_id(raw: str | None) -> int | None:
    """The numeric NCBI GeneID from either spelling MANE uses, or `None`.

    **The source spells its own key two ways**, and the whole point of the three tables is that they
    join: `summary` writes `GeneID:1029` while `changed_select_accessions` and
    `protein_coding_genes_not_in_mane` write a bare `1029`. One normalizer, tested on both raw
    spellings (`@one-normalizer-two-spellings`), so a cross-table question — *has this gene's frame
    ever moved?* — is a join on one integer rather than a string comparison that silently never
    matches. `None` for anything else: unreadable, never guessed.
    """
    text = (raw or "").strip().removeprefix("GeneID:").strip()
    return int(text) if text.isdigit() else None

parse_yes_no

parse_yes_no(raw: str | None) -> tuple[bool | None, bool]

(value, unparsable) for a Yes/No cell — three outcomes, as the house algebra requires.

An empty cell is a plain absence (None, not unparsable); anything that is neither Yes nor No is withheld and counted, because a token we cannot hold is a finding while silence is not.

Source code in enricher/src/just_dna_enricher/mane_build.py
def parse_yes_no(raw: str | None) -> tuple[bool | None, bool]:
    """`(value, unparsable)` for a `Yes`/`No` cell — three outcomes, as the house algebra requires.

    An empty cell is a plain absence (`None`, not unparsable); anything that is neither `Yes` nor
    `No` is withheld **and** counted, because a token we cannot hold is a finding while silence is
    not.
    """
    text = (raw or "").strip().casefold()
    if not text:
        return None, False
    if text == "yes":
        return True, False
    if text == "no":
        return False, False
    return None, True

download_mane_file

download_mane_file(dest: Path, url: str) -> ManeDownload

Stream one release file to dest (atomic .part rename), returning what it established.

Mirrors civic_build.download_civic_file, down to removing the partial file on failure: a .part left behind is the one residue a re-run would have to reason about.

Source code in enricher/src/just_dna_enricher/mane_build.py
def download_mane_file(dest: Path, url: str) -> ManeDownload:
    """Stream one release file to `dest` (atomic `.part` rename), returning what it established.

    Mirrors `civic_build.download_civic_file`, down to removing the partial file on failure: a
    `.part` left behind is the one residue a re-run would have to reason about.
    """
    streamed = stream_to_file(
        dest,
        url,
        error_cls=ManeUnavailable,
        what=f"the MANE file {Path(url).name}",
        timeout=120.0,
        remedy="Pass the local file instead of --download if you already hold a copy.",
    )
    return ManeDownload(
        path=streamed.path,
        sha256=streamed.sha256,
        url=url,
        etag=streamed.etag,
        last_modified=streamed.last_modified,
    )

discover_current_release

discover_current_release() -> str

The newest MANE version, read from current/README_versions.txt.

The one request that touches the mutable path, and it takes 96 bytes. Its answer is then resolved to a versioned URL before anything is downloaded, which is what makes a build pinnable after the fact: current/ moves, release_1.5/ does not.

Source code in enricher/src/just_dna_enricher/mane_build.py
def discover_current_release() -> str:
    """The newest MANE version, read from `current/README_versions.txt`.

    The **one** request that touches the mutable path, and it takes 96 bytes. Its answer is then
    resolved to a versioned URL before anything is downloaded, which is what makes a build pinnable
    after the fact: `current/` moves, `release_1.5/` does not.
    """
    url = mane_versions_url(None)
    try:
        response = httpx.get(url, timeout=60.0, follow_redirects=True)
        response.raise_for_status()
    except httpx.HTTPError as exc:
        raise ManeUnavailable(f"could not read {url}: {exc}") from exc
    release = parse_versions(response.text).get(MANE_VERSION_LABEL)
    if not release:
        raise ManeUnavailable(
            f"{url} does not state a {MANE_VERSION_LABEL!r}; refusing to guess a release from a "
            f"filename. Pass --release to pin one by hand."
        )
    return release

build_snapshot

build_snapshot(
    inputs: Mapping[str, Path],
    out_dir: Path,
    *,
    versions_file: Path | None = None,
    release: str | None = None,
    downloads: Mapping[str, ManeDownload] | None = None,
) -> ManeBuildResult

Convert one pinned MANE release into data/*.parquet + release.json.

inputs is keyed by MANE_TABLE_NAMES and must carry all three: the summary is the frame, the changed-accession list is the currency check, and the negative roster is what distinguishes "MANE has no answer for this gene" from "nobody asked". A snapshot missing any of them is the defect RM168 is about.

Rows are emitted in a total order (gene id, then accession), so a rebuild from the same inputs is byte-identical (Principle 7); release.json's built_at is the only per-run-varying byte and it lives outside the parquet.

release and every URL default to None rather than to this module's constants. Only a caller that actually fetched can say where the bytes came from, and only a README_versions.txt that was read can say which release they are. Defaulting either would write a provenance the build never established into the one file whose whole job is to pin it.

Source code in enricher/src/just_dna_enricher/mane_build.py
def build_snapshot(
    inputs: Mapping[str, Path],
    out_dir: Path,
    *,
    versions_file: Path | None = None,
    release: str | None = None,
    downloads: Mapping[str, ManeDownload] | None = None,
) -> ManeBuildResult:
    """Convert one pinned MANE release into `data/*.parquet` + `release.json`.

    `inputs` is keyed by `MANE_TABLE_NAMES` and must carry all three: the summary is the frame, the
    changed-accession list is the currency check, and the negative roster is what distinguishes "MANE
    has no answer for this gene" from "nobody asked". A snapshot missing any of them is the defect
    RM168 is about.

    Rows are emitted in a total order (gene id, then accession), so a rebuild from the same inputs is
    byte-identical (Principle 7); `release.json`'s `built_at` is the only per-run-varying byte and it
    lives outside the parquet.

    **`release` and every URL default to `None` rather than to this module's constants.** Only a
    caller that actually fetched can say where the bytes came from, and only a `README_versions.txt`
    that was read can say which release they are. Defaulting either would write a provenance the
    build never established into the one file whose whole job is to pin it.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the MANE snapshot; install the publisher/dev surface with "
            "`pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    missing = [name for name in MANE_TABLE_NAMES if name not in inputs]
    if missing:
        raise ManeBuildError(
            f"a MANE snapshot needs all three release files; missing {missing}. The summary alone is "
            f"a cache with nothing to notice it going stale, which is the defect this snapshot exists "
            f"to close."
        )
    paths = {name: Path(inputs[name]) for name in MANE_TABLE_NAMES}
    for name, path in paths.items():
        if not path.exists():
            raise FileNotFoundError(f"MANE {name} file not found at {path}")

    out_dir = Path(out_dir)
    data_dir = out_dir / SNAPSHOT_DATA_DIRNAME
    data_dir.mkdir(parents=True, exist_ok=True)

    result = ManeBuildResult(
        out_dir=out_dir,
        unparsable_gene_id=dict.fromkeys(MANE_TABLE_NAMES, 0),
    )
    result.versions = _read_versions(versions_file)
    result.release = result.versions.get(MANE_VERSION_LABEL) or release
    result.dataset = None if result.release is None else f"mane_grch38_v{result.release}"

    for table in MANE_TABLES:
        read, schema = _READERS[table.name]
        records = read(paths[table.name], result)
        frame = pl.DataFrame(records, schema=schema())
        parquet_file = data_dir / table.parquet
        frame.write_parquet(parquet_file, compression="zstd")
        result.parquet_files[table.name] = parquet_file
        result.rows[table.name] = frame.height

    given = downloads or {}
    result.source_sha256 = {
        name: (given[name].sha256 if name in given else _sha256_file(paths[name]))
        for name in MANE_TABLE_NAMES
    }
    # Always keyed, `None` when no `README_versions.txt` was read: a key that vanishes is a different
    # statement from a key whose value is null, and only the second is true here.
    result.source_sha256[MANE_VERSIONS_FILENAME] = (
        None if versions_file is None else _sha256_file(Path(versions_file))
    )

    _write_release_json(out_dir, result, downloads=given)
    logger.info(
        "Built the MANE snapshot (%s): %s → %s. %d gene(s) excluded over %d reason(s); %d changed "
        "MANE Select accession(s).",
        result.dataset or "release unknown",
        ", ".join(f"{name} {count}" for name, count in result.rows.items()),
        data_dir,
        result.rows["protein_coding_genes_not_in_mane"],
        len(result.excluded_reasons),
        result.rows["changed_select_accessions"],
    )
    return result