Skip to content

just_dna_enricher.mitomap_build

just_dna_enricher.mitomap_build

Build the MITOMAP snapshot the miss lane joins against (RM171).

The acquire stage is a plain curl of https://mitomap.org/downloads/mitomap.dump.sql.gz. That is worth stating because the web surface of the same host answers curl and the fetch tool with a Cloudflare 403 — which is why this source's terms had to be read from a browser — while the data surface has no interstitial at all. Two surfaces, two answers, and the negative belongs to the one it was measured on (@two-surfaces-two-denominators).

The build keeps six tables out of a 6.76 M-line dump and throws the other hundred away: the two curated variant tables, their two reference link tables, reference, and edit_date. Each variant row carries its source columns verbatim beside the four derived ones — the split status, the mapped clin_sig, the reason its alleles cannot be spelled as VCF where that applies, and the gene where locus names exactly one.

The dataset label comes from inside the dump, which is what makes a local build comparable. The dump carries edit_date, a per-table curation date (mMut, rtMut), and the label is both of them — the same choice ClinVar makes with ##fileDate rather than with Last-Modified. The header and the sha256 are recorded in release.json beside it, because provenance of the fetch is a different question from identity of the content, and a build from --dump <file> has the second and not the first.

DownloadedDump dataclass

DownloadedDump(
    path: Path,
    sha256: str,
    url: str,
    last_modified: str | None,
)

Where the dump landed and what the server said about it.

MitomapBuildResult dataclass

MitomapBuildResult(
    out_dir: Path,
    parquet_files: list[Path],
    rows: dict[str, int] = dict(),
    edit_dates: dict[str, str | None] = dict(),
    reference_rows: int = 0,
    references_with_pmid: int = 0,
    references_without_nlmid: int = 0,
    references_not_a_pmid: int = 0,
    citation_links: int = 0,
    unmintable: dict[str, int] = dict(),
    withheld_brackets: dict[str, int] = dict(),
    dataset: str | None = None,
    source_sha256: str | None = None,
    source_last_modified: str | None = None,
)

What the snapshot holds and what it came from.

download_mitomap_dump

download_mitomap_dump(
    dest: Path, url: str = DEFAULT_MITOMAP_URL
) -> DownloadedDump

Stream the MITOMAP dump to dest (atomic .part rename), hashing while streaming.

httpx's exceptions do not leave this function (@client-exception-contract): a Cloudflare 403 on a mistyped path has to reach the CLI as MITOMAP BUILD FAILED: … rather than as a raw HTTPStatusError, and the half-written .part goes with it so a failed run leaves the directory as it found it.

Source code in enricher/src/just_dna_enricher/mitomap_build.py
def download_mitomap_dump(dest: Path, url: str = DEFAULT_MITOMAP_URL) -> DownloadedDump:
    """Stream the MITOMAP dump to `dest` (atomic `.part` rename), hashing while streaming.

    **`httpx`'s exceptions do not leave this function** (`@client-exception-contract`): a Cloudflare
    403 on a mistyped path has to reach the CLI as `MITOMAP BUILD FAILED: …` rather than as a raw
    `HTTPStatusError`, and the half-written `.part` goes with it so a failed run leaves the directory
    as it found it.
    """
    streamed = stream_to_file(
        dest,
        url,
        error_cls=MitomapUnavailable,
        what="the MITOMAP dump",
    )
    return DownloadedDump(
        path=streamed.path,
        sha256=streamed.sha256,
        url=url,
        last_modified=streamed.last_modified,
    )

dataset_label

dataset_label(
    edit_dates: dict[str, str | None],
) -> str | None

mitomap_<mmutation date>+<rtmutation date> from the dump's own edit_date table.

In band, and that is the ClinVar precedent rather than a preference. ClinVar's dataset is the ##fileDate its VCF states about itself, not the Last-Modified its server states about the transfer — so the same label comes out whether the file was downloaded or handed over. The dump has the same property in edit_date, and taking it there is what lets mitomap build --dump from a copy on disk produce a snapshot that can be compared against a downloaded one. The header and the sha256 are still recorded in release.json; they are provenance of the fetch, which is a different question.

Two dates, because this lane adopts two tables and they are curated separately. A compound artifact gets a compound label — the same shape the derived miss lane uses for its two parents. Taking the later of the two would state that a rtmutation from August applies to an mmutation from October, and None where either is missing, because half a label is not a shorter label: it is one that cannot be compared.

Source code in enricher/src/just_dna_enricher/mitomap_build.py
def dataset_label(edit_dates: dict[str, str | None]) -> str | None:
    """`mitomap_<mmutation date>+<rtmutation date>` from the dump's own `edit_date` table.

    **In band, and that is the ClinVar precedent rather than a preference.** ClinVar's `dataset` is
    the `##fileDate` its VCF states about itself, not the `Last-Modified` its server states about the
    transfer — so the same label comes out whether the file was downloaded or handed over. The dump
    has the same property in `edit_date`, and taking it there is what lets `mitomap build --dump` from
    a copy on disk produce a snapshot that can be compared against a downloaded one. The header and
    the sha256 are still recorded in `release.json`; they are provenance of the *fetch*, which is a
    different question.

    **Two dates, because this lane adopts two tables and they are curated separately.** A compound
    artifact gets a compound label — the same shape the derived miss lane uses for its two parents.
    Taking the later of the two would state that a `rtmutation` from August applies to an `mmutation`
    from October, and `None` where either is missing, because half a label is not a shorter label: it
    is one that cannot be compared.
    """
    dates = [edit_dates.get(name) for name in VARIANT_TABLES]
    if not all(dates):
        return None
    return "mitomap_" + "+".join(str(date) for date in dates)

build_snapshot

build_snapshot(
    dump: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_MITOMAP_URL,
    source_sha256: str | None = None,
    source_last_modified: str | None = None,
) -> MitomapBuildResult

Turn a MITOMAP pg_dump into out_dir/data/*.parquet + release.json.

Rows are sorted by (start, ref, alt, record_id) per table so a rebuild from the same bytes is byte-identical (Principle 7); release.json's built_at is the only per-run-varying byte and lives outside the parquet.

A dump missing one of the six tables is refused, not built short. A snapshot silently lacking rtmutation would make every later miss count a lie about a table nobody looked at, and the single hardest thing to notice about a derived artifact is a denominator that changed.

Source code in enricher/src/just_dna_enricher/mitomap_build.py
def build_snapshot(
    dump: Path,
    out_dir: Path,
    *,
    source_url: str = DEFAULT_MITOMAP_URL,
    source_sha256: str | None = None,
    source_last_modified: str | None = None,
) -> MitomapBuildResult:
    """Turn a MITOMAP `pg_dump` into `out_dir/data/*.parquet` + `release.json`.

    Rows are sorted by `(start, ref, alt, record_id)` per table so a rebuild from the same bytes is
    byte-identical (Principle 7); `release.json`'s `built_at` is the only per-run-varying byte and
    lives outside the parquet.

    **A dump missing one of the six tables is refused, not built short.** A snapshot silently lacking
    `rtmutation` would make every later miss count a lie about a table nobody looked at, and the
    single hardest thing to notice about a derived artifact is a denominator that changed.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise ImportError(
            "polars is required to build the MITOMAP snapshot; install the publisher/dev surface "
            "with `pip install 'just-dna-enricher[dev]'` (or `uv sync --group dev`)."
        )
    dump = Path(dump)
    out_dir = Path(out_dir)
    tables = read_dump_tables(dump)
    missing = [name for name in DUMP_TABLES if name not in tables]
    if missing:
        raise MitomapError(
            f"{dump} carries no {', '.join(missing)} table. Refusing to build a snapshot that is "
            f"short a table rather than empty in one — a miss count derived from it would be "
            f"measured against a denominator nobody stated."
        )

    data_dir = out_dir / "data"
    data_dir.mkdir(parents=True, exist_ok=True)
    result = MitomapBuildResult(out_dir=out_dir, parquet_files=[])

    unmintable: Counter[str] = Counter()
    withheld: Counter[str] = Counter()
    for table in VARIANT_TABLES:
        cells = [_variant_cells(table, row) for row in tables[table]]
        for cell in cells:
            if cell["allele_defect"] is not None:
                unmintable[str(cell["allele_defect"])] += 1
            if cell["withheld_bracket"] is not None:
                withheld[str(cell["withheld_bracket"])] += 1
        frame = pl.DataFrame(cells, schema=_variant_schema()).sort(
            ["start", "ref", "alt", "record_id"], nulls_last=True
        )
        path = data_dir / VARIANT_PARQUET[table]
        frame.write_parquet(path)
        result.parquet_files.append(path)
        result.rows[table] = frame.height
    result.unmintable = dict(sorted(unmintable.items()))
    result.withheld_brackets = dict(sorted(withheld.items()))

    references = tables[REFERENCE_TABLE]
    pmid_by_reference: dict[str, str] = {}
    reference_cells: list[dict[str, object]] = []
    for row in references:
        ref_id = (row.get("id") or "").strip()
        pmid = nlmid_pmid(row.get("nlmid"))
        if pmid is not None and ref_id:
            pmid_by_reference[ref_id] = pmid
        raw = (row.get("nlmid") or "").strip()
        if not raw:
            result.references_without_nlmid += 1
        elif pmid is None:
            result.references_not_a_pmid += 1
        else:
            result.references_with_pmid += 1
        reference_cells.append(
            {
                "reference_id": ref_id or None,
                "authors": (row.get("authors") or "").strip() or None,
                "title": (row.get("title") or "").strip() or None,
                "publication": (row.get("publication") or "").strip() or None,
                "year": (row.get("date") or "").strip() or None,
                "nlmid": raw or None,
                "pmid": pmid,
            }
        )
    result.reference_rows = len(reference_cells)
    references_frame = pl.DataFrame(reference_cells, schema=_reference_schema()).sort(
        ["reference_id"], nulls_last=True
    )
    references_path = data_dir / REFERENCES_PARQUET
    references_frame.write_parquet(references_path)
    result.parquet_files.append(references_path)

    citations: list[dict[str, object]] = []
    for table, link_table in LINK_TABLES.items():
        id_column = f"{table}_id"
        for row in tables[link_table]:
            record_id = (row.get(id_column) or "").strip()
            reference_id = (row.get("reference_id") or "").strip()
            pmid = pmid_by_reference.get(reference_id)
            if not record_id or pmid is None:
                # A link into a reference with no PMID is a real link this tier cannot express as a
                # `StudyRow`. Dropped here and counted by the difference between `citation_links` and
                # the dump's own link total, which the drafter reports rather than implying that
                # every MITOMAP citation reached the module.
                continue
            citations.append(
                {
                    "table": table,
                    "record_id": record_id,
                    "reference_id": reference_id,
                    "pmid": pmid,
                }
            )
    citations_frame = (
        pl.DataFrame(citations, schema=_citation_schema())
        .unique(subset=["table", "record_id", "pmid"], keep="first")
        .sort(["table", "record_id", "pmid"])
    )
    citations_path = data_dir / CITATIONS_PARQUET
    citations_frame.write_parquet(citations_path)
    result.parquet_files.append(citations_path)
    result.citation_links = citations_frame.height

    dates = {
        (row.get("table_name") or "").strip(): (row.get("date") or "").strip() or None
        for row in tables[EDIT_DATE_TABLE]
    }
    result.edit_dates = {name: dates.get(EDIT_DATE_NAMES[name]) for name in VARIANT_TABLES}
    result.source_sha256 = source_sha256 or _sha256_file(dump)
    result.source_last_modified = source_last_modified
    result.dataset = dataset_label(result.edit_dates)
    _write_release_json(out_dir, result, source_url=source_url)
    logger.info(
        "Built MITOMAP snapshot: %s → %s",
        ", ".join(f"{name} {count}" for name, count in result.rows.items()),
        data_dir,
    )
    return result