Skip to content

just_dna_enricher.civic_build

just_dna_enricher.civic_build

Build the CIViC snapshot ([dev]) — derived, dated, and the first one that may be published.

CIViC (Griffith et al., Nat Genet 2017; civicdb.org) is a curated knowledgebase of variant interpretations in cancer. It is a source, of the same kind as ClinVar and PubMind: an authoritative annotation source. Nothing it produces may enter resolution.csv — CIViC states a clinical opinion about a locus something else resolved, and resolution.csv's authority column is a different word for a different thing (@source-vs-authority).

Two published surfaces disagree, and the dated download is the one a snapshot may use. The GraphQL API defaults to status: NON_REJECTED and serves 11,518 evidence items; the bulk ClinicalEvidenceSummaries TSV is 4,903 rows and every one is accepted. That is a 2.35x difference between two faces of one database, declared by neither. A snapshot has to be byte-reproducible from a pinned input, and only the download side has dated releases (01-Aug-2026/...), so this builder reads the TSV pair and records the basis explicitly. A figure from here is not comparable with a figure from the API, and release.json says so.

The direction axis is what CIViC has to offer here, and it is not clin_sig. Measured over the whole database, the germline subset carries five ACMG-tier calls and zero benign-class ones, so it can never make a clinical-significance disagreement sayable. What it does carry is Predisposition/Protectiveness crossed with Supports/Does Not Support — this format's direction axis (risk/protective). See docs/probes/CIVIC_SURVEY.md for every number.

"Does not support predisposition" is not "protective", and the null is the point. A refutation removes a claim; it does not establish the opposite one. So a Does Not Support row is kept, its raw words preserved, and its derived direction left null — the house's three-valued rule, where an unknown is withheld rather than negated. A drafter withholds those rows; it must not read them as protective.

Every drop is counted and the counts close. The origin filter alone removes about three quarters of the source, and a filter whose scope is narrower than its name is the defect this item was filed against. So input_rows == record_count + sum(dropped.values()) is asserted as an equality over a walked registry (@registry-completeness), and every reason lands in release.json (@dont-discard-computed).

Identity comes from what CIViC publishes, never from a liftover. Its coordinates are GRCh37 or absent — never GRCh38 — but the variant file also carries rsIDs and, for some records, RefSeq NC_ accessions on both builds. An rsID is build-independent and resolves through the ordinary chain, producing the independent second value resolution._verify cross-examines; a lifted coordinate would be the row's sole identity with nothing to check it against, which is what RM48 refused. So a record is kept when it carries an rsID or a GRCh38 accession, and dropped — counted — when it carries neither. allele_registry_id rides along so the dropped class stays addressable later.

Scoring an accession for build needs the per-chromosome map. NC_000001.11 is GRCh38 while NC_000002.11 is GRCh37: the version meaning "GRCh38" differs per chromosome. A first pass at this tested for ".11" or ".12" and overcounted reachable records five-fold.

Builder-only: polars is a guarded [dev] import, exactly as in the sibling builders.

CivicBuildError

Bases: RuntimeError

The build cannot proceed: a missing input, an unreadable file, a column that is not there.

CivicUnavailable

Bases: CivicBuildError

The release could not be fetched — a transport failure or a release that does not exist.

A subclass, so except CivicBuildError keeps catching everything it did, while a caller that wants to distinguish "the source did not answer" from "the source answered something we cannot build" can (@client-exception-contract). The subclassing makes a caller's except order load-bearing, which is why it is stated here rather than left to be discovered.

CivicDownload dataclass

CivicDownload(
    path: Path,
    sha256: str | None,
    url: str | None = None,
    etag: str | None = None,
    last_modified: str | None = None,
)

What a download established about one file's bytes — each half None when unstated.

CivicBuildResult dataclass

CivicBuildResult(
    out_dir: Path,
    parquet_file: Path,
    input_rows: int,
    record_count: int,
    dropped: dict[str, int] = dict(),
    unjoinable_submitted: int = 0,
    composite_profile_rows: int = 0,
    status_counts: dict[str, int] = dict(),
    status_basis: str = CIVIC_BULK_STATUS,
    vcf_evidence: dict[str, int] = dict(),
    curated_identities: dict[str, int] = dict(),
    identity_derivations: dict[str, int] = dict(),
    variants: int = 0,
    withheld_direction: int = 0,
    contested_variants: int = 0,
    unresolvable_with_caid: int = 0,
    unparsable_hgvs: int = 0,
    evidence_sha256: str | None = None,
    variant_sha256: str | None = None,
    profile_sha256: str | None = None,
    dataset: str | None = None,
)

Outcome of a build: the paths, the counts kept, and every count dropped.

civic_release_url

civic_release_url(release: str, filename: str) -> str

The URL of one file in a dated release.

CIViC's dated releases repeat the date in the filename (01-Aug-2026/01-Aug-2026-<file>), which is a shape a caller should not have to know.

Source code in enricher/src/just_dna_enricher/civic_build.py
def civic_release_url(release: str, filename: str) -> str:
    """The URL of one file in a dated release.

    CIViC's dated releases repeat the date in the filename (`01-Aug-2026/01-Aug-2026-<file>`), which
    is a shape a caller should not have to know.
    """
    return f"{CIVIC_DOWNLOAD_BASE}/{release}/{release}-{filename}"

download_civic_file

download_civic_file(dest: Path, url: str) -> CivicDownload

Stream one release file to dest (atomic .part rename).

Mirrors pubmind_build.download_pubmind_table, including keeping ETag and Last-Modified: a dated release should be immutable, and recording the headers is what would turn an upstream revision into a finding rather than a silent change of answer.

Source code in enricher/src/just_dna_enricher/civic_build.py
def download_civic_file(dest: Path, url: str) -> CivicDownload:
    """Stream one release file to `dest` (atomic `.part` rename).

    Mirrors `pubmind_build.download_pubmind_table`, including keeping `ETag` and `Last-Modified`: a
    dated release should be immutable, and recording the headers is what would turn an upstream
    revision into a finding rather than a silent change of answer.
    """
    streamed = stream_to_file(
        dest,
        url,
        error_cls=CivicUnavailable,
        what=f"the CIViC file {Path(url).name}",
        timeout=120.0,
    )
    return CivicDownload(
        path=streamed.path,
        sha256=streamed.sha256,
        url=url,
        etag=streamed.etag,
        last_modified=streamed.last_modified,
    )

parse_rsids

parse_rsids(aliases: str | None) -> list[str]

The rs-numbers in a variant_aliases cell, lowercased, in source order without duplicates.

CIViC's aliases are a comma-separated free-form list holding rs-numbers beside protein names and legacy labels, so the rs-numbers are selected by shape rather than by position.

Source code in enricher/src/just_dna_enricher/civic_build.py
def parse_rsids(aliases: str | None) -> list[str]:
    """The rs-numbers in a `variant_aliases` cell, lowercased, in source order without duplicates.

    CIViC's aliases are a comma-separated free-form list holding rs-numbers beside protein names and
    legacy labels, so the rs-numbers are selected by shape rather than by position.
    """
    seen: list[str] = []
    for token in re.split(r"[,\s]+", (aliases or "").strip()):
        if _RSID_RE.fullmatch(token) and token.lower() not in seen:
            seen.append(token.lower())
    return seen

variant_rsids

variant_rsids(variant: dict) -> list[str]

Every rs-number CIViC states for one variant, name first, then aliases.

Two fields, because CIViC uses both and neither is the documented one. The obvious place is variant_aliases, and a first version read only that — missing five variants in the germline direction set whose name is the rs-number itself (RS2736100, rs681673), sometimes with a protein alias beside it and sometimes with nothing. Those five needed no registry lookup and no conversion; the identity was published in plain sight, in the column a reader looks at first.

Name before aliases, because the name is the source's own primary label for the variant while an alias is a synonym. Order only decides which is stored as rsid; both are parsed either way.

Source code in enricher/src/just_dna_enricher/civic_build.py
def variant_rsids(variant: dict) -> list[str]:
    """Every rs-number CIViC states for one variant, name first, then aliases.

    **Two fields, because CIViC uses both and neither is the documented one.** The obvious place is
    `variant_aliases`, and a first version read only that — missing five variants in the germline
    direction set whose *name* is the rs-number itself (`RS2736100`, `rs681673`), sometimes with a
    protein alias beside it and sometimes with nothing. Those five needed no registry lookup and no
    conversion; the identity was published in plain sight, in the column a reader looks at first.

    Name before aliases, because the name is the source's own primary label for the variant while an
    alias is a synonym. Order only decides which is stored as `rsid`; both are parsed either way.
    """
    return parse_rsids(
        " , ".join(x for x in ((variant.get("variant") or ""), (variant.get("variant_aliases") or "")) if x)
    )

parse_grch38_substitution

parse_grch38_substitution(
    hgvs: str | None,
) -> tuple[str, int, str, str] | None

(chrom, start, ref, alt) from a GRCh38 genomic HGVS substitution, or None.

None covers three different things on purpose — no accession, a GRCh37 accession, and a GRCh38 accession in a form this parser does not read. The caller separates the third with has_unparsable_grch38, because "the source said nothing" and "the source said something we cannot hold" are different findings.

Source code in enricher/src/just_dna_enricher/civic_build.py
def parse_grch38_substitution(hgvs: str | None) -> tuple[str, int, str, str] | None:
    """`(chrom, start, ref, alt)` from a GRCh38 genomic HGVS substitution, or `None`.

    `None` covers three different things on purpose — no accession, a GRCh37 accession, and a GRCh38
    accession in a form this parser does not read. The caller separates the third with
    `has_unparsable_grch38`, because "the source said nothing" and "the source said something we
    cannot hold" are different findings.
    """
    for token in re.split(r"[,\s]+", (hgvs or "").strip()):
        match = _HGVS_SUB_RE.match(token)
        if match is None:
            continue
        accession, position, ref, alt = match.groups()
        chrom = _G38_ACCESSIONS.get(accession)
        if chrom is not None:
            return chrom, int(position), ref, alt
    return None

has_unparsable_grch38

has_unparsable_grch38(hgvs: str | None) -> bool

True when a GRCh38 accession is present but no substitution on it could be parsed.

Source code in enricher/src/just_dna_enricher/civic_build.py
def has_unparsable_grch38(hgvs: str | None) -> bool:
    """True when a GRCh38 accession is present but no substitution on it could be parsed."""
    if parse_grch38_substitution(hgvs) is not None:
        return False
    return any(
        token.split(":")[0] in _G38_ACCESSIONS
        for token in re.split(r"[,\s]+", (hgvs or "").strip())
        if ":" in token
    )

build_snapshot

build_snapshot(
    evidence_tsv: Path,
    variant_tsv: Path,
    profile_tsv: Path,
    out_dir: Path,
    *,
    release: str | None = None,
    evidence_sha256: str | None = None,
    variant_sha256: str | None = None,
    profile_sha256: str | None = None,
    vcf: Path | None = None,
    vcf_sha256: str | None = None,
) -> CivicBuildResult

Reduce the CIViC release pair to one parquet plus release.json.

Rows are emitted sorted by (chrom in karyotype order, start, ref, alt, variant_id, evidence_id), so a rebuild from the same release is byte-identical (Principle 7); release.json's built_at is the only per-run-varying byte and lives outside the parquet, exactly as in pubmind_build. Rows with no parsed GRCh38 coordinate sort after the placed ones, by variant_id — a deterministic position rather than wherever the dict landed them.

Every provenance argument defaults to None because only a caller that actually fetched can say where the bytes came from; a build off local disk records unknown rather than inventing a URL.

vcf widens the status basis and nothing else (RM169). Given the dated civic_accepted_and_submitted.vcf from the same release, every emitted row gains the evidence_status CIViC assigned it, and the submitted evidence items join the corpus. The TSV pair stays primary and every row is still built from TSV columns — the VCF contributes which items exist and their status, never an identity: its POS is GRCh37 and lifting it is refused (RM48). Omitted, the build is exactly what it was, on the accepted basis.

Source code in enricher/src/just_dna_enricher/civic_build.py
def build_snapshot(
    evidence_tsv: Path,
    variant_tsv: Path,
    profile_tsv: Path,
    out_dir: Path,
    *,
    release: str | None = None,
    evidence_sha256: str | None = None,
    variant_sha256: str | None = None,
    profile_sha256: str | None = None,
    vcf: Path | None = None,
    vcf_sha256: str | None = None,
) -> CivicBuildResult:
    """Reduce the CIViC release pair to one parquet plus `release.json`.

    Rows are emitted sorted by `(chrom in karyotype order, start, ref, alt, variant_id, evidence_id)`,
    so a rebuild from the same release is byte-identical (Principle 7); `release.json`'s `built_at` is
    the only per-run-varying byte and lives outside the parquet, exactly as in `pubmind_build`. Rows
    with no parsed GRCh38 coordinate sort after the placed ones, by `variant_id` — a deterministic
    position rather than wherever the dict landed them.

    Every provenance argument defaults to `None` because only a caller that actually fetched can say
    where the bytes came from; a build off local disk records unknown rather than inventing a URL.

    **`vcf` widens the status basis and nothing else (RM169).** Given the dated
    `civic_accepted_and_submitted.vcf` from the same release, every emitted row gains the
    `evidence_status` CIViC assigned it, and the submitted evidence items join the corpus. The TSV pair
    stays primary and every row is still built from TSV columns — the VCF contributes *which* items
    exist and their status, never an identity: its POS is GRCh37 and lifting it is refused (RM48).
    Omitted, the build is exactly what it was, on the `accepted` basis.
    """
    if pl is None:  # pragma: no cover - exercised only where the [dev] extra is absent
        raise CivicBuildError(
            "polars is required to build a snapshot; install the [dev] extra "
            "(`uv sync --extra dev`). Reading a built snapshot needs no such dependency."
        )
    evidence = _read_tsv(Path(evidence_tsv), _EVIDENCE_COLUMNS)
    variants = _read_tsv(Path(variant_tsv), _VARIANT_COLUMNS)
    profiles = _read_tsv(Path(profile_tsv), _PROFILE_COLUMNS)

    # The VCF half. Read before the loop so `evidence` is one list by the time anything walks it: a
    # submitted row that took a different code path from an accepted one would be a second parser,
    # and the drop registry could not close over both.
    dropped = dict.fromkeys(CIVIC_DROP_REASONS, 0)
    doid_by_disease = _doid_by_disease(evidence)
    status_basis = CIVIC_BULK_STATUS
    unjoinable_submitted = 0
    vcf_statuses: dict[str, int] = {}
    if vcf is not None:
        entries = read_vcf_entries(Path(vcf))
        assert_vocabulary_covers(entries)
        vcf_statuses = summarize(entries)
        status_basis = "accepted+submitted"
        by_variant_id = {
            (row.get("variant_id") or "").strip(): row
            for row in variants
            if (row.get("variant_id") or "").strip()
        }
        known = {int(row["evidence_id"]) for row in evidence if row["evidence_id"].isdigit()}
        for entry in entries:
            if entry.status != "submitted" or entry.evidence_id in known:
                continue
            variant = by_variant_id.get(str(entry.variant_id))
            if variant is None and (entry.allele_registry_id or entry.variant_aliases or entry.civic_hgvs):
                # `VariantSummaries.tsv` is accepted-only too, so most submitted evidence names a
                # variant it does not describe. The same CSQ entry carries the four identity cells the
                # TSV would have supplied, so the row is built from those and stamped `vcf_csq`.
                variant = _variant_row_from_csq(entry)
                by_variant_id[str(entry.variant_id)] = variant
                variants.append(variant)
            if variant is None:
                # A submitted item on a variant the TSV does not describe: no gene, no aliases, no
                # identity route, nothing to build from. Counted here rather than in the main loop
                # because the row never enters `evidence`, and a drop the walking loop cannot see is a
                # drop the registry cannot close over.
                # NOT counted in `dropped`: the registry's equality is over the rows that entered
                # `evidence`, and this item never did. Counting it there would make the input total
                # disagree with the list the loop walks — the guard below catches exactly that, and
                # catching it is what says the two halves are one accounting rather than two.
                unjoinable_submitted += 1
                continue
            evidence.append(_submitted_evidence_row(entry, variant, doid_by_disease))

    #: profile id → how many variants it names. A profile naming two is a combination genotype and
    #: is dropped as one; a profile this map has never heard of is a dangling reference.
    profile_arity: dict[str, int] = {
        (row.get("molecular_profile_id") or "").strip(): len(
            [v for v in re.split(r"[,\s]+", (row.get("variant_ids") or "").strip()) if v]
        )
        for row in profiles
    }
    by_profile: dict[str, dict[str, str]] = {
        row["single_variant_molecular_profile_id"]: row
        for row in variants
        if (row.get("single_variant_molecular_profile_id") or "").strip()
    }

    identity_derivations = dict.fromkeys(sorted(CIVIC_IDENTITY_DERIVATIONS), 0)
    curated = _classify_curated(variants)
    records: list[dict[str, object]] = []
    unparsable_hgvs = 0
    withheld_direction = 0
    unresolvable_caids: set[str] = set()

    for row in evidence:
        if row.get("variant_origin") not in CIVIC_GERMLINE_ORIGINS:
            dropped["non_germline_origin"] += 1
            continue
        significance = row.get("significance") or ""
        direction_raw = row.get("evidence_direction") or ""
        if (significance, direction_raw) not in CIVIC_DIRECTION_MAP:
            dropped["not_direction_axis"] += 1
            continue
        profile_id = (row.get("molecular_profile_id") or "").strip()
        variant = by_profile.get(profile_id)
        if variant is None:
            # The profile file is what separates these: a profile naming two or more variants is a
            # combination genotype, and anything else that fails to join is a dangling reference.
            # Inferring both from one failed join would report a real class and an integrity failure
            # under the same name.
            arity = profile_arity.get(profile_id, 0)
            dropped["combination_profile" if arity > 1 else "no_variant_record"] += 1
            continue

        rsids = variant_rsids(variant)
        coords = parse_grch38_substitution(variant.get("hgvs_descriptions"))
        if coords is None and has_unparsable_grch38(variant.get("hgvs_descriptions")):
            unparsable_hgvs += 1
        caid = (variant.get("allele_registry_id") or "").strip()
        caid = caid if caid and caid != "unregistered" else ""
        # The curated table is consulted only where the source itself says nothing, and only on an
        # exact name match. Both halves matter: `applied` below is the same predicate `_classify_
        # curated` counts with, so the registry closes, and a row CIViC has since filled never
        # reaches here at all.
        curated_row = _curated_for(variant) if not rsids and coords is None and not caid else None
        if not rsids and coords is None and not caid and curated_row is None:
            dropped["unresolvable_identity"] += 1
            continue

        if variant.get(_CSQ_SOURCED):
            # Stamped ahead of the published-identifier routes, and deliberately: those name *which*
            # identifier answered, while this names which **file** it was read from. For a variant the
            # TSV does not describe at all, the file is the fact a consumer cannot otherwise recover,
            # and the routes inside are visible in the row's own rsid/chrom/allele_registry_id cells.
            derivation = VCF_DERIVATION
            if caid and not rsids and coords is None:
                # A CAID-only row is one a later identity pass can recover, whichever file it came
                # from — so it belongs in this count. Accrued here as well as in the `else` branch
                # below, because taking the `vcf_csq` label first means such a row never reaches it,
                # and `unresolvable_with_caid` then understates exactly the class the number exists
                # to size (`@dont-discard-computed`).
                unresolvable_caids.add(caid)
        elif curated_row is not None:
            derivation = CURATED_DERIVATION
        elif rsids and coords is not None:
            derivation = "both"
        elif rsids:
            derivation = "rsid"
        elif coords is not None:
            derivation = "grch38_hgvs"
        else:
            derivation = "caid"
            unresolvable_caids.add(caid)
        identity_derivations[derivation] += 1

        direction = CIVIC_DIRECTION_MAP[(significance, direction_raw)]
        if direction is None:
            withheld_direction += 1

        if curated_row is not None:
            chrom, start, ref, alt = (curated_row.chrom, curated_row.start, curated_row.ref, curated_row.alt)
            rsids = [curated_row.rsid] if curated_row.rsid else rsids
        else:
            chrom, start, ref, alt = coords if coords is not None else (None, None, None, None)
        records.append(
            {
                "chrom": chrom,
                "start": start,
                "ref": ref,
                "alt": alt,
                "rsid": rsids[0] if rsids else None,
                "allele_registry_id": caid or None,
                "identity_derivation": derivation,
                "direction": direction,
                "significance_raw": significance,
                "evidence_direction_raw": direction_raw,
                "variant_id": int(variant["variant_id"]),
                "variant_name": variant.get("variant") or None,
                "gene": variant.get("gene") or None,
                "evidence_id": int(row["evidence_id"]),
                "molecular_profile_id": int(profile_id),
                # RM174. A TSV row's own profile IS `profile_id` — the TSV path drops a multi-variant
                # profile as `combination_profile` before it reaches here, so the two can only differ
                # on the VCF path, where a composite is fanned out and kept. The **name** is null for
                # a TSV row because `MolecularProfileSummaries.tsv` publishes none; filling it from
                # the variant's name would state a profile name the source never wrote.
                "evidence_molecular_profile_id": int(row.get("evidence_molecular_profile_id") or profile_id),
                "evidence_molecular_profile_name": (row.get("evidence_molecular_profile_name") or "").strip()
                or None,
                "evidence_level": row.get("evidence_level") or None,
                "rating": int(row["rating"]) if (row.get("rating") or "").isdigit() else None,
                "variant_origin": row.get("variant_origin") or None,
                "pmid": _pmid(row),
                "disease": row.get("disease") or None,
                "doid": row.get("doid") or None,
                "civic_grch37_chrom": _grch37_chrom(variant),
                "civic_grch37_start": _grch37_start(variant),
                # CIViC's own curation status, verbatim. On the `accepted` basis every row carries
                # `accepted` from the TSV's own column; on the wider basis it is what separates a row
                # an editor signed off from one a curator entered. Never translated into a house
                # grade: it is the source's instrument and naming it is the point (RM169).
                "evidence_status": (row.get("evidence_status") or "").strip() or None,
            }
        )

    composite_profile_rows = sum(
        1 for record in records if record["evidence_molecular_profile_id"] != record["molecular_profile_id"]
    )

    assert_registry_closes(len(evidence), len(records), dropped)
    assert_curation_closes(curated)

    records.sort(key=_sort_key)
    out_dir = Path(out_dir)
    data_dir = out_dir / SNAPSHOT_DATA_DIRNAME
    data_dir.mkdir(parents=True, exist_ok=True)
    parquet_file = data_dir / CIVIC_PARQUET
    frame = (
        pl.DataFrame(records, schema=_polars_schema()) if records else pl.DataFrame(schema=_polars_schema())
    )
    frame.write_parquet(parquet_file)

    camps: dict[int, set[str]] = {}
    for record in records:
        if record["direction"] is not None:
            camps.setdefault(int(record["variant_id"]), set()).add(str(record["direction"]))
    status_counts = dict(
        collections.Counter(str(r["evidence_status"]) for r in records if r["evidence_status"] is not None)
    )
    result = CivicBuildResult(
        out_dir=out_dir,
        parquet_file=parquet_file,
        input_rows=len(evidence),
        record_count=len(records),
        dropped=dropped,
        identity_derivations=identity_derivations,
        status_counts=status_counts,
        unjoinable_submitted=unjoinable_submitted,
        status_basis=status_basis,
        composite_profile_rows=composite_profile_rows,
        vcf_evidence=vcf_statuses,
        curated_identities=curated,
        variants=len({int(r["variant_id"]) for r in records}),
        withheld_direction=withheld_direction,
        contested_variants=sum(1 for v in camps.values() if len(v) > 1),
        unresolvable_with_caid=len(unresolvable_caids),
        unparsable_hgvs=unparsable_hgvs,
        evidence_sha256=evidence_sha256,
        variant_sha256=variant_sha256,
        profile_sha256=profile_sha256,
        dataset=f"civic_{release}" if release else None,
    )
    _write_release_json(out_dir, result, release=release)
    _write_license(out_dir)
    return result

assert_curation_closes

assert_curation_closes(states: dict[str, int]) -> None

Each curated row landed in exactly one state, and the states account for the whole table.

The same equality-over-a-walked-set the drop registry gets (@registry-completeness). A build that quietly stopped consulting the table would otherwise look identical to one where every row happened to be superseded.

Source code in enricher/src/just_dna_enricher/civic_build.py
def assert_curation_closes(states: dict[str, int]) -> None:
    """Each curated row landed in exactly one state, and the states account for the whole table.

    The same equality-over-a-walked-set the drop registry gets (`@registry-completeness`). A build
    that quietly stopped consulting the table would otherwise look identical to one where every row
    happened to be superseded.
    """
    total = sum(states.values())
    if total != len(CIVIC_NAME_IDENTITY_BY_VARIANT):
        raise CivicBuildError(
            f"the curated identity table does not close: {len(CIVIC_NAME_IDENTITY_BY_VARIANT)} rows, "
            f"{total} accounted for across {sorted(states)} (`@registry-completeness`)."
        )

assert_registry_closes

assert_registry_closes(
    input_rows: int, kept: int, dropped: dict[str, int]
) -> None

Every input row is either kept or counted under a reason, and nothing falls between.

An equality over the walked registry, never a floor (@registry-completeness). A filter added without a counter beside it would let the build truncate silently, and silent truncation reads as full coverage — which is the defect the whole drop registry exists to prevent.

Its own function so it can be exercised directly: a test that has to contrive a broken build in order to reach a guard usually ends up proving something else instead.

Source code in enricher/src/just_dna_enricher/civic_build.py
def assert_registry_closes(input_rows: int, kept: int, dropped: dict[str, int]) -> None:
    """Every input row is either kept or counted under a reason, and nothing falls between.

    An equality over the walked registry, never a floor (`@registry-completeness`). A filter added
    without a counter beside it would let the build truncate silently, and silent truncation reads as
    full coverage — which is the defect the whole drop registry exists to prevent.

    Its own function so it can be exercised directly: a test that has to contrive a broken build in
    order to reach a guard usually ends up proving something else instead.
    """
    total = sum(dropped.values())
    if input_rows != kept + total:
        raise CivicBuildError(
            f"the drop registry does not account for every input row: {input_rows} read, "
            f"{kept} kept, {total} dropped across {sorted(dropped)}. A row was filtered out without "
            f"a counter beside it (`@registry-completeness`)."
        )