Skip to content

just_dna_enricher.civic_draft

just_dna_enricher.civic_draft

Draft direction-axis rows from the CIViC snapshot (RM152).

The axis is direction, not clin_sig, and that is the whole reason this provider exists. S84 proposed CIViC as a clinical-significance authority and the measurement refused it: of CIViC's 3,103 germline evidence items, five carry an ACMG tier this format can receive and zero are benign-class, so a clinical-significance disagreement is unsayable. What the same measurement found is 1,458 germline items on Predisposition/Protectiveness — this format's direction (risk/protective). RM152 rejected a drafter partly because it would "write rows whose significance column is empty"; that is true of clin_sig, where 812 germline items are NA, and false of direction, where the NA count is zero. The rejection was measured on the axis the report aimed at rather than the one that survives, which is why this provider is that one.

A flag on the existing command, never a second command — draft-panel --source civic, the shape --source pubmind set. A separate command is for a provider writing different tables; this writes the ones every panel draft writes.

It reads the snapshot and never the network. civic build pins a dated release; this reads its parquet. That keeps the acquisition gate where it belongs and means a draft is reproducible from a named dataset rather than from whatever the API served that afternoon (@currency-asks-the-source-not-the-cache).

The skip guard is derived from VariantRow, never restated beside it. This is a recorded scar: pgx_draft restated the rule as "no rsID and no position" while the model wants rsID or chrom+start, and draft --gene CYP2C9 died on an unhandled pydantic error. CIViC supplies a live example of the shape that kills a restated guard — variant 1770 carries a build and a start and a referenceBases with no chromosome, which passes any "has a position?" test and is not a position. So the guard here asks the model, and the model's refusal is the answer.

A contested variant is withheld, never resolved. Where CIViC's own evidence puts a variant in both camps, picking one is mode() over an unsorted group. The rows stay in the snapshot, the variant gets no drafted row, and the count is reported.

Every excluded row is counted in the RESULT, not in this docstring. That was the reporter's own objection to a drafter and it is correct: a filter whose scope is narrower than its name hides what it removed. The snapshot already counted the somatic majority; this counts what it withholds on top.

CivicDraftError

Bases: RuntimeError

The draft cannot run: no snapshot, or one that is present and will not answer.

CivicDraftResult dataclass

CivicDraftResult(
    reports: list[DraftReport] = list(),
    warnings: list[str] = list(),
    withheld: dict[str, int] = dict(),
    candidates: int = 0,
    caid_resolved_by_rsid: int = 0,
    caid_resolved_by_coordinate: int = 0,
    caid_anchored_indels: int = 0,
    dataset: str | None = None,
    skipped: bool = False,
    refuted_beside_claim: list[
        tuple[str, str, str]
    ] = list(),
    refutation_basis: str | None = None,
)

What a CIViC draft did — and, in equal detail, what it did not write.

added property

added: int

Rows added across every table this run wrote — variants and their studies.

outcome_for

outcome_for(csv_name: str, outcome: str) -> int

How many variants.csv rows landed in one outcome — added, already_present, invalid.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def outcome_for(self, csv_name: str, outcome: str) -> int:
    """How many `variants.csv` rows landed in one outcome — added, already_present, invalid."""
    return sum(len(getattr(r, outcome)) for r in self.reports if r.csv_name == csv_name)

accounts_for_every_candidate

accounts_for_every_candidate() -> bool

Every admitted row is either drafted, already there, refused, or withheld by name.

An equality, not a floor. A provider that quietly drops a row it cannot handle looks exactly like one with nothing to say about it, and the difference is the whole of what a drafted module's author needs to know. invalid is inside the sum rather than outside it: a row the table refused is still a row this pass has to account for.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def accounts_for_every_candidate(self) -> bool:
    """Every admitted row is either drafted, already there, refused, or withheld by name.

    An equality, not a floor. A provider that quietly drops a row it cannot handle looks exactly
    like one with nothing to say about it, and the difference is the whole of what a drafted
    module's author needs to know. `invalid` is inside the sum rather than outside it: a row the
    table refused is still a row this pass has to account for.
    """
    landed = sum(
        self.outcome_for("variants.csv", outcome)
        for outcome in ("added", "already_present", "differs", "invalid")
    )
    withheld = sum(v for k, v in self.withheld.items() if k != "gene_not_requested")
    return self.candidates == landed + withheld

civic_dataset_label

civic_dataset_label(reference: Path | None) -> str | None

The snapshot's dataset, or None when there is no snapshot or it names none.

None is an unknown release, never a fabricated one — the same answer the sibling labels give.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def civic_dataset_label(reference: Path | None) -> str | None:
    """The snapshot's `dataset`, or `None` when there is no snapshot or it names none.

    `None` is an unknown release, never a fabricated one — the same answer the sibling labels give.
    """
    return _release_payload(reference).get("dataset")

identity_refused_by_model

identity_refused_by_model(cells: dict) -> str | None

None when VariantRow accepts these identity cells, else the model's own complaint.

Thin wrapper over drafting.identity_refused_by_model, kept because this name is what the withheld-reason vocabulary and this module's tests both use. The implementation moved to the scaffold in RM228, and with it went the message parsing: this function used to branch on "identifier" in message or "positional" in message or "chrom" in message, consuming pydantic's rendered text as an API — a string that moves on a dependency bump with nothing to notice. The probe now pre-fills every non-identity field with values the model is known to accept, so any ValidationError reaching it is an identity refusal and no inspection is needed.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def identity_refused_by_model(cells: dict) -> str | None:
    """`None` when `VariantRow` accepts these identity cells, else the model's own complaint.

    Thin wrapper over `drafting.identity_refused_by_model`, kept because this name is what the
    withheld-reason vocabulary and this module's tests both use. The implementation moved to the
    scaffold in RM228, and with it went **the message parsing**: this function used to branch on
    `"identifier" in message or "positional" in message or "chrom" in message`, consuming pydantic's
    rendered text as an API — a string that moves on a dependency bump with nothing to notice. The
    probe now pre-fills every non-identity field with values the model is known to accept, so any
    `ValidationError` reaching it **is** an identity refusal and no inspection is needed.
    """
    return scaffold_identity_refused(VariantRow, cells, _PROVIDER.table)

trait_curie

trait_curie(doid: str | None) -> str | None

CIViC's bare Disease Ontology number as an ontology CURIE, or None.

trait_efo_id is not an EFO-only column, and reading the name as though it were is a mistake this provider made once. The field takes any ontology CURIE — its own description says "EFO/MONDO/OBA/HP", the validator accepts any PREFIX:LOCAL token, and cells are multi-valued — so DOID:1612 belongs in it. The first version of this provider put the DOID in conclusion prose instead, reasoning that a DOID in an EFO column would be a wrong identifier; the premise was false, and the effect was to bury a structured id nothing could join on. Every CIViC germline direction row carries a DOID, so the cost was the whole column.

CIViC publishes the number bare (1612), which is not a CURIE; the prefix is added here rather than stored upstream, because a bare integer in this column would fail the validator and a reader cannot tell which ontology it came from.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def trait_curie(doid: str | None) -> str | None:
    """CIViC's bare Disease Ontology number as an ontology CURIE, or `None`.

    **`trait_efo_id` is not an EFO-only column**, and reading the name as though it were is a mistake
    this provider made once. The field takes any ontology CURIE — its own description says
    "EFO/MONDO/OBA/HP", the validator accepts any `PREFIX:LOCAL` token, and cells are multi-valued —
    so `DOID:1612` belongs in it. The first version of this provider put the DOID in `conclusion`
    prose instead, reasoning that a DOID in an EFO column would be a wrong identifier; the premise was
    false, and the effect was to bury a structured id nothing could join on. Every CIViC germline
    direction row carries a DOID, so the cost was the whole column.

    CIViC publishes the number bare (`1612`), which is not a CURIE; the prefix is added here rather
    than stored upstream, because a bare integer in this column would fail the validator and a reader
    cannot tell which ontology it came from.
    """
    doid = (doid or "").strip()
    return f"DOID:{doid}" if doid else None

draft_panel_from_civic

draft_panel_from_civic(
    spec_dir: Path,
    genes: Sequence[str] = (),
    *,
    snapshot: Path | None = None,
    declared_use: str = "unstated",
    offline: bool = False,
    registry: ClingenAlleleClient | None = None,
    dry_run: bool = False,
) -> CivicDraftResult

Append direction-axis rows from the CIViC snapshot, with everything withheld accounted for.

genes filters; empty drafts every variant in the snapshot. The gene filter is applied first and counted separately from the withholding, because "CIViC has nothing for this gene" and "CIViC has something and we would not write it" are different answers an author needs told apart.

The counts land on the result, never only in a log line. The snapshot already recorded the somatic majority it dropped; this records what it withheld on top, and accounts_for_every_candidate() is an equality over both.

Source code in enricher/src/just_dna_enricher/civic_draft.py
def draft_panel_from_civic(
    spec_dir: Path,
    genes: Sequence[str] = (),
    *,
    snapshot: Path | None = None,
    declared_use: str = "unstated",
    offline: bool = False,
    registry: ClingenAlleleClient | None = None,
    dry_run: bool = False,
) -> CivicDraftResult:
    """Append `direction`-axis rows from the CIViC snapshot, with everything withheld accounted for.

    `genes` filters; empty drafts every variant in the snapshot. The gene filter is applied first and
    counted separately from the withholding, because "CIViC has nothing for this gene" and "CIViC has
    something and we would not write it" are different answers an author needs told apart.

    **The counts land on the result, never only in a log line.** The snapshot already recorded the
    somatic majority it dropped; this records what it withheld on top, and
    `accounts_for_every_candidate()` is an equality over both.
    """
    result = CivicDraftResult(withheld=dict.fromkeys(CIVIC_WITHHELD_REASONS, 0))
    reference = snapshot if snapshot is not None else resolve_civic_reference()
    if reference is None:
        result.skipped = True
        result.warnings.append(
            "No CIViC snapshot found, so nothing was drafted from it. Build one with "
            "`just-dna-enricher civic build --release <date>`, or point at one with "
            "$JUST_DNA_CIVIC_CACHE. Nobody-asked is not the same as the source having nothing."
        )
        return result

    rows = _snapshot_rows(Path(reference))
    result.dataset = civic_dataset_label(Path(reference))
    wanted = {g.strip().upper() for g in genes if g.strip()}

    admitted: list[dict] = []
    for row in rows:
        if wanted and (row["gene"] or "").upper() not in wanted:
            result.withheld["gene_not_requested"] += 1
            continue
        admitted.append(row)
    result.candidates = len(admitted)

    # Contested is decided over the WHOLE admitted group, before any per-row filter runs: a variant
    # whose dissenting row a filter removed would read as uncontested, and the filter would have
    # picked the winner (`@filter-before-the-group-picks-a-winner`).
    camps: dict[int, set[str]] = {}
    for row in admitted:
        if row["direction"] is not None:
            camps.setdefault(int(row["variant_id"]), set()).add(str(row["direction"]))
    contested_ids = {vid for vid, seen in camps.items() if len(seen) > 1}
    contested_names = sorted(
        {str(r["variant_name"] or r["variant_id"]) for r in admitted if int(r["variant_id"]) in contested_ids}
    )

    # RM170 — the same group-first rule, one slot over. A refutation is **not** a camp: it withholds a
    # claim rather than making the opposite one, so it never enters `camps` and `contested_ids` is
    # correctly blind to it. What an author still needs to know is that the variant they are about to
    # be handed a `risk` row for is one the same snapshot also rebuts, so the pair is computed here,
    # over the whole admitted group, before any per-row filter can remove either side of it.
    refuting_evidence: dict[int, list[str]] = {}
    for row in admitted:
        if row["evidence_direction_raw"] == CIVIC_REFUTES:
            status = row.get("evidence_status")
            refuting_evidence.setdefault(int(row["variant_id"]), []).append(
                f"EID {row['evidence_id']}" + (f" ({status})" if status else "")
            )
    claimed_ids = {int(row["variant_id"]) for row in admitted if row["direction"] is not None}
    refuted_beside_claim = set(refuting_evidence) & claimed_ids
    result.refutation_basis = _snapshot_status_basis(Path(reference))

    # One client for the run, so its cache and its pacing gate are shared: a CAID appearing on several
    # evidence rows is looked up once, and the registry sees one paced stream rather than N.
    registry = registry if registry is not None else ClingenAlleleClient(offline=offline)
    consulted_registry = False

    # The anchor reader, built once per run so its cache is shared. `SequenceProxy` degrades to
    # `None` offline and on an unreachable service, which is what makes `anchor_base_unreadable` a
    # real outcome rather than a crash.
    sequences = SequenceProxy(offline=offline)

    def read_base(chrom: str, pos: int) -> str | None:
        """One GRCh38 reference base at a 1-based position, or `None`."""
        try:
            accession = refget_accession(chrom)
        except UnsupportedBuildError:  # pragma: no cover - GRCh38 is this snapshot's only build
            return None
        if accession is None:
            return None
        return sequences.subsequence(accession, pos - 1, pos)

    def read_window(chrom: str, start: int, end: int) -> str | None:
        """GRCh38 reference bases over 1-based `[start, end]`, or `None`."""
        try:
            accession = refget_accession(chrom)
        except UnsupportedBuildError:  # pragma: no cover - GRCh38 is this snapshot's only build
            return None
        if accession is None:
            return None
        return sequences.subsequence(accession, start - 1, end)

    variant_partials: list[PartialRow] = []
    study_partials: list[PartialRow] = []
    for row in admitted:
        if int(row["variant_id"]) in contested_ids:
            result.withheld["contested_variant"] += 1
            continue
        if row["direction"] is None:
            result.withheld["refutation_states_no_direction"] += 1
            continue
        if int(row["variant_id"]) in refuted_beside_claim:
            # The row IS drafted — the source supports this direction and the drafter writes what the
            # source said. Recorded, not withheld: withholding would be the tier choosing a winner
            # between two of the source's own statements, and the accounting equality over `withheld`
            # counts rows that were not written.
            _note_refuted(result, row, refuting_evidence)
        if _needs_the_registry(row):
            # The snapshot kept this row because it has a *route* to an identity rather than one.
            # Walking that route is what turns it into a drafted row, and the three outcomes stay
            # three: placed, established-absence, and nobody-asked.
            consulted_registry = True
            found = registry.resolve(row["allele_registry_id"])
            if found.outcome == "needs_anchor" and found.unanchored is not None:
                # An insertion or a deletion, which the registry states with one side empty. One
                # reference base turns it into a row; without one it stays withheld rather than
                # becoming a half-written coordinate.
                placed = anchor_indel(found.unanchored, read_base)
                # RM273: the registry's interbase point is HGVS's 3'-most one, so inside a repeat the
                # anchored row is right of VCF's spelling. Left-align it before it becomes a row, or the
                # drafted coordinate keys a different spelling from every caller's.
                if placed is not None:
                    placed = _left_align(*placed, read_window)
                if placed is None:
                    result.withheld["anchor_base_unreadable"] += 1
                    continue
                row = dict(row)
                row["chrom"], row["start"], row["ref"], row["alt"] = placed
                result.caid_anchored_indels += 1
                partial = _variant_row(row, dataset=result.dataset)
                if identity_refused_by_model(partial.cells) is not None:
                    result.withheld["identity_refused_by_model"] += 1
                    continue
                variant_partials.append(partial)
                study = _study_row(row)
                if study is not None:
                    study_partials.append(study)
                continue
            if not found.placeable:
                key = "caid_no_identity" if found.outcome == "no_identity" else "caid_unresolved"
                result.withheld[key] += 1
                continue
            row = dict(row)
            if found.rsid:
                row["rsid"] = found.rsid
                result.caid_resolved_by_rsid += 1
            elif found.coordinate:
                row["chrom"], row["start"], row["ref"], row["alt"] = found.coordinate
                result.caid_resolved_by_coordinate += 1
        partial = _variant_row(row, dataset=result.dataset)
        if identity_refused_by_model(partial.cells) is not None:
            result.withheld["identity_refused_by_model"] += 1
            continue
        variant_partials.append(partial)
        study = _study_row(row)
        if study is not None:
            study_partials.append(study)

    # The licence row lands inside each table's commit (RM232). The registry is listed only where it
    # was actually asked, the same predicate the tail call uses.
    commit_licence = licence_commit(
        sources=[CIVIC_SOURCE] + ([CLINGEN_ALLELE_REGISTRY_TERMS.source] if consulted_registry else []),
        spec_dir=spec_dir,
        dataset=result.dataset,
        declared_use=declared_use,
        error=CivicDraftError,
    )
    if variant_partials:
        result.reports.append(
            append_partial_rows(
                spec_dir, "variants.csv", variant_partials, dry_run=dry_run, before_commit=commit_licence
            )
        )
    if study_partials:
        result.reports.append(
            append_partial_rows(
                spec_dir, "studies.csv", study_partials, dry_run=dry_run, before_commit=commit_licence
            )
        )
    result.warnings.extend(_withheld_warnings(result.withheld, contested_names))
    result.warnings.extend(_refuted_warning(result.refuted_beside_claim, result.refutation_basis))

    # **A pass that consults a source writes its `SourceRow`; one that contributed nothing writes
    # none** (`@write-the-sourcerow`). The comment here said exactly that and the gate was `not
    # dry_run` and nothing else, so a `--gene` filter matching nothing still wrote a `civic` row —
    # a licence row claiming a module uses CIViC when it does not (RM222). `strchive_draft` and
    # `mitomap_draft` both gate on an outcome and `test_strchive_draft.py` refuses this shape on that
    # path; this is the same predicate, keyed on what *this run covered*: a locus is in the module's
    # table because of this provider, added now or recognised as `already_present` from an earlier run.
    covered = any(
        outcome.status in {"added", "already_present"}
        for report in result.reports
        for outcome in report.outcomes
    )
    if not dry_run and covered:
        # The registry is listed only where it was actually asked, which is why the flag is set at the
        # lookup rather than derived from the snapshot's contents.
        #
        # **`dataset` travels with the row** (RM222). It was computed for every row's `conclusion` and
        # then dropped on the floor, because `record_source_terms` had no way to carry one — so the
        # licence row read `dataset=''`, and `--verify-datasets` had nothing to compare and
        # `withdraw_stale_dataset` nothing to withdraw. A CIViC-drafted module sat outside the currency
        # check every other drafted module is inside.
        consulted = [CIVIC_SOURCE] + ([CLINGEN_ALLELE_REGISTRY_TERMS.source] if consulted_registry else [])
        result.warnings.extend(
            record_draft_provenance(
                provider=_PROVIDER,
                sources=consulted,
                spec_dir=spec_dir,
                dataset=result.dataset,
                covered=True,
                drafted=any(
                    outcome.status == "added" for report in result.reports for outcome in report.outcomes
                ),
                declared_use=declared_use,
                error=CivicDraftError,
            )
        )
    return result