Skip to content

just_dna_enricher.clinvar_draft

just_dna_enricher.clinvar_draft

Draft a gene panel's variants.csv rows from the ClinVar snapshot (0.5.1, RM26).

The provider that RM4 has been waiting for, in the shape the charter allows: it drafts rows a human then owns, with no compile-time reference materialization and no injected reference in the compile path. The compiler stays inject-only; this is the network tier, so the snapshot it reads is found on the usual cache ladder and provisioned from HuggingFace when absent (--offline forbids that, and an explicit --snapshot is taken as given). It used to be a required argument, which meant the published snapshot could not reach an author at all: they had to build 4.4M records from a 200 MB VCF first.

Why these rows are partial, and why that is the honest answer rather than a limitation. VariantRow.genotype is required and ClinVar publishes alleles, not genotypes. Whether carrying a pathogenic allele once is informative — a carrier, an affected proband, neither — is zygosity interpretation, and it follows from the condition's inheritance mode, not from the allele. ClinVar cannot supply it and this provider must not invent it: writing A/G because the alt is G would be a clinical claim the source never made. reference_examples/pathogenic_clinvar/ is a human having made exactly that call, by hand, per row.

So genotype is left carrying vocab.TEMPLATE_PLACEHOLDER, which no mode compiles: the file is authored, in place, in the right order, and loudly incomplete until a human has decided. Rows are placed into their gene's block (group_by=("gene",)), because a panel is read gene by gene and stubs stranded at the end of a 500-row file are what made this unpleasant enough to look like a blocker.

Identity is filled whole or not at all. With an rsID, only the rsID goes in; without one, the full chrom/start/ref/alts. Never a subset — a lone alts on a position-only row makes derive_variant_key mint a VRS ga4gh:VA.… id instead of chrom:start:ref, so a partial coordinate silently changes which variant the row is.

What it will not fill, each a rule rather than an omission:

  • weight, direction, effect_size, effect_measure, effect_allele — ClinVar publishes no effect statistic, and a weight is the author's model of the finding.
  • trait_efo_id — ClinVar's condition is free text and MedGen, not EFO. Mapping it is inference.
  • acmg_sf — a different list this package deliberately does not hold (see the roadmap).
  • curator, method — the spec's defaults: block owns those.

ClinVarDraftError

Bases: RuntimeError

A gene-panel draft could not be completed.

ClinVarDraftResult dataclass

ClinVarDraftResult(
    reports: list[DraftReport] = list(),
    warnings: list[str] = list(),
    skipped: bool = False,
)

What a panel draft did.

added property

added: int

Rows added across every table this run wrote — variants and their studies.

Ask added_for when you mean one of them: since the provider began drafting grounding evidence too, a bare total no longer answers "how many variants did I get".

added_for

added_for(csv_name: str) -> int

Rows added to one table.

Source code in enricher/src/just_dna_enricher/clinvar_draft.py
def added_for(self, csv_name: str) -> int:
    """Rows added to one table."""
    return sum(len(r.added) for r in self.reports if r.csv_name == csv_name)

multi_allelic_rsids

multi_allelic_rsids(records: Sequence[dict]) -> set[str]

rsIDs that name more than one distinct allele event in this selection.

An rsID is a position/multi-allelic-level tag, not a per-allele one: ClinVar lists rs773443949 in HFE as both G>A and G>T. An rsid-only row cannot say which, so drafting one row per record would write two identical rows — and de-duplicating them would silently drop a real allele. Found by drafting an actual panel; these records take the coordinate identity instead.

The event is (chrom, start, ref, alt), and keying the site on ref was the bug (S41). The predicate used to group by (rsid, chrom, start, ref) and fire on >1 alt inside that group, which reads as "more than one alt at one position" and is not: an ordinary ClinVar dup/del mirror pair — A>AT beside ATT>A at one position, the same event written from either side — lands in two groups of one alt each, so the rsID was never flagged, both records reduced to the same rsid signature, and append_partial_rows dropped the second as already_present. A differing ref breaks an rsid-only identity exactly as thoroughly as a differing alt; the docstring's own claim was the correct rule and the code was narrower than it. Measured on the 2026-06-27 snapshot over BRCA1/BRCA2/ATM/MLH1/MSH2: 942 rsIDs flagged before, 1,589 after, and the 647 newly flagged are exactly the 647 identities that were collapsing — 725 records, of which 187 dropped a better-reviewed record than the one kept, since select_by_gene orders by ref before review_stars DESC and the survivor is therefore an artifact of allele spelling.

Distinctness is over the whole event rather than over records, deliberately: two rows of the same allele (a re-submission under a second variation_id) are one claim written twice and collapsing them loses nothing, while coordinate identity would not separate them anyway. On that measurement the two readings coincide — every multi-record rsID is also multi-allele — but only this one is true by construction, and the other would flag rsIDs whose collapse the fix cannot repair.

Source code in enricher/src/just_dna_enricher/clinvar_draft.py
def multi_allelic_rsids(records: Sequence[dict]) -> set[str]:
    """rsIDs that name more than one distinct **allele event** in this selection.

    An rsID is a **position/multi-allelic-level** tag, not a per-allele one: ClinVar lists
    `rs773443949` in HFE as both `G>A` and `G>T`. An rsid-only row cannot say which, so drafting one
    row per record would write two identical rows — and de-duplicating them would silently drop a real
    allele. Found by drafting an actual panel; these records take the coordinate identity instead.

    **The event is `(chrom, start, ref, alt)`, and keying the *site* on `ref` was the bug (S41).**
    The predicate used to group by `(rsid, chrom, start, ref)` and fire on >1 alt inside that group,
    which reads as "more than one alt at one position" and is not: an ordinary ClinVar dup/del mirror
    pair — `A>AT` beside `ATT>A` at one position, the same event written from either side — lands in
    two groups of one alt each, so the rsID was never flagged, both records reduced to the same rsid
    signature, and `append_partial_rows` dropped the second as `already_present`. A differing `ref`
    breaks an rsid-only identity exactly as thoroughly as a differing `alt`; the docstring's own claim
    was the correct rule and the code was narrower than it. Measured on the `2026-06-27` snapshot over
    BRCA1/BRCA2/ATM/MLH1/MSH2: 942 rsIDs flagged before, 1,589 after, and the 647 newly flagged are
    exactly the 647 identities that were collapsing — 725 records, of which 187 dropped a
    *better-reviewed* record than the one kept, since `select_by_gene` orders by `ref` before
    `review_stars DESC` and the survivor is therefore an artifact of allele spelling.

    Distinctness is over the whole event rather than over records, deliberately: two rows of the same
    allele (a re-submission under a second `variation_id`) are one claim written twice and collapsing
    them loses nothing, while coordinate identity would not separate them anyway. On that measurement
    the two readings coincide — every multi-record rsID is also multi-allele — but only this one is
    true by construction, and the other would flag rsIDs whose collapse the fix cannot repair."""
    events_by_rsid: dict[str, set[tuple]] = {}
    for record in records:
        rsid = (record.get("rsid") or "").strip()
        if not rsid:
            continue
        event = (
            record.get("chrom"),
            record.get("start"),
            (record.get("ref") or "").strip(),
            (record.get("alt") or "").strip(),
        )
        events_by_rsid.setdefault(rsid, set()).add(event)
    return {rsid for rsid, events in events_by_rsid.items() if len(events) > 1}

sole_expressible_genotype

sole_expressible_genotype(record: dict) -> str | None

The one genotype a caller can emit at this locus, or None where zygosity is a real decision.

The placeholder exists to protect a judgement, and on a non-diploid contig there is none to protect (S6). Carrying a pathogenic allele is a carrier state or an affected one depending on the condition's inheritance mode, which ClinVar does not state — that is why genotype is stubbed at all. But the mitochondrial genome is haploid and chrY outside the pseudoautosomal regions is hemizygous: exactly one genotype is expressible per allele there, so the decision the stub is holding open does not exist, and every consumer of draft_gene_panel was left to rediscover independently that its natural "write both zygosities" fill is wrong for those rows. One did, at 264 mitochondrial loci in a genome-wide panel and 260 in a cardiac one, each asserting a second copy that is not there.

Y is decided per locus, three-valued, through the same predicate the compiler's ploidy check uses: PAR1 and PAR2 recombine with X and are diploid in every karyotype, and XG and SPRY3 straddle a boundary, so a gene- or contig-wide verdict is wrong for half of either. True (diploid) and None (no PAR table for this build) both keep the placeholder — an undecided question is not an answer, and the author still has one to make.

The allele written is the ALT: a variant row is an annotation about carrying the finding, and on a haploid contig carrying it is spelled with one allele. The heteroplasmy axis is a different question with its own table kind, which is why the aggregated notice at the draft's end says so rather than leaving the reading implicit.

Source code in enricher/src/just_dna_enricher/clinvar_draft.py
def sole_expressible_genotype(record: dict) -> str | None:
    """The one genotype a caller can emit at this locus, or `None` where zygosity is a real decision.

    **The placeholder exists to protect a judgement, and on a non-diploid contig there is none to
    protect** (S6). Carrying a pathogenic allele is a carrier state or an affected one depending on
    the condition's inheritance mode, which ClinVar does not state — that is why `genotype` is stubbed
    at all. But the mitochondrial genome is haploid and chrY outside the pseudoautosomal regions is
    hemizygous: exactly one genotype is expressible per allele there, so the decision the stub is
    holding open does not exist, and every consumer of `draft_gene_panel` was left to rediscover
    independently that its natural "write both zygosities" fill is wrong for those rows. One did, at
    264 mitochondrial loci in a genome-wide panel and 260 in a cardiac one, each asserting a second
    copy that is not there.

    Y is decided **per locus**, three-valued, through the same predicate the compiler's ploidy check
    uses: PAR1 and PAR2 recombine with X and are diploid in every karyotype, and `XG` and `SPRY3`
    straddle a boundary, so a gene- or contig-wide verdict is wrong for half of either. `True`
    (diploid) and `None` (no PAR table for this build) both keep the placeholder — an undecided
    question is not an answer, and the author still has one to make.

    The allele written is the **ALT**: a variant row is an annotation about carrying the finding, and
    on a haploid contig carrying it is spelled with one allele. The heteroplasmy axis is a different
    question with its own table kind, which is why the aggregated notice at the draft's end says so
    rather than leaving the reading implicit.
    """
    chrom = normalize_chrom(str(record.get("chrom") or ""))
    alt = (record.get("alt") or "").strip()
    start = record.get("start")
    if not alt or "," in alt or chrom not in ("MT", "Y"):
        return None
    if chrom == "Y" and (start is None or in_pseudoautosomal_region("Y", int(start)) is not False):
        return None
    return alt

draft_gene_panel

draft_gene_panel(
    spec_dir: Path,
    genes: Sequence[str],
    *,
    snapshot: Path | None = None,
    clin_sig: frozenset[str] = DEFAULT_CLIN_SIG,
    min_review_stars: int = 2,
    max_citations: int = DEFAULT_MAX_CITATIONS,
    declared_use: str = "unstated",
    offline: bool = False,
    download: bool = True,
    dry_run: bool = False,
) -> ClinVarDraftResult

Draft variants.csv rows for one or more genes, leaving genotype for the human.

Re-runnable and additive: run it per gene as a panel grows. A variant already in the file — stub or filled — is reported, never re-added, because a partial row keys on the identity columns rather than on a key that runs through the placeholder.

min_review_stars defaults to 2 (multiple submitters, no conflicts). A panel that silently mixes a 0-star "no assertion criteria" submission with a 3-star expert-panel review is worse than one that says which floor it drew from, and the floor belongs in the author's hands.

The snapshot is found, then provisioned — it used to be a required argument. snapshot=None resolves the usual cache ladder and, unless offline, downloads the published one (download.ensure_clinvar_snapshot), the same shape enrich() and the gene-metrics pass use. Before this, the published snapshot could not reach an author: they had to build 4.4M records from a 200 MB VCF themselves, or already know the cache path. That matters most for the citations, which are what make a drafted panel compilable at all (studies.csv is mandatory and the VCF carries no PMIDs) and which now travel with the published snapshot.

Source code in enricher/src/just_dna_enricher/clinvar_draft.py
def draft_gene_panel(
    spec_dir: Path,
    genes: Sequence[str],
    *,
    snapshot: Path | None = None,
    clin_sig: frozenset[str] = DEFAULT_CLIN_SIG,
    min_review_stars: int = 2,
    max_citations: int = DEFAULT_MAX_CITATIONS,
    declared_use: str = "unstated",
    offline: bool = False,
    download: bool = True,
    dry_run: bool = False,
) -> ClinVarDraftResult:
    """Draft `variants.csv` rows for one or more genes, leaving `genotype` for the human.

    Re-runnable and additive: run it per gene as a panel grows. A variant already in the file — stub
    or filled — is reported, never re-added, because a partial row keys on the identity columns rather
    than on a key that runs through the placeholder.

    `min_review_stars` defaults to 2 (multiple submitters, no conflicts). A panel that silently mixes
    a 0-star "no assertion criteria" submission with a 3-star expert-panel review is worse than one
    that says which floor it drew from, and the floor belongs in the author's hands.

    **The snapshot is found, then provisioned — it used to be a required argument.** `snapshot=None`
    resolves the usual cache ladder and, unless `offline`, downloads the published one
    (`download.ensure_clinvar_snapshot`), the same shape `enrich()` and the gene-metrics pass use. Before
    this, the published snapshot could not reach an author: they had to build 4.4M records from a 200 MB
    VCF themselves, or already know the cache path. That matters most for the *citations*, which are what
    make a drafted panel compilable at all (`studies.csv` is mandatory and the VCF carries no PMIDs) and
    which now travel with the published snapshot.
    """
    declared_use, declared_from = effective_declared_use(spec_dir, CLINVAR_TERMS, declared_use)  # S105
    skip_reason = check_declared_use(CLINVAR_TERMS, declared_use)
    if skip_reason:
        return ClinVarDraftResult(warnings=[skip_reason], skipped=True)

    # **Before the snapshot, not after it** (R2-2). `source_build_mismatch` raises `EnrichmentError`
    # on a present-but-unreadable `module_spec.yaml` — correctly, since a module whose declaration
    # cannot be read has no build to draft against — and asking it below meant a spec carrying only
    # `name:` failed *after* `_resolve_snapshot` had provisioned a published snapshot. It reads one
    # file beside the spec, so it costs nothing here. The warning is still appended in its old place.
    build_warning = source_build_mismatch(spec_dir, "the ClinVar snapshot", CLINVAR_GENOME_BUILD)

    reference, provisioning_warnings = _resolve_snapshot(snapshot, offline=offline, download=download)
    records = select_by_gene(reference, list(genes), clin_sig=clin_sig, min_review_stars=min_review_stars)
    warnings: list[str] = list(provisioning_warnings)
    # The snapshot is built from NCBI's `vcf_GRCh38/clinvar.vcf.gz`, and `_row_cells` writes the full
    # coordinate for any record this pass cannot key by rsID — so a module on another build is about
    # to record GRCh38 positions under its own declaration. See `enrich.source_build_mismatch`;
    # computed above, before the snapshot is resolved.
    if build_warning:
        warnings.append(build_warning)
    partials: list[PartialRow] = []
    unkeyable = 0
    ambiguous = multi_allelic_rsids(records)
    if ambiguous:
        # The list used to be printed whole, which was right at the one rsID that motivated it and is
        # not at the 1,589 a real cancer panel produces (S41 widened the predicate). `examples` is the
        # house aggregation, shared so this cannot drift from the other callers'.
        warnings.append(
            f"{len(ambiguous)} rsID(s) name more than one allele here "
            f"({examples(sorted(ambiguous))}) — written with their full coordinate, since an rsID "
            f"alone cannot say which allele the row is about."
        )
    record_by_signature: dict[tuple[str, ...], dict] = {}
    # signature -> the columns this run left for a human, so the reports below read the *decision*
    # `PartialRow` was given rather than restating which columns a stub can land in.
    stubbed_by_signature: dict[tuple[str, ...], tuple[str, ...]] = {}
    for record in records:
        cells = _row_cells(record, force_coordinate=(record.get("rsid") or "").strip() in ambiguous)
        if cells is None:
            unkeyable += 1
            continue
        # `state` is required and has no honest value for an undecided clinical call, so it is stubbed
        # for the same reason `genotype` usually is: the source did not say, and only a human can. See
        # `STATE_BY_CLIN_SIG` — `VALID_STATES` offers no "uncertain" member, and every candidate
        # asserts something ClinVar declined to (`neutral` says benign, `risk` says a direction).
        # Skipping the row instead would throw away the conclusion, phenotype, clin_sig and citations
        # already assembled for it.
        #
        # Derived from the cells rather than listed: `_row_cells` decides what it can state, and a
        # column it filled is not a column to stub. A row on a haploid contig therefore arrives
        # complete (`stubbed=()`), which `PartialRow` handles as the degenerate case — it validates
        # every cell and matches on the same identity columns.
        stubbed = tuple(column for column in ("genotype", "state") if column not in cells)
        signature = _signature(cells)
        record_by_signature[signature] = record
        stubbed_by_signature[signature] = stubbed
        partials.append(
            PartialRow(
                model=VariantRow,
                cells=cells,
                stubbed=stubbed,
                # Identity, not the natural key: the key runs through `genotype`, which is the stub.
                # Matching here means "this variant is already in the panel, however it was written".
                # `alts` is in the set because without it two rows of a multi-allelic site collapse
                # into one and a real allele is lost — which is exactly what drafting HFE did.
                match_on=_MATCH_ON,
            )
        )
    if unkeyable:
        warnings.append(
            f"{unkeyable} ClinVar record(s) skipped: neither an rsID nor a complete coordinate, so "
            f"nothing this format can key on."
        )
    if not partials:
        return ClinVarDraftResult(warnings=warnings + ["nothing matched; no rows drafted"])

    # The release label is read here rather than at the tail (RM232): the licence row now lands
    # inside each table's commit, so what the row states has to be known before the first write. The
    # warning for a snapshot that cannot state its release stays where it was, below.
    dataset = clinvar_dataset_label(reference)
    commit_licence = licence_commit(
        sources=[CLINVAR_TERMS.source],
        spec_dir=spec_dir,
        dataset=dataset,
        declared_use=declared_use,
        error=ClinVarDraftError,
    )
    report = append_partial_rows(
        spec_dir,
        "variants.csv",
        partials,
        group_by=("gene",),
        dry_run=dry_run,
        before_commit=commit_licence,
    )
    reports = [report]
    warnings.extend(_superseded_rsid_rows(report.path, ambiguous))

    # Grounding evidence, from ClinVar's own literature links. Without this a drafted panel could not
    # compile at all — `studies.csv` is mandatory and the VCF carries no PMIDs — so the provider
    # produced a module that needed evidence nobody could supply. Build it with `clinvar citations`.
    links = citations_for(reference, [str(r.get("variation_id") or "") for r in records])
    if not links:
        warnings.append(
            "no citations table in the snapshot, so no studies.csv rows were drafted — grounding "
            "evidence is mandatory, so add it by hand, or run `just-dna-enricher clinvar citations` "
            "(a published snapshot now carries the table, so re-provisioning an empty cache also gets it)."
        )
    else:
        studies, dropped, unusable = _study_rows(records, links, max_citations, ambiguous)
        if studies:
            reports.append(
                append_rows(
                    spec_dir,
                    "studies.csv",
                    studies,
                    group_by=("rsid",),
                    dry_run=dry_run,
                    before_commit=commit_licence,
                )
            )
        if dropped:
            warnings.append(
                f"{dropped} further ClinVar citation(s) not drafted (--max-citations {max_citations})."
            )
        # Kept apart from the cap above: one is a choice this run made, the other is ClinVar filing a
        # non-PMID under PubMed. Reporting them together would read as "you asked for fewer".
        if unusable:
            warnings.append(
                f"{unusable} ClinVar citation(s) skipped: the id ClinVar filed under PubMed is not a "
                f"PMID (nine digits, where PubMed is at eight). Rebuilding the snapshot with a current "
                f"`clinvar citations` drops them at the source."
            )
    warnings.extend(_refusal_summary(report.invalid))

    # ── What is still open, scoped to the FILE rather than to this run (RM71) ────────────────────
    #
    # Both stub reports below hang off `_open_stubs`, so a re-run of the command answers the question
    # it answered the first time and `--dry-run` answers it without appending. They have to move
    # together: the `state` line reads as "these rows *also*", so a file-wide genotype denominator
    # beside a run-scoped `state` one would name two different sets in consecutive lines.
    genotype_stubs = _open_stubs(report, "genotype", stubbed_by_signature)
    if genotype_stubs:
        warnings.append(
            f"{len(genotype_stubs)} row(s) carry an unreplaced genotype placeholder and will not "
            f"compile until you decide the zygosity each finding is about."
        )
        warnings.extend(
            _genotype_worklist(
                [
                    record_by_signature[signature]
                    for signature, _ in genotype_stubs
                    if signature in record_by_signature
                ]
            )
        )
        # The rows this run cannot describe: drafted for another gene, or under a filter this run did
        # not use, so nothing here holds their alleles. Counted and their genes named — an unknown is
        # withheld, never guessed, and never silently dropped from a count the file can see.
        withheld = [
            (row or {}).get("gene") or ""
            for signature, row in genotype_stubs
            if signature not in record_by_signature or not _publishes_alleles(record_by_signature[signature])
        ]
        if withheld:
            genes = examples(sorted({gene.strip() for gene in withheld if gene.strip()}))
            # States what it knows and stops. It does NOT say *why* this run holds no record — the
            # usual reason is another gene or a tighter --clin-sig/--min-review-stars, but a record
            # that publishes no alleles lands here too, and naming a cause the pass did not establish
            # would be a guess dressed as a finding.
            warnings.append(
                f"{len(withheld)} row(s) of that list carry alleles this run cannot state, because "
                f"nothing it selected covers them "
                f"(gene: {genes or 'none recorded'}). The alleles are withheld rather than guessed: "
                f"draft-panel for the gene each row records, or `hint variant`, will state them."
            )

    warnings.extend(
        _state_stub_warnings(_open_stubs(report, "state", stubbed_by_signature), record_by_signature)
    )

    # The non-diploid notice stays scoped to what this run WROTE, deliberately: it reports a reading
    # the provider committed to on rows it just filled, not work anybody has left to do. One line, not
    # one per row — at panel scale this is hundreds of loci and the reading is identical for all.
    written_records = [
        record_by_signature[o.key]
        for o in report.added
        if o.key in record_by_signature and "genotype" not in stubbed_by_signature.get(o.key, ())
    ]
    if written_records:
        contigs = ", ".join(sorted({normalize_chrom(str(r.get("chrom") or "")) for r in written_records}))
        warnings.append(
            f"{len(written_records)} row(s) on non-diploid contigs ({contigs}) were written with "
            f"a single-allele genotype: exactly one is expressible there, so no zygosity decision "
            f"was pre-empted. They read as homoplasmic/hemizygous — a heteroplasmic level is a "
            f"different question and belongs in heteroplasmy.csv."
        )
    # WHICH release the rows were copied out of, in the column that exists to say so (RM4). The draft
    # was machined, so the marker is machined too: `clinical.tautology_reason` reads this back and can
    # then skip a check that a drafted module makes structurally unfailable, without the author having
    # to maintain a `panel:` block by hand for the sole benefit of one check. A snapshot that cannot
    # state its release leaves the cell empty rather than carrying a guess — and says so, because the
    # consequence is invisible otherwise.
    if dataset is None:
        warnings.append(
            "this snapshot does not say which ClinVar release it carries (no readable release.json), "
            "so the licence row records no dataset: nothing downstream can tell these rows were copied "
            "out of it, and the clin_sig cross-check will compare every one of them against it in full. "
            "That is the conservative outcome rather than a defect in the module. Build the snapshot "
            "with `just-dna-enricher clinvar build`, which writes the release.json this reads."
        )
    if not dry_run:
        # A source that rows were copied out of must be recorded, permissive terms or not: the compile
        # gate and `manifest.sources` read the licence table and nothing else.
        #
        # **`covered=True` reproduces this provider's behaviour exactly and is a question, not an
        # endorsement.** Every other drafter gates the row on this run having covered something; this
        # one writes it whenever the run was not a dry run, which is the shape RM222 found wrong in
        # `civic_draft` — a `--gene` filter matching nothing would write a licence row claiming a
        # module uses ClinVar when no ClinVar row reached it. RM228 is a behaviour-preserving
        # migration, so the gate is left as it was and the question is recorded rather than answered
        # under cover of a refactor.
        warnings.extend(
            record_draft_provenance(
                provider=_PROVIDER,
                sources=[CLINVAR_TERMS.source],
                spec_dir=spec_dir,
                dataset=dataset,
                covered=True,
                drafted=bool(report.added),
                declared_use=declared_use,
                error=ClinVarDraftError,
                stale_warning=lambda superseded, ds: (
                    f"this module already recorded rows drafted from {superseded}, and these came from "
                    f"{ds or 'a snapshot that does not state its release'} — so the licence row's "
                    f"dataset has been cleared rather than re-labelled: it cannot name two releases, "
                    f"and naming one would be a claim about rows that did not come from it. The "
                    f"consequence is that the clin_sig cross-check now runs over the whole table again."
                ),
            )
        )
    return ClinVarDraftResult(reports=reports, warnings=warnings)