Skip to content

just_dna_enricher.strchive_draft

just_dna_enricher.strchive_draft

Draft repeat_alleles.csv identity rows from the STRchive catalogue (RM165, the other half).

The band half of this source is checked and never written (strchive.check_repeat_bands). This is the half that is written: the facts the catalogue owns and an author would otherwise transcribe by hand.

And it is a much thinner row than the item expected, which is the finding this module carries. The proposal listed chrom/start_hg38/stop_hg38, locus_structure, ref_copies and the disease identifiers as the identity half — and RepeatAlleleRow has no column for any of them except the trait. That is not an oversight in the model: repeat coordinates are RM65, deferred, and the motif structure is RM66, parked. So what a drafted row can actually carry is the gene, the motif, the trait CURIE and the fixed measure_kind, with conclusion left as a placeholder for the human. The rest of the catalogue's identity half is reported and counted rather than silently dropped — a number this pass computes and discards is one every reader has to recompute.

What that thin row is still worth: the motif orientation is a real trap (a minus-strand locus is published in two spellings), the MONDO id is a real transcription task, and the list of loci is itself the thing an author starting a repeat module does not have. A drafted row is a locus the author has been told exists, with its identity spelled the way the catalogue spells it and every interpretive cell left blank.

Never drafted: the bands. measure_min/measure_max/measure_tiling, and the per-band direction/clin_sig/phenotype that go with them. Measured against both corpus modules and wrong in both: STRchive is a band coarser than fmr1_cgg_repeat at the premutation threshold, and its pathogenic_max would invent a ceiling htt_repeat_expansion deliberately leaves open. The withheld set is derived from the drafted one rather than listed beside it, so a column added to the model lands on the withheld side by default.

Appends, never mutates (@draft-appends). A locus already in the table is reported already_present and nothing is rewritten — matched on (gene, repeat_unit), which is the identity that survives the human filling in the bands and splitting one drafted row into four.

StrchiveDraftError

Bases: RuntimeError

A draft could not be attempted — an unreadable catalogue, or an unwritable spec directory.

StrchiveDraftResult dataclass

StrchiveDraftResult(
    report: DraftReport | None = None,
    candidates: int = 0,
    withheld: dict[str, int] = dict(),
    contested: list[tuple[str, str]] = list(),
    fractional_ref_copies: int = 0,
    with_locus_structure: int = 0,
    warnings: list[str] = list(),
    dataset: str | None = None,
    skipped: bool = False,
)

What was drafted, and an account of every catalogue locus that was not.

accounts_for_every_candidate

accounts_for_every_candidate() -> bool

Every admitted locus is either a partial row this run offered, or a counted withholding.

Source code in enricher/src/just_dna_enricher/strchive_draft.py
def accounts_for_every_candidate(self) -> bool:
    """Every admitted locus is either a partial row this run offered, or a counted withholding."""
    offered = len(self.report.outcomes) if self.report else 0
    return offered + sum(self.withheld.values()) == self.candidates

draft_repeat_loci

draft_repeat_loci(
    spec_dir: Path,
    genes: Sequence[str] = (),
    *,
    catalogue: Path | StrchiveCatalogue | None = None,
    declared_use: str = "unstated",
    dry_run: bool = False,
) -> StrchiveDraftResult

Append one identity row per STRchive locus into the module's repeat_alleles.csv.

genes filters; empty drafts every locus in the catalogue. The filter runs first and is counted separately from the withholding, because "the catalogue has nothing for this gene" and "it has something and this provider would not write it" are different answers.

Contestation is decided over the whole admitted set before any row is built, so a filter cannot remove the second claimant and leave the first reading as uncontested (@filter-before-the-group-picks-a-winner).

Source code in enricher/src/just_dna_enricher/strchive_draft.py
def draft_repeat_loci(
    spec_dir: Path,
    genes: Sequence[str] = (),
    *,
    catalogue: Path | StrchiveCatalogue | None = None,
    declared_use: str = "unstated",
    dry_run: bool = False,
) -> StrchiveDraftResult:
    """Append one identity row per STRchive locus into the module's `repeat_alleles.csv`.

    `genes` filters; empty drafts every locus in the catalogue. The filter runs **first** and is
    counted separately from the withholding, because "the catalogue has nothing for this gene" and
    "it has something and this provider would not write it" are different answers.

    Contestation is decided over the whole admitted set before any row is built, so a filter cannot
    remove the second claimant and leave the first reading as uncontested
    (`@filter-before-the-group-picks-a-winner`).
    """
    spec_dir = Path(spec_dir)
    result = StrchiveDraftResult(withheld=dict.fromkeys(WITHHELD_REASONS, 0))

    if catalogue is None:
        result.skipped = True
        result.warnings.append(
            "No STRchive catalogue was provisioned, so nothing was drafted from it. Build one with "
            "`just-dna-enricher strchive build --release <tag>`, or pass a downloaded "
            "STRchive-loci.json. Nobody-asked is not the same as the source having nothing."
        )
        return result

    # MIT grants the fetch under every declaration, so this always returns `None` — kept because the
    # gate is per source and a caller reading this file should see which answer it gives, not have to
    # infer that the question was never asked.
    declared_use, declared_from = effective_declared_use(spec_dir, STRCHIVE_TERMS, declared_use)  # S105
    refusal = check_declared_use(STRCHIVE_TERMS, declared_use)
    if refusal is not None:  # pragma: no cover - unreachable while the terms stay permissive
        raise StrchiveDraftError(refusal)

    loaded = catalogue if isinstance(catalogue, StrchiveCatalogue) else load_strchive_catalogue(catalogue)
    result.dataset = loaded.dataset

    wanted = {g.strip().upper() for g in genes if g.strip()}
    admitted = [locus for locus in loaded.loci if not wanted or locus.gene.upper() in wanted]
    result.candidates = len(admitted)

    # Over the admitted set, before any row is built.
    claims: dict[tuple[str, str], list[str]] = {}
    for locus in admitted:
        for motif in locus.motifs:
            claims.setdefault((locus.gene, motif), []).append(locus.locus_id)
    contested = {key for key, ids in claims.items() if len(ids) > 1}
    result.contested = sorted(contested)

    partials: list[PartialRow] = []
    offered: list[StrchiveLocus] = []
    incomplete: list[str] = []
    for locus in admitted:
        result.fractional_ref_copies += int(
            locus.ref_copies is not None and not float(locus.ref_copies).is_integer()
        )
        result.with_locus_structure += int(bool(locus.locus_structure))
        # The reference spelling is the one drafted: a locus published in two orientations is one
        # locus, and writing both would put two rows in the table for one place on the genome. The
        # check reads the union so an author who wrote the other spelling still joins.
        motif = locus.motifs[0] if locus.motifs else ""
        if not motif:
            result.withheld["no_motif"] += 1
            continue
        if (locus.gene, motif) in contested:
            result.withheld["contested_key"] += 1
            continue
        row, missing = _partial(locus, motif)
        if row is None:
            result.withheld["incomplete_row"] += 1
            incomplete.append(f"{locus.locus_id} ({', '.join(missing)})")
            continue
        partials.append(row)
        offered.append(locus)

    # The licence row lands inside the table's commit (RM232).
    commit_licence = licence_commit(
        sources=[STRCHIVE_TERMS.source],
        spec_dir=spec_dir,
        dataset=result.dataset,
        declared_use=declared_use,
        error=StrchiveDraftError,
    )
    if partials:
        result.report = append_partial_rows(
            spec_dir, REPEAT_ALLELES_CSV, partials, dry_run=dry_run, before_commit=commit_licence
        )

    # Carried on the result, not logged: every caller here renders `warnings` itself, and
    # `civic_draft`/`clinvar_draft` do the same. Logging them as well printed each note twice.
    result.warnings.extend(_notes(result, incomplete))
    result.warnings.extend(_evidence_notes(offered, _evidence_by_locus(loaded.path)))

    # **A pass that consults a source writes its `SourceRow`; one that contributed nothing writes
    # none** (`@write-the-sourcerow`). The key is what *this run covered*: a row is written when at
    # least one locus is in the module's table because of this provider — added now, or added by an
    # earlier run and recognised as `already_present`. A `--gene` filter that matched nothing, or a
    # catalogue whose every locus was withheld, leaves the module untouched, and a licence row saying
    # "this module uses STRchive" would then be a claim about a module that does not.
    covered = result.report is not None and any(
        outcome.status in {"added", "already_present"} for outcome in result.report.outcomes
    )
    if not dry_run:
        # The covered-predicate stays here because it is this provider's own reading of "contributed
        # something"; what the scaffold owns is what follows from it, which is where the halves used
        # to come apart (RM228).
        result.warnings.extend(
            record_draft_provenance(
                provider=_PROVIDER,
                sources=[STRCHIVE_TERMS.source],
                spec_dir=spec_dir,
                dataset=result.dataset,
                covered=covered,
                drafted=bool(result.drafted),
                declared_use=declared_use,
                error=StrchiveDraftError,
            )
        )
    return result