just_dna_enricher.drafting¶
just_dna_enricher.drafting ¶
The shared drafting scaffold: what every *_draft.py provider does the same way (RM228).
Drafting grew bottom-up, one provider at a time, and seven modules ended up reimplementing the same
four decisions. The tell is that the two implementations which did derive their skip rule derived
it differently, and the better-looking one derived it from an incomplete oracle (see
identity_refused_by_model). This module is the mechanism those seven were each approximating.
The split this module exists to enforce. A provider's skip rule is two different rules that were being written as one hand-kept list:
- The model's requirement — what
VariantRow(or whichever row model) will accept. Derivable, identical for every provider, and it must never be restated:pgx_draftonce restated "no rsID and no position" whereHaplotypeRowwants rsID or chrom+start, anddraft --gene CYP2C9died on an unhandled pydantic error. - The provider's own source precondition — a true fact about that snapshot, such as ClinPGx carrying no coordinate. Legitimate, provider-specific, and it needs a stated reason.
Mashed into one list nobody can tell them apart, which is how mitomap_draft came to gate identity
on clin_sig — a column that is not an identity requirement at all — without anyone noticing. A
SourcePrecondition carries its reason as a field, so the two halves can never merge again.
What is deliberately not here. DRAFT_PROJECTIONS in provenance stays where it is and is now
derived from this registry rather than restated beside it — it answers a narrower question (which
sources cross-check a column they drafted) about a subset of these providers.
SourcePrecondition
dataclass
¶
Cells this provider's source requires, beyond whatever the model requires.
reason is not documentation — it is the field that makes the precondition auditable. A
condition with no reason is indistinguishable from a misread of the model, which is exactly the
state mitomap_draft's clin_sig clause was in.
DraftProvider
dataclass
¶
DraftProvider(
name: str,
module: str,
table: str,
match_on: tuple[str, ...],
kind: str,
precondition: SourcePrecondition | None = None,
checked: tuple[str, ...] = (),
projection_identity: tuple[str, ...] | None = None,
projection_identity_reason: str = "",
withheld_reasons: tuple[str, ...] = (),
)
One drafting provider, and everything the scaffold needs to treat it uniformly.
identity is per provider and not per table, deliberately. append_partial_rows builds its
covered-set from a single match_on tuple for the whole batch, so handing two providers that
write the same table one shared tuple would make the second one's rows stop matching on lap 2 and
be re-added every run (@match-on-is-per-batch). MITOMAP's differs from CIViC's because MT
variants carry no rsIDs, which is a fact about the source and so is recorded as one.
identity
property
¶
The cross-check's identity — match_on unless the provider states otherwise.
identity_refused_by_model ¶
None when model accepts these identity cells, else the model's own complaint.
Constructing the model is the oracle, and authoring_requirements is not. That was measured
rather than assumed: authoring_requirements("variants.csv") answers
any_of: [['rsid'], ['chrom','start']], a grammar that cannot express VariantRow's third
clause — ref/alts require chrom and start. So a guard built on it accepts
{"rsid": "rs1", "alts": "G"}, which the model refuses and a compile would refuse, and the
partial coordinate rides through (@identity-whole-or-none). One provider had exactly that
shape. authoring_requirements still answers the human-readable question below; it is not the
verdict.
No message parsing. Every non-identity field is pre-filled from IDENTITY_PROBE_FILLER with
values the model is known to accept, so any ValidationError reaching here is an identity
refusal and needs no inspection. The previous implementation branched on
"identifier" in message or "positional" in message or "chrom" in message — pydantic's rendered
text as an API, which moves on a dependency bump with no warning.
Source code in enricher/src/just_dna_enricher/drafting.py
missing_required ¶
missing_required(
table: str,
cells: Mapping[str, object],
stubbed: Sequence[str] = (),
) -> list[str]
The required cells this row does not carry, named for a human.
Reported, never used as the verdict — identity_refused_by_model is the verdict. This exists so
a warning can say which cells are absent, which the model's own message does not always spell.
Source code in enricher/src/just_dna_enricher/drafting.py
skip_reason ¶
Why this row is skipped, or None to draft it — the model's rule and the source's, in order.
The model is asked first so that a row failing both is reported as the identity problem it is, rather than as whatever the provider additionally wanted.
Source code in enricher/src/just_dna_enricher/drafting.py
licence_commit ¶
licence_commit(
*,
sources: Sequence[str],
spec_dir: Path,
dataset: str | None,
declared_use: str | None,
error: type[Exception],
layer: str = "annotation",
extra_datasets: Mapping[str, str] | None = None,
license_texts: Mapping[str, str] | None = None,
) -> Callable[[], None]
The licence merge as a callable, for draft.append_*'s before_commit (RM232).
The merge half of record_draft_provenance, and nothing else — no stale-label withdrawal, no
projection restamp. Those two need drafted and the provider's kind, both of which are answers
about the run as a whole rather than about one table, so they stay where they were, at the tail.
Why it is a factory and not a second call site. A drafter appends its tables through the
compiler's writer, which renames each one into place on its own; the licence row was recorded
afterwards, so a refused merge — a scaffold's <<REPLACE>> placeholder is enough — left drafted
rows in the author's tables with no licence record and the compile gate, which reads
sources.csv and nothing else, nothing to refuse on. Handing this closure to every append binds
the row to the commit of each table it licenses. One body rather than a copy per drafter, because
a copy per drafter is what RM228 existed to remove.
Never-clobber, so calling it once per table records one row.
The pre-flight is here rather than in each drafter (S98, RM231): building the closure reads
the licence table through require_sources_file, so a licensing.csv that does not load — a
scaffold's unreplaced <<REPLACE>> row is the case S98 was filed on — refuses before the first
append instead of after it. One place, so a new drafter inherits it rather than remembering it.
This makes a dry run refuse where it used to report, and that is intended. A drafter builds
the closure unconditionally, so --dry-run against a module with an unreadable licence table now
raises instead of printing what it would have written. A dry run exists to say what the real run
will do, and a dry run that passes while the real one refuses says the opposite; RM231 made the
same trade at the same seam.
Source code in enricher/src/just_dna_enricher/drafting.py
record_draft_provenance ¶
record_draft_provenance(
*,
provider: DraftProvider,
sources: Sequence[str],
spec_dir: Path,
dataset: str | None,
covered: bool,
drafted: bool,
declared_use: str | None,
error: type[Exception],
layer: str = "annotation",
extra_datasets: Mapping[str, str] | None = None,
license_texts: Mapping[str, str] | None = None,
stale_warning: Callable[[str, str | None], str]
| None = None,
) -> list[str]
Write this run's SourceRows if it covered anything, and withdraw a stale release label.
One function because they are one decision made twice. Three providers wrote the licence row
themselves; two of those never called withdraw_stale_dataset at all, so a CPIC or ClinPGx
re-curation left a row naming the older release with nothing able to notice (8.10 #5). Splitting
them is what let one half be forgotten, so the scaffold does not offer the halves separately.
sources is plural because a draft really can consult several. civic_draft records CIViC
and, when it asked it, the ClinGen Allele Registry — so the singular shape would have forced that
provider to keep writing its own rows and stay outside the scaffold, which is how the patchwork
started. This delegates to record_source_terms, the general primitive that already takes a
datasets map (RM222), rather than reaching past it to merge_sources_file.
covered is what this run covered — at least one row in the module's table because of this
provider, added now or recognised as already_present. A pass that contributes nothing writes
no row (@write-the-sourcerow): a --gene filter matching nothing leaves the module untouched,
and a licence row saying "this module uses X" would be a claim about a module that does not. That
is RM222, and it was written wrong here once already.
drafted gates the withdrawal separately and narrowly: a re-draft that added no row changed
nothing to be honest about, so there is no stale label to withdraw. The withdrawal targets the
first source only — the provider's own — because a secondary registry consulted along the way
does not own this module's release label.
Widening a module from a newer snapshot leaves the row naming the older release, because the
merge is never-clobber — right for a curator's hand-written terms and a false claim for
dataset. The stale label is withdrawn, never re-labelled: a module carrying rows from two
releases has no honest single label, and unknown is withheld.
Returns the warnings to append, rather than logging, so a caller keeps one reporting path.
Source code in enricher/src/just_dna_enricher/drafting.py
354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 | |
draft_digest ¶
Hash source's drafted table as it stands on disk, or None when there is nothing to hash.
None for a source that drafts nothing, for a table this module does not carry, and for one that
carries no rows — all three are "no copy has been established", which is the state that leaves a
check running.
Source code in enricher/src/just_dna_enricher/drafting.py
stamp_draft_digest ¶
stamp_draft_digest(
spec_dir: Path,
source: str,
layer: str,
*,
error: type[Exception],
) -> str | None
Record the current digest onto the (source, layer) licence row. Returns what it wrote.
This exists because merge_sources_file is never-clobber, which is right for terms a curator
may have hand-written and wrong for a machine-stamped cell that must track the table. Without an
explicit restamp a second draft's digest is silently dropped, the recorded one stays behind
naming a table that has since grown, and the skip dies permanently — in the safe direction, and
invisibly, which is what would make it a trap rather than a bug. withdraw_stale_dataset is the
same lesson on the neighbouring column, and it is the precedent for overwriting here at all.
Unlike dataset, this one re-labels rather than withdraws, and the difference is real: a
release label cannot name two releases, so a module spanning two has no honest value and unknown
is withheld. A digest has no such problem — it describes the table as it now stands, whatever
mixture of releases and hands produced it, so recomputing is always the honest answer.
Called unconditionally by a provider that wrote anything: a run that appended no row leaves the projection unchanged, so the restamp is a no-op rather than a special case to guard.
Source code in enricher/src/just_dna_enricher/drafting.py
drafted_unchanged ¶
Has every checked cell stayed as the drafter wrote it? Tri-state.
None— nothing recorded a digest for this source, so nothing was ever established. A module nobody drafted, one drafted before this shipped, or a table that has since been deleted.False— a checked value has moved since the draft. The row stopped being a copy, whoever moved it, and the cross-check has something real to compare.True— the projection still hashes to what the drafter recorded.
True alone is not grounds to skip a check. The digest describes this module's table, not
the source's release, so it is silent about currency: a matching digest against a newer
snapshot is a genuine comparison, not a tautology. The caller conjoins this with the release
check (clinical.tautology_reason's existing rule) and skips only when both hold.