just_dna_format.overrides¶
just_dna_format.overrides ¶
overrides.csv — the authored overlay that lies on top of a derived table (RM124, 0.7).
Every derived sidecar this workspace writes is machine-produced and, until 0.7, also hand-editable: the enricher merges rather than clobbers, so an existing row is authoritative and a re-run adds to it instead of replacing it. That rule exists because these tables are human-overridable by design, and its consequence is the single most important operational fact about a second pass — to re-derive a sidecar you delete it first, and deleting it discards every hand-curated row along with the stale ones. The 2026-08-12 cost amendment (CONSTITUTION Principle 9) names the class in its own words: a derived table that is both machine-written and human-overridable can be edited into a state that is not merely stale but a false claim, which wants a mechanism rather than a convention.
This is the mechanism. An overlay row records a correction beside the derived table rather than
inside it, so the derived files become pure build products — derived = f(source, overlay) — and
four things follow. Nothing is hand-edited, so re-derivation is non-destructive by construction
rather than by a wrapper being careful. A difference between a fresh row and a previous one means the
source revised, full stop. The reason for a correction travels with the module, which is what
makes this a record rather than a knob. And the terminal state becomes detectable: an overlay row
that no longer changes anything means the source caught up.
The key is (table, subject, member, field). subject names the group a derived row belongs to
and member discriminates within it — locus_index for resolution.csv, population for
frequencies.csv, assertion_id for gene_validity.csv, empty for a table whose subject already
identifies exactly one row. One member column serves every covered table; a key that differs by
table is a rule every reader re-derives, differently.
Empty member on a grouped table is group-scoped for update, and refused for insert and
suppress. The asymmetry is deliberate. A group-scoped correction is a coherent thing to want —
every locus this key resolves to has the wrong source — and it is recoverable if wrong. A
group-scoped suppression silently drops every locus for one variant_key when the author almost
certainly meant one, and it is not recoverable by reading the result, because the rows are simply
absent. An insert under an empty member on a grouped table is not merely destructive but
incoherent: the row it would create has no member value to carry, so nothing could ever match it
again. The destructive and the incoherent operations refuse the wildcard; an author who genuinely
wants a whole group gone writes one row per member, and the count is small by construction.
The dead end that asymmetry leaves, stated rather than papered over. A derived row whose member
column is null cannot be named by any override, because an empty member cell already means "the
whole group". A gene_validity.csv row the source published no assertion_id for, and a
clinical_assertions.csv not_found row — whose null variation_id that model's own docstring calls
a value, the record that the archive was consulted and holds nothing — are both reachable by a
group-scoped update and unreachable by a suppress. A sentinel spelling for null was not invented
for it: that is a second key grammar in the column whose whole purpose is that there is only one, and
the remedy an author actually has (correct the row, or re-derive the table) costs them nothing. The
refusal message names the case rather than advising the impossible.
Dependency-light like every other module here: pydantic plus the leaf vocabularies, no polars.
OverlayTarget
dataclass
¶
OverlayTarget(
model: type[BaseModel],
subject_field: str,
member_field: str | None,
vindication: str,
vindication_reason: str,
)
What an overlay row means when it names one derived table.
subject_field is the column whose value an overlay's subject carries; member_field is the
within-group discriminator, or None for a table whose subject already identifies exactly one
row. Both are column names on model, checked by a walked test rather than trusted — a
registry is only as good as the guard that walks it (@registry-completeness).
vindication is a member of VINDICATION_READINGS, and vindication_reason says why, as a field
rather than a comment, so the decision for each table travels with the table (RM290).
OverrideRow ¶
Bases: AuthoredModel
One authored correction to one cell (or one row) of one derived table.
An AuthoredModel and not a fact model, which is the whole distinction: a human writes this,
every other row in the tables it names is machine-written. So it carries the reserved-
namespace guard and the shared field validators, and it is inside content_signature — a
correction is part of what the module says.
Inside it by its first six columns only (S87). table/subject/member/field/
operation/value say what the correction is, and two modules differing in any of them
assert different things. reason/decided_by/decided_at say why, who and when — provenance
beside the claim, which nothing in the compiler or enricher reads and which is exactly the cell
an author improves on a second pass. They are marked OUTSIDE_CONTENT_IDENTITY, which
content_signature alone reads: rewording a reason is a patch, as fixing a README caveat is
(S25), rather than a new content identity for byte-identical data. Everything else still
carries them — overrides.parquet, the raw-bytes hash of overrides.csv in manifest.inputs
and therefore the verification binding (editing a reason still un-closes a module, and still
moves artifact.digest), reverse, and every model_dump() writer. Not exclude=True:
the consumer's candidate was the stamped-column mechanism, and it would have emptied reason
in draft._authored_dump and every other writer that serializes a row through model_dump()
— a drafted overlay row the compiler then refuses for the blank it itself wrote.
The line is drawn here and not on variants.csv, where curator/method are inside the
signature: those are folded from defaults: as content (RM37) and moving them would re-key
every published module, so that asymmetry is carried rather than repaired. Decided with the
maintainer on 2026-09-03, before 0.7 was cut, because no published module carries an overlay
and the window in which excluding these moves nothing is the release.
overlay_coherence_errors ¶
The file-level rules, which no single row can answer: one operation per key group.
A duplicate (table, subject, member, field) is the model's own _KEY_FIELDS and is caught by
the compiler's ordinary duplicate-row check, so it is not repeated here.
Source code in schema/src/just_dna_format/overrides.py
apply_overrides ¶
apply_overrides(
table: str,
rows: Sequence[BaseModel],
overrides: Sequence[OverrideRow],
*,
defer_unmatched: bool = False,
) -> tuple[list[BaseModel], list[str], list[str]]
Lay the overlay rows naming table over its derived rows. Returns (rows, errors, warnings).
The returned list is the table as the module asserts it, which is what every downstream
reader wants: the parquet is built from it, the fact signatures and resolution_signature are
over it, and reverse_module emits it. Overlay rows naming another table are ignored, so a
caller passes the whole file and this picks its own out.
Order is load-bearing (parquet bytes depend on it), so:
updateedits a row in place and moves nothing;suppressremoves a row and reorders nothing that remains;insertappends at the end of its subject's group — after the last row already carrying that subject, or at the end of the table when the subject has no group yet — in the order the overlay rows appear. Placement is a function of the overlay's own authored order, which the round trip already preserves, rather than of a sort over values a corrected cell could move.
A no-op is not a finding, and that is forced rather than tidy. After reverse_module the
derived table is post-overlay, so on the second lap update-already-equal, insert-already-present
and suppress-already-absent are all three true of a perfectly healthy module. Reporting any of
them would make a module and its own round trip disagree on manifest.compilation.warnings, a
published field. An update matching no row is the one mismatch an overlay operation cannot
manufacture for itself, because an update never creates a row.
Two warnings come out of here, and the second is a record rather than a mismatch. A suppress
removes a row and leaves no trace of the removal anywhere in the build product, so RM131 has it
say so — counted over the overlay's rows and aggregated by reason, which is what keeps it from
being the lap-1-only line the paragraph above rules out. _suppression_warnings argues both.
The overlay is stable; the derived table under it was not, and RM137 is that repair. For the
two tables in LOSSY_OVERLAY_TABLES the compiler rebuilds from something narrower than the file
it read, so an update naming a dropped row matched on lap 1 and reported on lap 2 — the exact
disagreement this function exists to avoid, arriving through the derivation rather than the
overlay. Applying the overlay after the drop is still refused (the checks must see what the module
asserts) and reverse still has no source for rows the artifact does not hold. What changed is the
finding: defer_unmatched=True suppresses the warning here, and the caller splits it with
update_targets + classify_update_targets once it can answer could this module carry that row
at all — a property of the module, so it says the same thing on both laps. The six other tables
rebuild whole and keep the message below unchanged.
Source code in schema/src/just_dna_format/overrides.py
645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 | |
update_targets ¶
update_targets(
table: str,
rows: Sequence[BaseModel],
overrides: Sequence[OverrideRow],
) -> list[tuple[tuple[str, str], bool]]
Every update target on table, with whether it reached a row — the raw finding (RM137).
Matched targets are returned too, and that is the whole point. The unreachable finding has to fire whether or not the row is there today, or it is lap-dependent all over again: on lap 1 an uncited literature row is present and the update matches, and on lap 2 the row is gone. Reporting only the unmatched ones reproduces exactly the defect this item is about.
Exists so a caller can classify the finding where the inputs to do so are in scope, which is
not where the overlay is applied. apply_overrides runs early, before any check reads a row; the
question could the artifact ever have carried this row needs studies.csv and the citing tables,
which both validate_spec and compile_module load later. Pass defer_unmatched=True to
apply_overrides and call this instead, on the pre-overlay rows — apply rebinds its input, and
an insert earlier in the same overlay would otherwise make a later update look matched.
Matching is apply_overrides' own, through the same _canonical_key_cell/_cell pair rather than
restated here: two statements of one rule drift, and this one would drift silently — the caller
would report a subject as unreached that the apply had just corrected.
Source code in schema/src/just_dna_format/overrides.py
classify_vindicated_answers ¶
classify_vindicated_answers(
table: str,
targets: Sequence[tuple[tuple[str, str], bool]],
) -> list[str]
An overlay answer whose conflict the archive has since resolved — the author was right (RM117).
This is the one trust signal in the format that is available nowhere else, and it costs nobody a decision. An author records a judgement against a contested subject; later the authorities agree, the enricher stops writing that subject into the record, and the overlay row reaches nothing. On every other table that state is ambiguous. Here it is not: the record holds contested subjects only and is rewritten whole, so leaving it means the contest ended.
It replaces a message that was actively misleading, which is why it is worth a code of its own. The generic finding offers the subject may be mistyped, or the correction may be aimed at a row the compiler drops — put to an author in the one case where their judgement was vindicated.
It says nothing about who was right about the biology, and the wording is careful: the authorities agreed, and the author's row is now unnecessary. That the archive moved toward the author is an observation about the record, not a verdict — the same restraint the concordance tables keep everywhere else.
Source code in schema/src/just_dna_format/overrides.py
classify_update_targets ¶
classify_update_targets(
table: str,
targets: Sequence[tuple[tuple[str, str], bool]],
target_survives: Callable[[str], bool] | None = None,
) -> list[str]
Split "this correction reached no row" into its two answerable readings (RM137).
The problem this solves is that the old single warning was not stable across a round trip. An
update naming a row the compiler drops matched on lap 1 and reported on lap 2, so a module and
its own compile → reverse → compile disagreed on manifest.compilation.warnings — a published
field, and one RM126 has since made load-bearing.
"Count it over the overlay's own rows" is the decision, and this is what that means in code.
Counting the overlay's update rows outright is a tautology that fires on every healthy module
(@tautology-zero); counting the ones that reached nothing is the lap-dependent original. The
stable quantity is a property of the target: can an artifact of this module carry that row at
all? That is computable from data which survives the round trip — studies.csv for a citation,
the positioned rows for a locus — so it answers the same on both laps whether or not the row is
there to be matched.
target_survives is the caller's predicate over the subject, and it is three-valued by absence:
None means the caller could not ask, and then this behaves exactly as before — one warning naming
every reading, withholding rather than accusing. That is the path the six non-lossy tables take.
Neither reading is "a typo", and that is the correction this makes to the entry's own framing.
A mistyped pmid is also an uncited one and a mistyped variant_key is also an unpositioned one, so
a mistake lands in the unreachable bucket rather than the reachable one. What the reachable
bucket really means is narrower and more useful: the subject is cited or positioned, so the
artifact could carry the row, and the sidecar simply does not have it — re-run the enricher.