Skip to content

0.7 design thread — the deferrals that were waiting on a decision, not on the world

What this is. Stage 3 of the design cycle for the 0.7 line. ROADMAP_0_7.md holds items that are legal in a minor and were not taken into 0.6, each waiting on a design question, a corpus, or a caller. That taxonomy is the whole of this round's sort: an item waiting on a corpus or a caller is not ours to decide, and an item waiting on a design question has been waiting for somebody to make one. This document makes them, per item, with the maintainer.

Same convention as PROPOSAL_0_6_PT2.md: each item records the problem in plain terms, the facts established while deciding it — four of which change what the roadmap entry says — the decision, the repairs rejected and why, and the charter check. Two items dissolved into others while being decided, and one grew an obligation nobody had filed.

Status. Decided 2026-08-28, on the 0.7 branch, before implementation. All twelve items have since shipped, across four waves on that branch; this document is kept as the decision record and is not rewritten to match. Where it and the code disagree the code won, and the corrections made during the build are marked inline where they apply.

Concluded on 2026-08-31, and it no longer wins over anything. While items were landing, this document beat whatever ROADMAP_0_7.md and ROADMAP.md still said about them. All twelve have landed now, so every entry is in ROADMAP_HISTORY.md and that is the file to read for what shipped; the last four moved there on 2026-08-31, having sat in the forward-only files with a SHIPPED banner. ROADMAP_0_7.md closed the same day into history/, and what was still waiting became ROADMAP_0_8.md. This document stays here beside the other five concluded threads: it is the decision record, kept for the reasoning and the refusals, and is not rewritten to match the code — where the two disagree the code won.

The release. 0.7 is the minor that was left uncut on 2026-08-27 — the S63–S74 batch already on this branch, whose next-minor markers resolve here. Everything below joins that batch and cuts with it as 0.7.0 across all three packages. Nothing in this document is gated on a version: every item is additive under Principles 3 and 8, and legality stopped nothing this round.

Scope. Twelve items decided: six from ROADMAP_0_7.md and six that were sitting in ROADMAP.md as a minor, release undecided. Eleven build and one closes into another — RM83, whose premise stopped holding. The succession filed alongside them is RM135 and is not one of the twelve.

Two items were decided after the round closed, and both are addenda rather than a thirteenth and a fourteenth — RM140 and RM152. Each arrived on 2026-08-31 against a 0.7.0 that was bumped and not yet cut, so each ships inside the same number. They are recorded here, in this file's idiom, because a decision of the same shape as the twelve belongs beside them rather than in a thread of its own — and they are dated and set apart because the round itself is closed and the count above is a fact about that round. Every "twelve" in this document means the 2026-08-27/28 round.

And the practice itself was reviewed when the second one arrived, because a closed file had started to read as a closed door. The maintainer's objection, recorded because it outlives this item: while the release is an open construction site — three pyproject.toml files at 0.7.0, last tag v0.6.6 — anything landing in it is to be decided here, and a deferral to the next minor is a case-by-case call rather than the default a concluded document quietly imposes. Otherwise unblocked minor work accumulates deferrals for a reason that is bookkeeping rather than design. The rule, stated once: a proposal file is closed against re-opening its own decisions, never against recording a new one taken inside the same uncut release. An addendum is how that is written down, and there may be as many as the release earns.

RM134 was pulled in on 2026-08-28, after the other eleven were decided, and it is the largest single piece of work here. It also reaches back into two of them: its concordance machinery supersedes RM130's record shape, and its new checks are what expose the warnings problem in RM126. Both are amended below rather than left to collide during implementation. Nine further items stay deferred on a gate that is not ours to open, restated per item in Not taken. One new item (RM135) was filed by the round itself, and RM134 arrived alongside it from an unrelated thread and is out of scope here.


The sort, and the rule it used

One rule, and it is the file's own taxonomy read as a scheduling instruction:

An item waiting on a corpus or a caller is not a candidate. An item waiting on a design question is a candidate, and the decision is the work.

That is narrower than PT2's three rules and does the same job, because PT2's rule 2 — a new expensive feature waits for the demand that would fix its shape — is a statement about which open question an item has, not a separate test. Demand fixes a shape nobody has fixed. Where the shape is already fixed by a decision one release old, demand has nothing left to contribute, and RM132 is the item that makes the distinction load-bearing: a full-cost authored column that would normally wait is taken here because RM47 already decided its shape for a structurally identical table, so the risk P9 prices was spent a release ago.

What sizes the release. Legality, as always, and it settles nothing: every item is a new optional column, a new optional table, a new report, or a flag. Severity orders the queue inside it, and by severity this round is led by a charter debt (RM126) and a mechanism the charter itself named as missing (RM124).

The two dissolutions are the most valuable output of the round, and both went the same way: an item was filed as a missing operation, and deciding a neighbour removed the condition that made the operation necessary. RM83 is one and RM128's central ask is the other. Neither was argued down — the premise stopped holding.


Decisions

RM126 — nothing tells a consumer what a release changed about compiled output

Severity medium-high · Owner format (record + needs_recompile + roster) + compiler (the sweep) · Entry ROADMAP_HISTORY.md § RM126 · Motivating case S62 (just-dna-registry), narrowed by S65

The problem

Principle 3, as amended 2026-08-21, says a corrected derivation may ship in any release but never silently — each release declares its corrections, readable offline and without recompiling. No such channel exists, so the charter names a surface that is not there. This is the one item in the round that is owed rather than offered.

A consumer holding a stored artifact can ask two questions and not the third. Is the stored input still legal? — validate_spec answers ok. Was this compiled under a contract-incompatible compiler? — compare versions, and a patch is compatible. Neither is would recompiling this artifact produce different output than the stored one? Answering it today means enriching into a scratch directory and recompiling, which is the operation rather than a triage for it.

Facts established while deciding it

The sweep in the entry stands and is the design's own prototype: all sixteen reference_examples/ compiled under v0.6.1 and 0.6.6 from byte-identical spec inputs, an interval that is entirely patch releases — 16/16 changed at least one published manifest field, 10/16 moved artifact.digest, 0/16 moved content_signature. The digest movement is RM120's authored curator column growing studies.parquet by 257 bytes, so a parquet schema moved across a patch interval. Six changed a published, indexed manifest field with both hashes byte-identical.

Authored identity held throughout, which is the charter working as designed and is exactly why no existing surface can see any of this: a digest comparison, a signature comparison and a revalidate all correctly report no change while an indexed field goes stale.

The decision

Build the shape settled in the S62 thread, in full, plus the roster S65 asked for.

  • just-dna-format carries the record and the function. A static table of per-release records and a pure needs_recompile(compiled_under, current) over it. Four axes per record — parquet schema, parquet bytes, content_signature, and the set of changed manifest fields — plus the declared correction-versus-addition split, which is the half no diff can compute: only the person fixing the bug knows whether the stored value was wrong (stats.genes) or merely absent (curator). A static table and a function is pydantic-only work, and format is the tier every consumer has.
  • Intervals compose as a union over the releases in (a, b]. Storage is linear rather than O(releases²), and moved-and-moved-back still counts as moved, which is the right reading for staleness. The interval shape is also what gives S65's convergence requirement for free — the interval from a version to itself is empty, so an automated sweep cannot mint a fresh PATCH every run forever. That property is load-bearing and is recorded as such, because a field-keyed or latest-known-defect shape would not have it.
  • Unknown is a state, never an empty result. Asked about an interval the installed package has no record of, the answer is cannot say. Without it the surface is worse than nothing, because a consumer would stop recompiling on the strength of a silence. House tri-state, and None is never False.
  • just-dna-compiler carries the sweep instrument, since producing a record means compiling. Backfill 0.6.1→0.6.6 by measurement with the harness that produced the numbers above; older intervals stay honestly unknown.
  • The gate runs in the bump → uv sync → tag sequence, not as an ordinary test, because it needs the previous release actually installed. A release whose sweep shows a changed field with no declaration covering it fails. That is what stops this becoming the hand-kept map everyone agrees it must not be. A release where nothing moved records a measured zero with its evidence, never silence.
  • The roster ships beside it, and it is the half that shrinks the item rather than growing it: a published set naming which manifest fields are pure functions of the authored rows, so a consumer can recompute the current answer from stored inputs using spec_tables (RM116) and module_stats (RM121) rather than consulting the table at all. Its boundary is conditional and the condition must be published with it: validate_spec computes stats over the full row set while compile_module re-derives over the survivors only when the symbolic-allele drop removed something, so a recomputation from authored rows is the pre-drop side and manifest.stats legitimately disagrees with it, permanently, for a module that lost the sole row naming a gene. The condition is now checkable, because compilation.dropped_rows shipped 2026-08-24.
  • Scope v1 to compiler-derived outputs and say so. Enricher-side outputs stay unmeasured, which is not the same claim as unchanged.
  • compilation.warnings gets its own axis and is outside the set that drives recompile. Found while folding in RM134, and it would have swamped the instrument on its first release: compilation.warnings is a published manifest field, RM131 restructures that channel, and RM134 adds checks — so the sweep would report a manifest field changed on essentially every module in 0.7, and a registry acting on it mints a PATCH across the whole catalogue for a message change. That is the false-positive class S65 made convergence a hard requirement over. compilation.compiled_at is already excluded as noise and warnings cannot simply join it, because a new warning can be a real signal — hence a separate, declared axis. RM131's carried split is the discriminator that makes it decidable: a finding the author cannot clear moving is noise, one they can clear appearing is not. The two items are coupled and the implementation order below reflects it.

What this does not touch: SpecRow.needs_upgrade is self.upgraded() != self, computed over authored row content, and it is a hard filter in the marketplace — a needs_upgrade version is not offered to a consumer at all. A warning never touches an authored row, so no warning change can flip it. Verified rather than assumed, because the failure mode would have been modules silently vanishing from a catalogue.

Repairs rejected

  • A should_rebuild verdict. The same fact costs a registry an immutable PATCH and just-dna-lite a free rebuild, so the decision is the consumer's and only the fact is ours. The declared correction flag is not this: it is a fact about whether a value we published was wrong, which is upstream knowledge only this repo holds, and the per-axis breakdown stays exposed underneath it. A bare boolean with nothing under it would deserve the objection exactly.
  • A hand-kept per-release map. The defect wearing a public name, and @registry-completeness is the tag: five of the six RM104–RM111 fixes were a derived value restated by hand. The measurement forces the declaration rather than the author remembering to write one.
  • A field-keyed table. Fails S65's convergence requirement — a hint firing for a version compiled by the exact compiler now installed mints a version number every run, forever.
  • Replacing the consumer's recomputation probes. Scoped for coexistence at the reporter's request: we state what a release did, they check what a specific stored artifact says, and the two fail differently.

Charter check

P3 — this is the amendment's channel, so the item exists to discharge a rule rather than to test one. P4 — reads identity, moves neither hash. P6 — the axis names are a closed vocabulary, frozenset[str] plus a validator. P9 — a static table plus a pure function in format, and an instrument in the compiler: no authored layer is touched at all, so this is the cheapest layer the charter prices.


RM124 — an author's correction to a derived table has nowhere to live except inside it

Severity medium-high · Owner format (schema) + compiler (apply + reverse) + enricher (the derived-table writers) · Entry ROADMAP_0_7.md § RM124 · Motivating case S60 (just-module-creator)

This is the keystone of the round. RM83 closes into it, RM130 was blocked on one of its questions, and RM128's central ask thins because of it.

The problem

Every derived sidecar is merge-not-clobber: an existing row is authoritative and a re-run adds to it rather than replacing it. The rule exists because these tables are human-overridable by design, and its consequence is the single most important operational fact about a second pass — to re-derive a sidecar you delete it first, and deleting it discards every hand-curated row in it along with the stale ones. reference_examples/cyp2c9_warfarin_grch37/ carries three hand-authored source="manual" rows in resolution.csv that no re-run reproduces; literature.csv's loss includes a curator's deliberate blank, which a merge cannot distinguish from an absent value in the first place.

The 2026-08-12 cost amendment names this class in its own words — a derived table that is both machine-written and human-overridable can be edited into a state that is not merely stale but a false claim, which wants a mechanism rather than a convention. RM45 discharged it for exactly one table by making verification.json unwritable by hand. Nothing discharges it for the seven where overriding is the intended feature.

A consumer built the conservative exit — a non-destructive capture / verify / delete / re-derive / classify / reapply wrapper — and it stops precisely where the problem predicted: a subject present in both copies with a differing fact is either a cell the author edited or a revision the source published, and with two data points there is no third to separate them. Their tool reports and refuses to resolve, which is the honest outcome and is a symptom.

The decision

overrides.csv: an authored overlay table that lies on top of a derived one and is never merged into it. The derived files become pure build products — derived = f(source, overlay) — and the four consequences the reporter named all hold: nothing is hand-edited so re-derivation is non-destructive by construction rather than by a wrapper being careful; a difference between a fresh row and a previous one means the source revised, full stop; the reason for a correction travels with the module; and the terminal state becomes detectable for free, since an overlay row that no longer changes anything means the source caught up — evidence that an authored judgement was later vindicated, available nowhere else in this format.

Columns. table, subject, member, field, operation, value, reason, decided_by, decided_at. reason is required, and that is what makes this a record rather than a knob.

Operations are update, insert and suppress — a closed vocabulary, frozenset[str] plus a validator per P6. update corrects one field of a derived row; insert supplies a row the source has no answer for, which is what the three source="manual" rows in the flagship module actually are; suppress removes a derived row the author rejects. An inserted row is written as several overlay rows sharing (table, subject, member), one per field, so the table has one shape rather than two.

The keying decision, which the entry left open and the operations make sharper

(table, subject, field) is not enough, and the table it fails on is the flagship case. RM115 published resolution.csv's merge rule as subject, not a uniqueness constraint: one variant_key legitimately resolves onto several loci, locus_index orders them, and a pass replaces the group whole. So a subject names a group of rows and an overlay keyed on it cannot say which locus it corrects. gene_validity.csv needs the same care in the other direction, its key having two levels (assertion_id else the gene's grain).

The key is (table, subject, member, field), where member is an optional within-group discriminator whose meaning is the named table's own ordering or sub-key column — locus_index for resolution.csv, assertion_id for gene_validity.csv, empty for every table whose subject already identifies exactly one row. That single column serves all three tables without a per-table key grammar, which is the thing to avoid: a key that differs by table is a rule every reader re-derives.

An empty member on a grouped table is group-scoped for update, and refused for suppress. The asymmetry is deliberate and is the whole reason the operations made this question sharper. A group-scoped correction is a coherent thing to want — every locus this key resolves to has the wrong source — and it is recoverable if wrong. A group-scoped suppression silently drops every locus for a variant_key when the author almost certainly meant one, and it is not recoverable by reading the result, because the rows are simply absent. The destructive operation refuses the wildcard; an author who genuinely wants a whole group gone writes one row per member, and the count is small by construction.

What P7 makes of a build product, and why no previous_value column is needed

reverse_module emits the post-overlay derived table plus the overlay. The entry flagged that this means the overlay applies twice and the fixed point has to be checked rather than assumed — which is right, and checking it is what P7 requires of every derivation anyway. It passes, because all three operations are idempotent set operations by construction: update to a value already present is a no-op, insert of a row already keyed (subject, member) is a no-op, suppress of a row already absent is a no-op. row.upgraded().upgraded() == row.upgraded() is the same property the charter already demands, and it holds here without a new column.

The alternative — reverse emitting the pre-overlay table so the apply happens exactly once — would require the overlay to record the value it replaced, so reverse could un-apply. That is a derived cell inside an authored table, and it rots the moment the source moves: precisely the staleness licensing.withdraw_stale_dataset had to be built for, and the reason RM71 refuses a comment column. Idempotency buys the round trip at no schema cost.

Idempotency proves value identity; the round trip is compared on bytes, and row order is load-bearing — parquet bytes depend on it, and authored row order is preserved through compile → reverse → recompile. So insert owes a placement rule, and it is stated here rather than invented by the implementer: an inserted row is appended at the end of its subject's group, in the order the overlay rows appear. That makes placement a function of the overlay's own authored order, which the round trip already preserves, rather than of a sort over values that a corrected cell could move. update and suppress need no rule — one edits a row in place and the other removes one, and neither reorders what remains.

The fact signatures and resolution_signature are over the derived tables as they stand, therefore post-overlay, which is right: a signature should describe what the module actually asserts. The overlay is authored input, so it is inside content_signature — and no published module's content_signature moves, because an absent optional table contributes nothing, exactly as an unset optional column does.

Whether merge-not-clobber survives: it does not, and what replaces it

Dropped for the seven covered derived tables. That is the prize, and it is what makes this 0.7 rather than a smaller thing: it changes what a re-run does to every module already published.

The covered set, enumerated — and the count was wrong until 2026-08-28. This section and the entry it inherits from both said six and neither ever listed them, which is the shape of defect this repo has a tag for: a number restated by hand beside the thing that could have produced it. Code has seven merge-not-clobber derived sidecars, and a build against the uncounted number would have left merge-not-clobber alive in exactly one table — silently, since nothing computes the set.

table writer
resolution.csv enrich._write_resolution_csv
frequencies.csv frequencies._write_frequencies_csv
gene_metrics.csv gene_metrics._write_gene_metrics_csv (two passes share it — enrich_gene_metrics and clingen.enrich_dosage_sensitivity)
gene_validity.csv gene_validity._write_gene_validity_csv
clinical_assertions.csv assertions._write_assertions_csv
literature.csv literature._write_literature_csv
gwas_effects.csv gwas._write_gwas_csv

sources.csv/licensing.csv is out: it has its own merge path (licensing.merge_sources_file) and is not a table where overriding is the intended feature. gwas_effects.csv is the one the miscount most likely dropped — it is the newest pass (RM90) and its ENRICHER.md section was appended last — which is exactly why the set is enumerated here rather than counted again. Derive the covered set from the writers, never from this table.

What replaces it is narrower than "every run re-asks", which would cost the full resolution time on every pass and is not what dropping the rule means:

  • Gap-filling stays the default. A re-run fills subjects with no row and leaves recorded rows alone, as today — the difference is that a recorded row now carries no authored content, so leaving it alone risks nothing.
  • Full re-derivation becomes free, because the author's corrections are in the overlay and cannot be lost. rm plus a re-run is the crude form and now costs nothing; enrich --rederive is the form that keeps a baseline, and it is where RM83's residue lands (below).

The enricher's docstrings are part of this change, not a follow-up. Every writer of a covered table states, in the docstring a reader outside this repo actually has, that the artifact is now purely derived, that hand-authoring it is not expected, and that overrides.csv is where a correction goes. A printed contract is the surface this repo has measured the cost of getting wrong, at 3,038 rows for the 0-versus-1-based start description, and the sweep belongs in the same commit as the behaviour.

The overlap with provenance.json, and the succession

ProvenanceItem.outranks: dict[str, str] — {column: why}, an authored cell outranking a source, with prose — is the same concept one table over. The entry says to decide before either grows a second field, and the decision is a succession rather than a merge:

Both mechanisms stand in 0.7, the duplication is stated rather than hidden, and the unification is filed for 1.0. The rule that survives is the simpler one: the existence of an override in an authored table auto-beats the derived value, with no separate declaration to write. By Venn diagram the new logic is a partial superset of what outranks allows and reaches it more directly, and once an author controls both derived overrides and authored tables the outranks knob's need evaporates. Filed to the 1.0 cleanup tracker, since a removal is major-only under P3.

0.7 emits no deprecation warning, and that is a P3 decision rather than caution. A deprecation belongs in a minor only where its audience can act on it. outranks covers an authored cell beating a source; the overlay covers derived tables. Until the overlay's semantics reach authored tables — which is the 1.0 work — an author warned off outranks has nowhere to go, and a warning nobody can clear is noise rather than notice. The succession is documented in SCHEMAS and on the tracker; the warning follows the replacement.

Repairs rejected

  • Per-consumer application of the overlay. If each downstream tool applies its own, two consumers compiling one spec directory disagree about what the module says and the artifact stops being a function of the spec. The reporter's tier argument was accepted and is not an open question.
  • Update-only, with manual rows left inside the derived file. The sidecar is then not a pure build product after all, and merge-not-clobber has to survive for exactly those rows — which re-opens what this item exists to close.
  • A per-table key grammar. Each table's own key spelled into the overlay reads as precise and is a rule every consumer re-derives, differently. One member column, whose meaning the named table fixes, is the same expressiveness with one thing to learn.
  • Merging outranks into the overlay now. Legal only if the old form survives as a working derived alias (P3), which means shipping the unification and its compatibility shim in a release where the overlay has no users yet. The succession costs a documented duplication for one major line and buys the decision being made against real usage.

Charter check

P1 — data, not code. P2 — authored input, not a fetch; the compiler reading it is doing what it already does with every other authored table. P3 — a new optional table is additive and minor-legal, and no published module's identity moves. P5 — the overlap with outranks is the principle's own question, and it is answered by a dated succession rather than left to erode. P7 — the round trip holds by idempotency, proven by test rather than asserted. P8 — demotes nothing. P9 — full cost, and the amendment is what invited it: a mechanism rather than a convention for the class the amendment names.

The coordination step, which is real and routine

just-dna-registry rebuilds a spec directory from RECOGNIZED_SPEC_FILES, built from SPEC_DATA_FILES — a hand-kept mirror of our table constants — and a name missing there is a file dropped on the next re-publish, which is how licensing.csv was lost before their 0.16.2. overrides.csv needs one entry added there. It is the same one-line change every new table kind already needs, and it goes in the integration notes rather than being discovered.


RM83 — a derived sidecar can only be refreshed by deleting it

Severity medium-high · Owner enricher · Entry ROADMAP_0_7.md § RM83 · Disposition closed as filed, contingent on RM124 landing

The decision, and it is not a build

RM83 is dissolved by RM124 rather than argued down, and both of its halves go the same way. The --refresh command with its three open questions is not built; nothing named in the entry — no proposed table beside the current one, no diffs file, no new command with a lifecycle — is built either.

  • The refresh half stops existing. The entry's problem is "to re-derive a sidecar you delete it first, and deleting it discards the curator's rows." Once the derived files are pure build products with the overrides in the overlay, there is nothing inside a sidecar to preserve, so rm costs nothing and needs no command wrapped around it to be safe.
  • The drift half stops being unperformable. The entry's problem is that merge-not-clobber means a re-run never re-asks about a recorded row, so a source that silently revised an answer moves no fetched_at, no fact signature and no digest — making §5.1's canary an instrument that cannot fire on its own, because detecting drift is the delete-and-re-derive that discards the overrides. With the discard harmless, a full re-derivation is an ordinary operation and the canary fires from it.

The blocking question is answered rather than deferred. The entry named it: on most sidecars nothing records that a row was overridden, so "re-derive the machine rows and keep the overrides" is not implementable, because the tier cannot tell a curator's edit from what the source said last time. Under the overlay the tier never has to tell them apart — the edit is recorded by construction and the derived row carries no authored content at all.

The residue, and it is a flag rather than a command

A full re-derivation that keeps a baseline can report what moved, and this is where the round's honesty has to be exact rather than comfortable. The report is only free where a baseline exists. rm followed by a re-run destroys the old values before the fresh ones arrive, so nothing holds both sides and no report is possible — that path re-derives silently and correctly, and the proposal says so.

enrich --rederive is the path that keeps one, and it composes with RM128's transaction below rather than adding machinery of its own: the run stages a complete fresh table beside the current one and commits by rename, so both files exist at the commit boundary and the diff is genuinely free. What it reports is which recorded subjects changed value — the canary, performed.

This is a flag on an existing command, not RM83's --refresh. The distinction is not cosmetic: what made RM83 a design item was three open questions — which tier owns it, what it does with a difference, and the blocking one about recording overrides. All three are now answered by decisions made elsewhere in this document (the enricher, because it fetches; report and commit the fresh side, because the author's corrections are safe; and the overlay). What is left is a re-derivation mode on the command that already does the derivation, which is the shape the entry would have reached had its blocker not existed.

Bookkeeping

The entry moves to ROADMAP_HISTORY as closed, not shipped, in the commit that lands the overlay — not before. Until RM124 is in the tree the problem it names is still real, and a closure recorded against an unlanded dependency is the kind of bookkeeping that makes a ledger untrustworthy.

Repairs rejected

  • A diffs file or table tracking what moved between passes. It is version control with no consumer, beside the version control the author already has, over a file that is now regenerable from source plus overlay.
  • A pass that applies the newer value. Rewriting an authored or curator-set cell destroys the evidence of the upstream change. Still the rule, and the overlay does not soften it: an overlay row is the author's answer to a difference, never the tier's.
  • Re-asking every subject on every run. Dropping merge-not-clobber does not mean this, and reading it that way would put the full resolution time on every pass to buy drift detection nobody asked to run continuously.

RM130 — a check's findings are counted and not kept, so a conflict has no name to act on

Severity medium · Owner enricher · Entry ROADMAP_HISTORY.md § RM130 (shipped 2026-08-28; it left ROADMAP.md when it landed) · Motivating case S70 (just-module-creator)

The problem, and what was blocking it

The observability half shipped 2026-08-24: clinical_significance writes a detail grouped on opposed, and _findings_warning says at validate/compile that a record reports non-zero findings. What is left is the sidecar — a derived table carrying the authored value, the source's value, and whether the two are opposed or merely different — and it was deliberately unshipped because it is a third table in the overlay/outranks family, so shipping it would pre-empt RM124's question 2.

That question is now answered, so this ships.

The decision

A derived conflict table, keyed (variant_key, genotype), positioned as the overlay's input side. The lifetime follows directly from the succession decided above: a conflict is a question and an overlay row is the answer. An author reads the conflict table, decides, and writes an overrides.csv row — which is also what makes the terminal state visible, since a conflict that stops being reported means the archive caught up with the author.

Amended 2026-08-28 when RM134 was pulled in: this table is RM134's concordance record, and it ships in that shape rather than in a ClinVar-shaped one. The row carries authority_concordance and authored_position, and each authority's own call lives in a paired detail table keyed (variant_key, genotype, authority) — the shape argued in § RM134, which survives five authorities without a key change. The version this section originally decided — one row carrying "the authored value, the source's value" — is superseded, and the reason is worth keeping: ClinSigConflict names its authority in a field (clinical.py:74, clinvar: str), so shipping that shape and meeting a second authority one item later would have cost a key change or a retype, which P3 reserves for a major. Two items landing in one release is what caught it.

It promotes the mechanism that survives to 1.0. The table's own documentation, and the warning that points at it, name overrides.csv as where an answer goes — never outranks. Steering authors onto the succession's winning side from the start is the cheapest form of a migration, and it costs a sentence.

The key cannot be a bare variant_key, and this was settled rather than re-derived: compare_clin_sig compares an authored call for a genotype, and annotations.parquet keys on genotype for the same reason. A table keyed on the variant alone would collapse two authored calls that disagree with the archive differently.

Repairs rejected

  • Folding the conflict into the overlay as an evidence column. A conflict nobody has answered yet has no overlay row to live on, so unanswered conflicts — the entire point — would have nowhere to be.
  • Escalation, a strict matter, or auto-correction. Out of scope on the reporter's own scoping and ours. A conflict is a question and half the time the archive is the stale side, which is why the ClinVar cross-check does not escalate under strict and why the shipped warning says so in its text.

Charter check

P2 — derived from an injected comparison, no fetch in the compile path. P3 — a new optional derived table, additive. P5 — the concordance and authored axes are separated rather than overloaded, which is what keeps the vocabulary arity-free. P9 — half cost: machine-written, and under RM124 it is not human-overridable at all, since the answer goes in the overlay rather than into this file.


RM134 — PubMind as a literature-derived annotation authority, and a ClinVar concordance check

Severity low-medium for the snapshot, medium-high for the concordance machinery, which reaches RM130 · Owner enricher, with the concordance record shared with RM130 · Entry ROADMAP.md § RM134 · Assessment PUBMIND_ASSESSMENT.md, which stays the evidence document · Motivating case the PubMind paper (doi:10.1038/s41467-026-76834-4) and a maintainer direction 2026-08-28

Why it is here, given this round's own sort rule

The sort excludes an item waiting on a caller, and sections B and C have a precondition that looks exactly like one: the ANNOVAR-distributed table publishes no data terms and only CHOP can say what they are. Reconciled the same way RM83's closure was, by separating the question from what it actually gates.

The licensing answer governs what a module may do with the values, not whether the machinery exists. Unknown terms warn and never gate (@no-named-licence), and the compile gate keys on sources.csv and nothing else — so a module carrying PubMind values records None on every term and behaves exactly as it does for every other source whose terms are unestablished. Nothing about building the snapshot, running the check, or drafting locally turns on the answer. What does turn on it is publishing a module carrying those bytes, and that is RM27's undesigned redistribution axis rather than this item. The ask-CHOP action stays in the entry as the thing that would turn None into an answer.

pubmind publish is refused by design, on the PharmVar precedent (@gated-source-caches): a bulk file arriving under terms we cannot establish is not a file we may pass on. The command exists and refuses with its reason, because a missing command reads as an oversight somebody will helpfully add.

The assessment stands; eight things changed while deciding

The assessment was written before this round and re-read against the eleven decisions. Its structure — four sections, the snapshot / the cross-check / drafting / the hint — is taken as written. Eight corrections, of which two were defects that would have shipped:

  1. The concordance record is shared with RM130, and RM130's shape changes because of it (below). ClinSigConflict names its authority in a field (clinical.py:74: clinvar: str), so a second authority would have forced a key change or a retype — major-only under P3, one item after shipping.
  2. One normalizer, not two, and it needs two fixes first — measured, not inferred. The assessment proposes a second hand-written map citing _CLIN_SIG_MAP as the precedent. Two maps for one vocabulary drift, and the check's entire output is a comparison of two normalizations, so a drift makes it report discordant on its own mapping rather than on the authorities. But reusing _normalize_clin_sig unchanged is also wrong today: its map keys are underscored and PubMind's tokens are spaced, and two of PubMind's six tokens fall through to other —
PubMind token _normalize_clin_sig today correct
Uncertain significance other uncertain_significance
Conflicting other conflicting

Both are concepts ClinVar normalizes correctly (Uncertain_significance, Conflicting_classifications_of_pathogenicity), so the check would manufacture a disagreement between two sources that agree — on PubMind's largest disagreeing class, since the corpus join reports Uncertain as 32 of 51 disagreements. The fix is a whitespace→underscore pre-step in the tokenizer, which is a no-op on ClinVar's already-underscored tokens, plus a bare conflicting key. _CLIN_SIG_SEVERITY's existing assert set(...) == VALID_CLIN_SIG guards the addition. 3. pubmind_sig/pubmind_sig_raw become clin_sig/clin_sig_raw. The ClinVar snapshot already uses the unprefixed names (clinvar_build.py:74): the source is the file, so naming a column after it restates the filename. With one column vocabulary across every snapshot, an N-authority check needs no per-source mapping at all — which is what makes the stress test below come out clean. 4. derivation's P5 open question closes. P3 makes a published name permanent because outside consumers key on it. pubmind publish is refused, so that parquet never leaves the machine that built it — not in an artifact, not attested in artifact.files, not a manifest field, not joinable by any consumer. Renaming it costs a rebuild the same command performs. The audit belongs on the names that really are one-way doors: --min-confidence, and the DRAFT_PROJECTIONS key pubmind, which is matched against SourceRow.source and therefore written into sources.csv, an authored published file. 5. The three-way check subsumes the two-way rather than running beside it. With no snapshot PubMind's side reads unchecked, which is what the tri-state is for, and the degenerate case is precisely today's finding — so nothing is lost and no author sees one disagreement reported twice. The pinned warning phrases are preserved deliberately, since @warning-text-is-api. 6. RM71 lands before section C. Both touch draft-panel in this release: RM71 widens the worklist past if report.added: and adds --dry-run, and --source pubmind must inherit the fixed worklist rather than reproducing the once-only behaviour. DRAFT_PROJECTIONS["pubmind"].identity cannot copy ClinVar's tuple, since most PubMind rows carry no rsID. 7. Rows withheld at draft time are named at draft time. Build-time drop counts land in release.json, but a draft that silently omits the derivation=indel rows is @unreachable-not-absent — write no row, and say separately that you wrote none and why. 8. The new checks are inside RM131's audit, and they are what forced RM126's warnings axis. Both amended in their own sections above rather than recorded only here.

The concordance record, and the stress test that shaped it

The maintainer's framing, and it is better than the authority-keyed record first proposed: concordance and discordance drive a composite algebra, so the record is about the agreement state and not about one authority. Keying on the authority gives N rows per variant and leaves the agreement state with nowhere to live.

The record. Keyed (variant_key, genotype) — one row per contested subject, stable at any number of authorities. Each authority's own call sits in a paired detail table keyed (variant_key, genotype, authority), carrying that authority's normalized clin_sig, its raw token, and its own confidence in its own units. Two flat CSVs rather than a nested cell: the arity objection was to the agreement state having no home, and on this split the state lives on the parent row while the per-authority facts are ordinary rows. Confidence is not normalized across authorities — ClinVar's review_stars and PubMind's 0–3 are different instruments, and folding them into one number is three axes in one field, the same objection the assessment itself makes against a single pubmind_score.

Two orthogonal fields, and the orthogonality is what makes them arity-free.

Field Members
authority_concordance concordant / discordant / single / none / unchecked
authored_position matches_all / matches_some / matches_none / absent / unchecked

The stress test, run at N=3 and N=5 on the maintainer's instruction, and the draft failed it. The assessment's seven outcomes name the authority inside the vocabulary member — pubmind_only, clinvar_only — so a third source needs a third member and five sources need every subset. The root cause is that one field carried two axes: concordant was defined as both authorities agree and the authored row agrees with them, with authored_dissents as a sibling member. That is the P5 anti-pattern, and the combinatorial explosion is its symptom. Split, both vocabularies are five members at N=2 and five at N=5, and which authority spoke is data in the detail table where it belongs.

The case that decided authored_position's definition. At five sources with a declared order E>B>D>C>A, suppose E and A agree and B, C and D agree against them. Lexicographic resolution says E; majority says B/C/D; choosing between those rules is a judgement about how authority rank trades against agreement count, and it needs a weighting model. This repo has refused to invent one three times — should_rebuild in RM126 (the same fact costs consumers differently, so the decision is theirs and only the fact is ours), @clinsig-never-escalates, and RM16's PRS weights, parked because combining weights needs a model nobody has; @weight-has-no-unit is the same rule again. So nothing resolves a split. authored_position is therefore a relation to the set rather than to a resolved call, which is computable with no weights and true at any topology: the E+A case reads discordant + matches_some, and a consumer holding its own weighting model computes whatever it likes from the detail rows.

The declared authority precedence is recorded and computed with by nothing. An ordered list beside weighting:/authorship: in module_spec.yaml, stating whose call the author weighted while curating — machine-readable so a consumer can see the stance, never an input our tier resolves with. We record the stance, report the facts, and resolve nothing. One optional module_spec.yaml field, full cost under P9, priced here rather than allowed to ride in unpriced: it is authored, a human writes it, and it earns that price by making a methodological stance readable to a consumer instead of leaving it inferable only by reading N contested rows. The per-variant exception is an overrides.csv row, so RM124 carries it and no parallel mechanism appears.

Because nothing resolves, the retype risk raised against the precedence dissolves with it: a field nothing computes with has no arity to outgrow.

Severity, and what the check must never become

Warning-tier in both modes, never escalating under strict (@clinsig-never-escalates), with more force here than for ClinVar alone: a disagreement with an LLM's aggregate over the literature is a statement about the extraction's limits at least as often as about the module, and 62 % agreement is not a number a gate is built on. discordant is not a defect in the module — it is a fact about the field — and reporting it as one would be wrong. opposed versus merely different reuses _CLIN_SIG_CAMP rather than inventing a second distinction, and it generalizes: at N authorities, opposed means two of them sit in opposite camps.

Absence is not disagreement and is not one thing: a variant missing from PubMind means no paper in the corpus was kept by the triage stage, not that the literature is silent and certainly not that the variant is benign. --offline with no snapshot is the third state (@unreachable-not-absent), and a check that could not run reports no zero (@tautology-zero).

Repairs rejected

  • An authored pubmind_score column. Full price on the authored layer, and P5 rules on it before the price matters: one number covering how many papers spoke, how confidently, and in which direction is three axes in one field, made permanent by P3. The snapshot keeps them apart.
  • Anything PubMind produces entering resolution.csv. Its coordinates are PyEnsembl back-mappings of extracted text, codon-enumerated and multi-PVID. "Authority" here means an authoritative annotation source as ClinVar is one, never resolution.csv's authority column (@source-vs-authority). PubMind annotates loci something else resolved.
  • Collapsing multi-PVID coordinates to one winner. 68,744 coordinates carry several, worst case 35, and their verdicts disagree. Choosing a winner means an ordering nobody defined — mode() over an unsorted group, which the deterministic-ordering rule bans outright. The multiplicity is a finding.
  • A majority or consensus field on the record. The E+A case above is the argument: it is a judgement needing weights, and precomputing it is should_rebuild wearing a different name. The detail rows carry everything a consumer needs to compute their own.
  • A separate draft-pubmind command. It writes the same tables from the same gene argument, so it would duplicate the genotype worklist, the placeholder guard, the dedup pass and the refusal summary — the parts that are hard to get right. draft-clinpgx is separate because it writes different tables.
  • Running the PubMind pipeline ourselves. Barred rather than descoped: the Constitution's non-goals forbid LLM SDKs in any tier, and AI-assisted authoring is just-dna-pipelines' job by name.

Charter check

P1 — the snapshot is data and the check is a comparison; no predicate language, no code in a cell. P2 — all of it in the enricher, the only tier permitted to fetch, and the compile path imports none of it. P3/P8 — a new derived table, a new optional module_spec.yaml field and new commands: additive, nothing demoted, nothing retyped, no published module invalidated. P5 — the two-axis split is the principle applied directly, and it is what survives five authorities. P6 — both vocabularies are frozenset[str] plus a validator. P7 — nothing changes what compile → reverse → compile carries; a drafted row is an ordinary authored row and the snapshot is not part of the artifact. P9 — the snapshot is the free layer, the concordance tables are half, and the one authored field is priced above.

The non-goal settles the competition question rather than an argument doing it. This workspace does no gene–disease inference; PubMind is an inference engine over the literature. The thing it does best is the thing we are constitutionally not in, which is what makes it a good source and a poor competitor.


RM128 — enrich() persists nothing until its tail, so a run killed at minute 29 has written nothing

Severity medium · Owner enricher · Entry ROADMAP_HISTORY.md § RM128 (shipped; it left ROADMAP.md when it landed) · Motivating case S66 (just-module-creator)

The problem, and the ask that dissolved

The truncation half closed 2026-08-24: nine sidecar writers now go through layout.atomic_writer, so a killed process leaves the previous table rather than a short one. Three asks remained, and the central one was incremental or checkpointed persistence, which turned on a promise nobody had written down: is a strict refusal allowed to leave rows behind? Today enrich(mode="strict") raises before the single if write: block, so a refused run leaves the module exactly as it was — a property somebody may be relying on, unwritten, which is exactly the state in which it gets broken by accident.

The question dissolves, because checkpointing was the wrong lever. The choice as filed was between keeping the promise and recovering the thirty minutes. It is not a choice: the run becomes a transaction, which keeps the promise absolutely and recovers the work as well.

The decision

Durable staging beside the target, plus an atomic commit at the gate.

  • The run stages resolved rows to disk, in the target's own directory, as it goes — the editor .swp pattern, and layout.atomic_writer already stages exactly there, so this extends a shipped primitive from one file to a whole run rather than inventing one.
  • Same-directory staging is not a convenience, it is the correctness condition: a rename within one filesystem is atomic, while shutil.move across a partition degrades to copy-then-delete and is not. Staging beside the target makes a cross-device move structurally impossible rather than merely avoided.
  • The gate commits. A strict refusal commits nothing, so "a refused strict run changes nothing" becomes a written promise rather than an accident of statement order — the item's actual question, answered in the direction that breaks nothing.
  • A kill at minute 29 leaves the staged work, and the next run resumes from it. That is the incident, and it is fixed without the mode-conditional behaviour that was refused in advance: write means one thing in every mode, and staging is not writing.
  • A flag keeps the staging files after a successful commit, for debugging. The default removes them.

Not mode-conditional, and the refusal in the entry stands: write=True meaning "at the end" under strict and "as we go" under best_effort is a flag that does not mean the same thing in every function that takes one. Under a transaction the flag means the same thing everywhere, because committing is the only write.

RM124 thins what the promise has to protect. What a staged, uncommitted table can contain is now provably machine-derived and never an authored value, since authored corrections live in the overlay. The two decisions were reached independently and reinforce each other, which is worth recording.

Ask 3 — an advisory lock, taken

flock on the spec directory for the read-modify-write window. The transaction does not close the concurrency window: two runs can each stage and each commit, and the last writer wins over a merge with neither knowing. The reported incident is the sharp form — a client-side kill did not stop the worker, a zombie run reached the write and overwrote a restored 330-row table with 162 rows, and the module then validated, closed and compiled green. Nothing downstream could see it, because the three branches that deliberately write no row for an unanswerable subject make a shorter table indistinguishable from a module whose author resolved less. Those branches are correct and are not in scope.

flock is the shape because it has neither of the alternative's problems: a lockfile left by exactly the kill this item is about would block every subsequent run, which is a worse unattended failure than the one it prevents, and the staleness rule that would fix it is a clock — which this repo has refused before (guard the plan, not the clock). flock dies with the process. It is untested here on the network filesystems a consumer may use, and the implementation owes a documented degradation rather than a silent one.

Ask 4 — a progress callback, taken, and the unit is decided here

progress: Callable[[int, int], None] | None = None, reporting (done, total) over subjects.

The entry filed this rather than shipping it because the resolver chain is batched inside resolver.py rather than being a per-subject loop, so the unit reported is a design choice — and a leaf shipped against a guess is one P3 keeps working forever. The guess is therefore argued rather than made:

  • The incident is an idle timeout. Both reported runs died at 1800 s with essentially every variant resolved. What the caller needs first is a keepalive with monotonic progress, which rules out phases: a 29-minute phase emits nothing and the timeout fires anyway.
  • total must be known up front for the number to mean anything to a caller rendering it. The subject count is; the link count is not, since it depends on what resolution finds.
  • Subjects are the only unit the author's mental model already has. Links are an implementation detail of the batched resolver, and publishing one makes a refactor of resolver.py a contract change — the rename P3 forbids arriving through the back door.

The reporter explicitly did not ask for a protocol, and none is added: two integers, no object, no event vocabulary to keep working forever.

Charter check

P2 — enricher-only; the compile path imports none of it. P3 — a keyword argument with a default and a staging directory are additive; no schema, no manifest field, no vocabulary. P7 — a committed run produces the table an uninterrupted run produces, which the resume path owes a test.


RM131 — warnings is a flat list[str], and the discriminator that would make it readable is discarded

✅ BUILT — shipped in 0.7, both halves in one release as this section sequenced them. Sixty-eight codes, nine carried, and three things this section did not anticipate. The emission surface is larger than "roughly 29 append sites and 16 returning helpers": the count misses the findings/messages collectors, the two .extend sites reaching into the schema tier, and the deprecated DuckDB resolver in the enricher, whose warnings land in manifest.compilation.warnings like everything else — so the audit reached three tiers rather than one. The carried list beside warnings is right, and the str subclass that makes it cheap leaks the code at exactly two places — a pydantic field flattens it, and any reformat returns plain prose — so compile_module/close_module read a classified list beside the result and three prefixing sites go through findings.restate; both are pinned by tests rather than left to a comment. And the suppression record has to be counted over the overlay, never over the rows it removed, or it says a number on lap 1 and vanishes on lap 2, which is the published-field disagreement the overlay's own no-op rule exists to prevent.

Severity medium · Owner compiler · Entry ROADMAP_HISTORY.md § RM131 · Motivating case S68 (just-module-creator)

The problem

ValidationResult, ClosureResult and CompilationResult all carry warnings: list[str] with no code, no count, and no way to tell a finding an author can clear from one they cannot. A compile of a 190-row module returned roughly 14 kB of warnings. strict=false does not help — it changes what counts as an error, not how much prose the channel carries — and every document on both sides of this seam tells an author that warnings on a green run are the real output, which is followable only if the output can be read.

We already compute the answer and spend it on severity alone. _BLAME_TIER/_BLAME_ROW is literally whose limit this is, and its own comment says "blame decides severity and nothing else"; _closure_warning reaches the same distinction from the other end. The actionability of each finding exists at the point it is built and is dropped on the way out.

Facts established while deciding it

Two, both of which shrink the item:

  • manifest.compilation.warnings already exists and already ships. The channel is not missing; it is unreadable. This is a structure change to a published surface rather than a new surface.
  • artifact_digest is computed over the parquet FileEntry list, so manifest.json sits outside the digest and appending to it moves no hash. A richer findings channel is therefore free of the identity cost that would otherwise price it. There is no separate metainfo artifact to build, and filing one would put a second home beside a shipped one.

The compiler emits zero logger.info calls, so there is no discarded informational tier either — the findings that exist are the warnings, and they are already persisted.

The decision

Both halves, sequenced, in 0.7.

  1. Actionability first, because it needs no vocabulary. A carried list beside warnings, holding the subset of findings the author cannot clear. warnings is unchanged and stays the complete list, so nothing that greps it breaks; a consumer subtracts to get the actionable set. This is the reporter's own fallback shape — a list beside, rather than a field on each — it invents no permanent names, and it answers the question they actually asked: can I do anything about this? blame and the closure branch already classify two families; what it needs is for every emission site to say which side it is on.
  2. Then codes, named by each check where it is built, feeding warnings_summary: dict[str, int]. Most work, most stable, and it has a precedent in this repo that is already a closed vocabulary a consumer keys on: VALID_VERIFICATION_CHECKS.

The audit is the shared groundwork and is done once. Roughly 29 append sites and 16 returning helpers must each declare their side and their code; doing it twice is the actual cost being avoided by sequencing rather than splitting across releases. RM134's new checks are part of that audit, not a follow-up to it — they are emission sites like any other, and classifying them after the fact is the second pass this sequencing exists to avoid.

The carried split is load-bearing beyond readability, which was not visible when this item was decided alone. RM126's sweep needs to tell a warning that moved because a message was reworded from one that moved because a finding appeared, and carried is that discriminator. So the classification has a second consumer, and getting it wrong costs more than an unreadable channel.

The suppression record rides this channel. A row suppressed by overrides.csv is invisible in the build product — absent, with no trace of why — which is the one thing RM124's suppress operation costs. It emits a finding into the same channel, aggregated by reason rather than one per row, so a module suppressing forty rows produces one line with a count. Same rule the tier already follows for repeated warnings, and it needs no new home precisely because the facts above say the home exists.

Repairs rejected

  • warnings_summary with a plausible code set, shipped unattended. The container is free and the code is not: a published vocabulary is permanent within a major under P3 and P6, so the first set shipped is the one every consumer keys on forever. That is the whole of the original deferral, and it is answered by deriving the set deliberately, not by declining the field.
  • Codes derived from the pinned catalogue. Honest, and covers exactly what consumers already match on — but partial by construction, and a summary that silently omits unpinned findings is worse than none, because a consumer reading a digest believes it complete.
  • Codes derived from the emission site. Mechanical and complete, and it keys on where the code lives rather than on what the finding is, so a refactor renames a published key.
  • A cap, a truncation, or a verbosity flag. All three hide findings rather than organising them, and the author with the most warnings is the one who most needs the hidden ones.
  • A new metainfo artifact. The channel ships already and is outside the digest; a second home would be one more thing to keep in step with the first.

Charter check

P3 — carried and warnings_summary are additive fields; warnings keeps its meaning and its text, and @warning-text-is-api is not disturbed. P5 — actionability and severity are separated rather than overloaded onto blame, which is the principle applied to a field that was quietly carrying two axes. P6 — the code vocabulary is frozenset[str] plus a validator. P9 — nothing authored is touched.


RM132 — pharm_variants.csv makes a clinical claim per row and cites per variant

✅ BUILT — shipped in 0.7 as decided, provenance_quote included: it does not follow, and both cross-check sites landed in the same release. One thing this section left as an implementation detail turned out to be the durable part: rather than teaching the two sites about a third table, the citing kinds became a derived registry (_CITING_TABLE_KINDS is every _TABLE_KINDS model declaring a pmid) behind two new public symbols, load_citing_rows/table_citations. So the RM40/RM41 requirement is met structurally rather than by discipline — a fourth citation site is read by both tiers by declaring the column — and an ast walk over the enricher's own source pins that no roster is kept there. load_binning_rows/binning_citations stay and stay narrow.

Severity medium · Owner format (schema) + compiler + enricher · Entry ROADMAP_HISTORY.md § RM132 · Motivating case S73 (just-module-creator)

The problem

A ClinPGx-drafted module carried 1,482 drug-response rows and had nowhere to cite any of them. Sixteen model fields, thirteen authored, none a PMID or DOI.

The shape was settled 2026-08-24 by the RM47 precedent, one release old and reached for a structurally identical table: a row cites when its claim is finer-grained than studies.csv's key. studies.csv keys on (variant_key, pmid), so a study row attaches to a variant; pharm_variants.csv keys on (variant_key, drug, genotype, phenotype_category, annotation_id), so one study row would attach to every drug, genotype and phenotype category recorded for that variant. evidence_level is not the provenance handle — it points at somebody else's grading of evidence rather than at the evidence — and the licence row's source/dataset state redistribution terms, not grounding.

Why a full-cost authored column is taken rather than deferred

This is the item that makes the round's sort rule load-bearing, so the argument is recorded rather than assumed.

What P9 prices is not the byte. An authored column is full cost because a human must learn it and P3 keeps it working forever, so the risk being priced is getting the shape wrong. That risk was spent a release ago. PharmVariantRow.pmid is a copy of two shipped fields under the same grammar (StudyRow.pmid, MeasureBinRow.pmid); an author who has met either learns nothing new. Demand is what fixes an unfixed shape, and there is no shape here left for demand to fix.

It is closer to a half-defect than to a new capability. The table already makes a clinical claim per genotype and structurally cannot ground one. That is a hole in an existing concern, not a new concern added to a table — the distinction the one concern per table gate actually turns on.

And the second half gets dearer with delay. Both literature cross-check sites must learn the new citation site in the same release — _cross_check_literature in the compiler and enrich_literature in the enricher — and that obligation is RM47's recorded lesson in its own words: shipping the column without both would be evidence the format never checks, which is worse than the gap. Every citation from the new site would otherwise read as a stale orphan in one direction and be invisible in the other. The cheapest moment to discharge a recorded lesson is while someone has just read it.

The decision

  • PharmVariantRow.pmid, optional, free-form under the same grammar as StudyRow.pmid and MeasureBinRow.pmid.
  • Both cross-check sites learn the site in the same release. The enricher reaches the rows through public compiler symbols, as load_binning_rows/binning_citations already do for the bins; a second copy of the table roster in the enricher is the RM40/RM41 shape and is not repeated.
  • provenance_quote does not follow, and the entry says so rather than leaving it implied. On the binning side it deliberately did not: the bin row cites and studies.csv/literature.csv describe, which is what stops StudyRow's whole provenance column set migrating one column at a time. A 1,482-row body of clinical claims is exactly where somebody will ask next, which is the reason to state the line rather than the reason to cross it.

Repairs rejected

  • Widening studies.csv's key. The repair that looks obvious and the one RM47 already refused: it would make a study row's subject depend on which table read it.
  • Treating evidence_level as the provenance handle. It grades evidence rather than pointing at it.

Charter check

P3 — a new optional column, additive; no published module is invalidated. P5 — citation and grading are separate axes and stay on separate columns. P8 — optional with respect to every published module. P9 — full cost, taken with the argument above rather than by weighing file count.


RM70 — requires_callable is VariantRow-only, so no PGx table can state CPIC's core assumption

✅ BUILT — shipped in 0.7 as decided, with no compiler change; reference_examples/cyp2c9_warfarin_grch37 populates the column on both tables. One thing this section did not anticipate: the two locus tables can name one place and legitimately disagree, because a haplotype row's claim is about assigning the reference haplotype and a pharm row's is about matching that row's genotype. So no cross-table equality check was added, and a test compiles the disagreeing pair clean to keep one from being added.

Severity medium · Owner format (schema) · Entry ROADMAP_HISTORY.md § RM70 · Found by dogfooding 2026-08-13, reference_examples/cyp2c9_warfarin_grch37/

The problem

CPIC's star-allele system assumes a position not called is reference — literally requires_callable=false — and haplotypes.csv, pharm_variants.csv and diplotypes.csv carry no such column; requires_callable and its companion callable_from are on VariantRow alone. So a star-allele module cannot record the single assumption a consumer most needs before trusting a *1/*1 result, and the check that exists for it is unreachable from the module kind whose upstream states the assumption in prose.

Not gated on RM65/RM66's caller VCF. Those wait on what a caller emits, because the shape of the data decides the schema. This is about what a curator asserts, and the assertion already exists in prose. The adjacency is that both ask whether a non-variants.csv table should carry something variants.csv has, not that they share a blocker.

The decision

requires_callable on HaplotypeRow and PharmVariantRow, optional. Not on DiplotypeRow. callable_from does not travel with them.

Two tables and not three because those two name loci — they are two of RM43's three positional tables — so a callability claim on either is about a position the row states, which is exactly what the column means on VariantRow. diplotypes.csv names a star-allele pair, not a locus, so the same column there could only mean "the variants defining these two haplotypes were callable", which is a fact about haplotypes.csv's rows restated one table over, where it drifts the moment a definition is edited. One concept, one home.

callable_from waits because the two are different axes and the cheaper answer is to add it when a module needs to say where the proof lives: callable_from says where the proof is, requires_callable says a proof is required, and a row may legitimately require one and not know where the evidence sits.

Repairs rejected

  • The column on all three PGx tables. Full cost three times and wrong on the third, per above.
  • Declaring it once in module_spec.yaml. The verdict is per locus and this repo has twice paid for assuming otherwise — RM36 rejected per-CSV build declaration because two files could disagree about one fact, RM32 rejected a gene-scoped PAR verdict because XG and SPRY3 straddle a boundary. CPIC's own assumption is not uniform: a gene whose common alleles are single SNPs and one defined partly by a structural event do not have the same callability requirement, and CYP2D6 has both inside one gene.
  • Deriving it from callable_from. Starts by adding the more expensive of the two columns, and collapses two questions into one: deriving requiredness from the presence of a pointer is an axis overload.
  • A stamped, compiler-managed parquet column. Nearly free under P9 and it cannot work — this is a curator's claim about what the annotation assumes, and a stamped column carries only what the compiler derives. There is nothing to compute.
  • Authoring the defining positions a second time in variants.csv. Two tables then name one locus, the shadow rows move artifact.digest while asserting nothing new, and it re-opens a star allele can be used without being defined from the other end, with two definitions instead of none.

Charter check

P3/P8 — new optional columns, additive, optional with respect to every published module. P5 — the home question is the principle, and the answer is the table that names the locus the claim is about. P9 — full cost on two authored tables, and the scoping decision is what keeps it from being three.


RM71 — the alleles a drafted genotype stub must be written from are in no file

Severity medium · Owner enricher (clinvar_draft) + compiler (draft) · Entry ROADMAP_HISTORY.md § RM71 · Found by dogfooding 2026-08-13, reference_examples/hboc_palb2/

The problem

draft-panel leaves genotype as vocab.TEMPLATE_PLACEHOLDER, correctly — ClinVar publishes alleles, not genotypes. The alleles the author must write the genotype from are in no file: a drafted row is rsID-only, so rs118203998 arrives with empty ref/alts, and the pair is stated once, in the warning stream. At the 16 rows PALB2 yields at ClinVar's 3-star floor this is a transcription exercise; at the 761 the same command drafts at the 2-star floor it is not one.

And it is emitted exactly once. The worklist is built inside if report.added: and scoped to added_records — itself a correct earlier repair, since it used to name rows the model had refused and rows already in the file. The consequence is that re-running draft-panel after the first draft adds nothing and therefore prints no worklist at all.

Corrected 2026-08-28, during implementation: draft-panel does have --dry-run. This paragraph said it did not, and the decision below was written to add one. It shipped in 0.5.1 (e40e118, 2026-08-03) with help text identical to draft's, threaded through both appends and gating the licence write, and the /create-module skill's command table already listed it. The claim was never checked against --help — the failure mode the command tables are known for, arriving this time in a proposal rather than in a table. What was actually missing is a test: nothing pinned the promise, which is why the flag could be believed absent by a reader of the code. The information cannot be re-requested from the command that produced it.

The decision

Make it re-requestable from the command that produced it. Two changes, neither of which touches a schema:

  • The worklist covers every stubbed row in the file, not only the rows this run added. The if report.added: guard and the added_records scope are what make a second run silent; widening the worklist to the file's current placeholder rows keeps the earlier repair's benefit — refused rows and settled rows stay out — while removing the once-only property.
  • ~~draft-panel gains the --dry-run that draft already has~~ — already shipped in 0.5.1; see the correction above. What lands instead is the test that pins it, so the worklist is obtainable without appending anything. A flag that means the same thing in both commands, which is the standing rule.

This answers the entry's actual open question — where does an author do this work — with in the command they already ran, rather than by inventing a third place. It changes no schema, fills no cell, moves no digest, and re-answers a question the drafting run answered once.

Repairs rejected

Every column-shaped repair, and the entry's own surviving candidate.

  • Writing ref/alts into the drafted row. A provider fills identity whole or not at all, and the model forbids ref/alts without a coordinate — so this means writing the full coordinate, discarding the rsID identity the provider deliberately chose as stabler and more legible. alts is also REDUNDANCY_BEARING: the allele-membership check compares the author's genotype against it and keeps its force because the two were authored independently. Filling it makes the compiler compare ClinVar with ClinVar.
  • A comment column on VariantRow. Full cost on the most expensive table, carrying text dead the moment the stub is replaced, and a provenance claim with no machine reader: "ClinVar publishes G>T" is a statement about a snapshot release, and a re-draft from a newer one leaves it naming the old alleles.
  • A sidecar the author reads beside the CSV. Half cost and still wrong three ways: its only reader is a human, its join key either says nothing a hint variant call does not or contains the placeholder, and an unknown file in a spec directory is tolerated but not read, hashed or listed in artifact.files — so a worklist file becomes a permanent resident with one more name for _check_misspelled_tables to learn.
  • Having the enricher fill it after resolution. Both cells are REDUNDANCY_BEARING, and filling a cell a Class-2 check cross-examines makes the comparison vacuous. The dependency also runs the other way: enrich refuses to load a file containing a placeholder, correctly.
  • Making genotype optional. Barred by P8, and wrong at 1.0 too: the zygosity decision is what the stub protects.
  • A bulk read-only advisory command. The entry's surviving candidate, and it puts the worklist in a third place while the author's complaint is that it is not in the one they are editing. The chosen repair is strictly smaller and lands the information where they already look.

Charter check

P3 — a new flag and a widened report; no schema, no vocabulary, no manifest field. P7 — untouched, since nothing authored changes.


RM85 — the origin of a module predicts the shape of its second pass, and nothing records it

✅ BUILT — shipped in 0.7 as currency.check_dataset_currency, a sixth enrich() check attested as dataset_currency, switched by --verify-datasets. Three things this section left open were settled in the build. "The source's current release" is not uniformly askable: a dataset label is minted by whichever pass wrote the row, and only ClinVar's has a live counterpart readable in the same namespace, so one probe ships in a derived registry and every other source is honestly unsupported rather than quietly current. Comparability is a withhold of its own beside the tri-state named here — clinvar_dataset_label's digest form and its stated-date form name one release space in two spellings, so a digest against a date is uncomparable, not behind. And strict refuses over the superseded set alone: severity follows the mode as decided, but an unreachable source and an --offline run both leave every leg unchecked, and escalating those would make --offline --strict impossible forever over something no author can edit.

Severity low-medium · Owner enricher · Entry ROADMAP_HISTORY.md § RM85

The problem

A module drafted from a source inherits that source's release cadence and needs a source-refresh pass; a module built from one paper inherits the literature's cadence and needs an evidence pass. Nothing tells an author their source has moved on. The tautology skip reads the release the module was drafted from and withdraw_stale_dataset handles a module that ends up mixing two, but neither answers "ClinVar has published since you drafted this". SourceRow.dataset records the release, so the fact is nearly there; what is missing is anything that acts on it.

The decision

An enricher check comparing SourceRow.dataset against the source's current release, reporting the gap. It needs the network, so it is an enricher check by the validation-ceiling rule, and it reads rather than writes — which is what keeps it out of the column-shaped repair the entry refuses.

It is RM83's neighbour: the same has the world moved question, asked about the release label instead of about the rows. With RM83's residue landing as a re-derivation that reports, the two compose — the label check is the cheap one that runs offline-adjacent and tells an author whether the expensive one is worth running.

Tri-state, and --offline is where it bites. Unreachable is not absent: a source whose current release could not be fetched is unchecked, never up to date. Severity follows the mode, as every enricher check does, and it reports rather than repairing.

Repairs rejected

  • A column saying what this module was made from and what would age it. Refused one table over in RM71, on the same grounds: it restates what dataset already states and rots where dataset is maintained — exactly the staleness withdraw_stale_dataset had to be built for, on a column where nothing could notice.
  • A publish-time or catalog-side signal. Puts the notice where a reader is rather than where an author is, and is out of these packages' scope. Recorded as an ask rather than built.
  • Nothing, deliberately. The status quo, defensible only while a module has one author who remembers, and it fails exactly when a module outlives its curator — which is the case the whole lifecycle document was written for.

Charter check

P2 — the enricher is the only tier that may fetch, and the check lives there. P3 — a new check with a new record member; the vocabulary member is additive. P9 — nothing authored, nothing derived stored.


RM133 — a card subtitle has no amendable home

Severity low-medium · Owner format (+ registry, for their half) · Entry ROADMAP.md § RM133 · Motivating case S64 (just-module-creator)

The problem, and the half that is already closed

Measured by the reporter: editing module.description from 44 words to 11 moves no content_signature, no artifact.digest and no fact signature — and drops the closure, because manifest.inputs covers the raw bytes of module_spec.yaml. README.md, by a wide margin the longer prose, is outside inputs and freely amendable. The shortest fixable prose in the system was the one that could not be fixed.

The binding stays as it is and that question is closed — it is the reviewer's claim rather than the artifact's, and a partition drawn along content_signature's line would make a closure transferable across a rename. Not re-proposed here. What is left is where a bounded short_description lives so that it lands amendable.

The decision

Registry-owned, beside the module, following the standing authority-key precedent.

  • A short_description key joins the registry-owned family alongside normalize.IDENTITY_AUTHORITY_KEYS (namespace, owner, canonical_id) — as a separate frozenset, since ownership and presentation are different reasons for a key to be registry-owned, and stripped by the same function so there is one path rather than two.
  • strip_authority_keys learns it, so a registry-injected value passes our validator and never reaches the stored bytes. manifest.inputs still matches, verify_manifest still passes, and the closure stands. Nothing about the binding moves, which is what makes this the route that actually unblocks the item.
  • The calibration ships as a constant on our side rather than being guessed at downstream: ~120 characters, against a measured 71 (comfortable) and 467 (the case that prompted it). A field that exists to fit a fixed layout is specified by that layout; it refuses nothing anyone has written, and absent it everything behaves as today.

Repairs rejected

  • short_description on ModuleInfo. Under the answered binding question every field in module_spec.yaml is on the un-amendable side, so this reproduces the defect in a new place. The reporter's own objection, and it is correct.
  • Splitting the binding along content_signature's line. Closed above; the two hashes cannot share a partition because they answer opposite questions about the same fields.
  • A field-aware six-field binding. Recorded as the better form of the idea and still refused: it inherits the cost of hashing a parse of the yaml, and RM82 is the precedent that prices it — the last change to the binding turned on being a byte transform needing no loader, no parse and no schema knowledge.
  • Render-time truncation or folding. Hides prose an author chose to write and leaves the spec wrong.
  • Anything retroactive to the seven published modules. They met every requirement that existed.

Charter check

P3 — a recognised registry-owned key and a published constant; no authored field, no schema change, nothing invalidated. P4 — the binding is untouched, deliberately. P9 — zero cost on the authored layer, which is the point of choosing this route over the natural-looking one.


Addendum — RM140, decided 2026-08-31 after the round closed

Severity medium · Owner format + compiler · Entry ROADMAP_HISTORY.md § RM140 · Motivating case S75 (just-module-creator)

Not one of the twelve, and not a reopening. The round above closed on 2026-08-31 and this arrived the same day, against a 0.7.0 whose three pyproject.toml files were bumped and whose tag was not cut. A new optional column is what sizes a release, and the number was already decided, so there was nothing to schedule: it lands inside the batch. What made it worth writing here rather than only in ROADMAP_HISTORY is that its shape is this document's shape — a column whose obvious companion gate is refused, and a check that had to learn one thing to stay true.

The problem

A reproducibility benchmark, run by the reporter: two agents, byte-identical prompts, the same three DOIs, one module each. They overlapped on exactly one row — rs117385980 from PMID 41249831 — and disagreed, p_value 0.36 against 0.75, with the same effect_size of 1.42.

Neither was a misreading. The paper reports two analyses of one association: an allelic Fisher's exact test on the 2×2 allele table (OR 1.4, p 0.36) and a univariate logistic regression (OR 1.42, 95% CI 0.18–11.67, p 0.75), with five adjusted models after it. One run's row was internally consistent with the second. The other paired the second's effect_size with the first's p_value — one analysis's estimate beside another's p-value, on a row that asserts the two belong together.

Everything was green, and that is the item. validate_module(strict), compile_module(strict) and audit_module all passed, and quotes_found was satisfied: the provenance quote is verbatim and correct, because it grounds the significance verdict and contains no statistic at all. A quote cannot witness a number it does not contain, so quote verification is structurally blind to this class of error — which is worth stating plainly, since the attestation column is the surface a reader would expect to catch it.

StudyRow had study_design, which describes the study. Nothing described the analysis. So a correct row and a mispaired one were byte-indistinguishable, and no check could be written at all: the facts it would compare were recorded nowhere.

Facts established while deciding it

The reporter's second claim does not reproduce, and the correction matters. They read key.columns = (variant_key, pmid) as meaning a paper's several analyses can be represented by exactly one, the rest dropped by silent choice. Probed on a real spec: two rows sharing a variant and a PMID both reach studies.parquet, and duplicate_study_citation is a warning that does not escalate under strict. So the capability was already there and only the legibility was missing — which narrows the item to the column and rules the key change out on the merits as well as on P3.

But the warning then contradicts itself. _cross_validate_studies's own docstring calls a repeated key the same claim written twice. Once statistical_test exists, two rows naming two analyses are not that, and an author using the column exactly as intended would be warned for it.

The decision

One optional free-form column, and one narrowing of the check that would otherwise contradict it.

  • StudyRow.statistical_test — plain str | None, shaped like study_design: the test or model, and what it was adjusted for. No vocabulary marker and no RECOMMENDED_* set; the space is open and a recommended set is additive later if a corpus ever shows a shape. What the column has to do first is make two rows distinguishable.
  • duplicate_study_citation suppresses only on both stated and different. An absent statistical_test is unknown, and unknown against a stated value cannot establish that two rows describe separate work. Kleene, not a != b — the naive form suppresses on every absent cell and silently retires the check for every module written before the column existed. Neither stated, both the same, or one stated and one blank in either order: warns exactly as before, same code, same bytes.
  • _KEY_FIELDS stays (variant_key, pmid). The check restates the pair rather than reading the tuple, so the split is contained in one function and reaches nothing else.

Repairs rejected

  • A validator requiring the pair to come from one analysis. The reporter argued this against their own ask, and the argument is correct: nothing on either side of the boundary knows what test a number came from until the column exists, and adding the column and the gate in one step would make every existing published row retroactively incomplete. Column first, and possibly never a gate — what a gate needs is a second recorded fact per number, not a stricter reading of one.
  • Widening _KEY_FIELDS to carry the analysis. Scoped out by the reporter, and independently the wrong move: the tuple drives hints.key_fields and the key.columns an authoring surface publishes, and re-keying a shipped authored table changes what an identity key means — major-only under P3. @dedup-key-decides-rows is the other half: the key decides which columns may become several rows, and that is a drafting-wide consequence for a warning-shaped problem.
  • A RECOMMENDED_STATISTICAL_TESTS frozenset. A recommended set is a claim about what a corpus contains, and one module is not a corpus. Open now, additive whenever there is something to recommend.
  • A second warning code for the half-stated pair — one row naming an analysis, the other blank. A permanent key for what is a transient state of an author mid-adoption; the existing message plus the rule in the compiler reference covers it. File it if someone reports being stuck there.
  • Exercising the column in a reference example. Moves digests for no gain and manufactures RM139's one side only case at the next cut, since no previous release can compile a spec using a column it does not have. Test fixtures only.

Charter check

P3 — a new optional column: additive, minor-legal, inside the decided 0.7.0. P8 — optional with respect to every published module, pinned by a test asserting that two specs differing only in the presence of the column hash to the same content_signature. P5 — study_design and statistical_test are separate axes, and a future analysis_covariates sits beside this one rather than inside it. P7 — the round trip carries the value and the digest is a fixed point, watched failing on each of the two reverse touch points in turn. P9 — full cost, an authored column, and the answer is the one this round gave twice already: the burden is on the rare author, the rare author is the one asking, and an unset cell burdens nobody.

The suppression is a loosening of a warning, never of validity: nothing that compiled stops compiling, nothing silent starts warning, and the message and code are byte-identical for every case that still reports.


Addendum — RM152, decided 2026-08-31 after the round closed

Severity low-medium · Owner enricher · Entry ROADMAP_HISTORY.md § RM152 · Evidence CIVIC_SURVEY.md, which stays the measurement document · Motivating case S84 (just-module-creator), 2026-08-31

Not one of the twelve, and not a reopening of RM134. It arrived as a consumer measurement against the same uncut 0.7.0 that RM140 landed inside, and it is here for the reason RM140 is: its shape is this document's shape — a candidate refused on evidence, a second candidate refused on better evidence, and a third route nobody proposed that turned out to cost nothing. What makes it worth reading beside the twelve rather than only in ROADMAP_HISTORY is that it is the first item in this line where a measurement, not an argument, changed the release class — it was filed deliberately carrying none.

The problem

RM134 shipped an N-authority clin_sig concordance record whose two vocabularies were stress-tested to hold at five authorities — synthetically. No real third authority had ever been run through it. S84 proposed CIViC as that authority and RM152 refused it on arithmetic: of CIViC's 3,103 germline evidence items the ACMG five-tier covers 599, 594 of them uncertain, leaving four pathogenic, one likely-pathogenic and zero benign-class. The check's entire product is opposition, and an authority with five calls in one camp and none in the other cannot make discordant sayable.

RM152 then left one question open — does direction warrant the apparatus clin_sig has? — and named the probe that would answer it, noting the reporter had declined to claim it.

Facts established while deciding it

The probe was run. All of RM152's figures reproduced; four things it did not know did not survive.

  1. Its own contested number was wrong in both directions. Grouped by molecular profile the count is 1, and that grouping is the mistake — a profile may name several variants and a variant sit in several profiles, so a two-variant profile's refuting row propagates to both members while each carries supporting evidence on its own. At variant granularity, which is what identity uses, it is 3. Two of the original four are lone refutations, which contested does not describe. And genuine opposition — risk against protective — is 0, at every scope and under every status basis.
  2. Widening past the germline origin filter adds nothing. All 620 variants re-swept unfiltered: 2,811 items, 11 newly camp-bearing, all SUPPORTS on variants already carrying risk, 0 new contested.
  3. The assertions table is structurally incapable of the axis. AssertionSignificance is a different 16-member enum without PREDISPOSITION or PROTECTIVENESS, so filtering by them is a GraphQL type error, not an empty result. This is stronger than a count: no assertion can ever hold a direction call, however the database grows.
  4. The denominator every record quotes was never declared. Both connections default to status: NON_REJECTED; ACCEPTED is 4,904 of 11,518, and the dated bulk TSV is accepted-only at 4,903 rows. Two published surfaces of one source, 2.35× apart, neither saying so.

And one nobody was looking for: CIViC publishes no GRCh38 coordinates — 2,189 GRCh37, 2,433 with no build, two GRCh38. That reads as fatal for a GRCh38-only format and is not, because coordinates are not the only identity in the file.

The decision

Refuse the apparatus; build the source anyway, on the axis that survives.

  • No direction-axis concordance record. Opposition is 0, the three contested variants are claim-against-refutation and all three dissolve under ACCEPTED, and the one real case is intra-authority — which (variant_key, genotype, authority) cannot represent as a disagreement and fold_authority_records already collapses on the sibling axis. Decisively: nothing in the enricher fills direction at all (swept). clinvar_draft's fold targets state, the legacy axis, and @axes-passthrough bars crossing them; VALID_EFFECT_DIRECTIONS is a deliberately separate axis. A concordance record needs two authorities and this one has one.
  • civic build — a dated release reduced to one parquet plus release.json, byte-reproducible. The bulk TSVs, not the API, because only the download side has dated releases and a snapshot that cannot name its input cannot be reproduced; the accepted basis is recorded so a count from it is never silently compared with one from the API.
  • draft-panel --source civic, writing direction and never clin_sig.
  • Identity from what CIViC publishes — an rsID, or a GRCh38 accession inside ClinVar's HGVS — and never a liftover. This is RM48's rule applied rather than circumvented, and better than RM48 hoped: a published rs-number is an independent value ordinary resolution cross-examines, where a recovered one would have resolution verify Ensembl against Ensembl.
  • A refutation is kept and its direction withheld. "Does not support predisposition" removes a claim without establishing its opposite; None is never False.

Repairs rejected

  • CIViC as a clin_sig concordance authority. S84's own preference, and the one the arithmetic kills. Unchanged by this round.
  • A direction-axis concordance apparatus. The open question, answered no — above.
  • Closing the item as a non-issue. Refused before this round on the grounds that 1,458 germline rows sit on an axis nobody had asked CIViC about. That refusal is now vindicated rather than merely argued: those rows are what shipped.
  • draft_from_civic on clin_sig. Still refused; the silent somatic drop is now a counted drop. But half its stated reason did not survive — "it would write rows with an empty significance column" is true of clin_sig (812 NA) and false of direction (0 of 1,458). The rejection had been measured on the axis the report aimed at, not the one RM152 itself named as surviving.
  • Liftover. Reopened at the maintainer's instruction with new balance weights, and the weights moved against it: the fallback branch RM48 left open is 131 variants, most recoverable through a ClinGen CAID without lifting anything, while the hazard is unchanged. Filed as RM153.
  • Resolving CAIDs inside civic build. A build that fetched would forfeit the offline reproducibility that is the whole reason it reads a dated file. RM153 carries it as an enrich-time question instead.

Charter check

P1 — a snapshot is data and a drafted row is an ordinary authored row; no predicate language, no code in a cell. P2 — all of it in the enricher, the only tier permitted to fetch; the compile path imports none of it. P3/P8 — no schema change of any kind: no column, no table, no vocabulary member, nothing demoted or retyped, no published module invalidated. The entire adoption rides on direction and state, which have existed since 0.3, which is also why it could land inside an uncut release without sizing it. P5 — direction and clin_sig stay separate axes, which is not incidental here but the whole finding: the source is useless on one and substantial on the other. P6 — no vocabulary was added; the measurement is the argument for not adding one. P7 — a rebuild is byte-identical and a re-draft is a no-op, both pinned. P9 — the snapshot is the free layer and the drafter writes only authored columns that already exist, so the authored surface is priced at zero.

One thing this round did that the twelve did not: dogfooding the new provider found a defect in shared drafting code — append_partial_rows built its covered-set from partials[0].match_on while comparing each row against its own, so a mixed-arity batch re-added rows on every lap. Invisible on a first run, and it grows a file by the same rows each time thereafter. Turning the tool on the work just done is what found it, which is the standing rule and not a lucky catch.


Filed by this round

Two items the decisions above created. Neither is 0.7 work.

  • RM135 — the outranks succession → the 1.0 cleanup tracker. Retire ProvenanceItem.outranks in favour of a single mechanism in which the existence of an override in an authored table auto-beats the derived value. A removal, hence major-only under P3. 0.7 documents the succession and emits no deprecation warning, because the replacement does not yet cover outranks's own case and a warning nobody can clear is noise rather than notice. Numbered rather than left as a prose note, because an unnumbered filing is one RM_TOC cannot index and nobody can find. Written into ROADMAP.md's 1.0 cleanup tracker and indexed in RM_TOC in the same commit as this correction.
  • Nothing else. The suppressed-row record was considered as its own item and folded into RM131 instead: manifest.compilation.warnings already ships and sits outside the digest, so the home exists and a second one would be one more thing to keep in step.

RM134 was out of scope and then was not — superseded 2026-08-28. This paragraph originally recorded it as belonging to the next round, on the reasoning that it was filed after the eleven-item scope was fixed. The maintainer pulled it in the same day. It is decided in § RM134 below, and the note is kept rather than deleted because the boundary it drew was real for part of a day and the decision to cross it was deliberate.


Not taken, and the gate in each case

Nine items, each waiting on something that is not a decision this repo can make. Restated so the deferral is a reason rather than an omission.

Item Gate
RM122 — the measure lookup as a public function A caller. The signature's open choices — one row or one per trait, None or a three-state result — are exactly what a first consumer implementing the lookup would settle. A leaf shipped against a hypothesis is one P3 keeps working forever.
RM68 — drafting on a non-GRCh38 module A real author with a GRCh37 module saying which outcome they wanted. Refusal makes such a module undraftable; stripping drops the 36 CPIC defining variants that have a position and no rsID. The warning shipped in 0.6.
RM16 — authored PRS weights A real consumer combining authored weights into a score. Not derivable, so full cost, and a score's shape is exactly what a first case would dictate.
RM56 (policy) — a measurement spanning bins Real caller output, so the vocabulary is fixed against what callers emit rather than against a guess.
RM65 — coordinates on the positional repeat/CNV tables A real repeat-caller or CNV VCF. Same gate as RM56. Would take RM43's coordinate lane from three tables to five.
RM66 — several motifs at one repeat locus Same gate as RM65, and it is a keying change on a shipped table, the expensive kind.
RM23 — predictor scores as a table Two blockers unmoved: per-transcript grain, and the acquisition measurement. Licensing is not the blocker.
RM28 — meta-conclusions, predicate half A corpus. The injected-cofactor half dissolved on 2026-08-13; what is left is small and genuinely unsolved.
RM67 — polyploid / partially-phased genotypes Not work — a documented divergence, numbered so it is findable and not re-probed.

RM126 is the counter-example that makes the table honest: it is in this document rather than this table because it waits on nobody, which is what "owed rather than offered" means.


Implementation plan

Ordering

RM124 is the keystone and goes first. RM83's closure, RM130's sidecar and RM128's thinned promise all hang off it, and its enricher docstring sweep belongs in its own commit rather than being scheduled separately.

RM126 is fully independent and may run first or in parallel — it touches no file the other lanes touch, and its release gate is the only piece that lands outside the packages.

Lane Items Tier
A RM126 — record + needs_recompile + roster; the sweep; the release gate format, compiler, release sequence
B RM124 — the overlay: model, apply, reverse, merge-not-clobber drop, docstring sweep format, compiler, enricher
B1 RM83 — the closure bookkeeping and enrich --rederive enricher, docs
B2 RM130 — the conflict table enricher
C RM128 — the transaction, flock, progress enricher
D RM131 — actionability, then codes; the suppression finding compiler
E RM132 — PharmVariantRow.pmid + both cross-check sites format, compiler, enricher
F RM70 — two requires_callable columns format, compiler
G RM71 — the widened worklist and draft-panel --dry-run enricher, compiler
H RM85 — the dataset-currency check enricher
I RM133 — the registry-owned key and the constant format
J RM134 — pubmind build/publish, the shared normalizer fix, the N-authority check, --source pubmind, the hint enricher (+ format for the precedence field)
K RM140 — StudyRow.statistical_test + the analysis-aware dedup check (the addendum; not one of the twelve) format, compiler
L RM152 — civic build, draft-panel --source civic, CIVIC_TERMS (the second addendum; not one of the twelve) enricher

B before B1 and B2, by construction. C before B1, because enrich --rederive composes with the transaction rather than reimplementing staging. D before the suppression finding lands, which is inside D anyway.

J is the largest lane and it is sequenced by four dependencies, none of them optional. B2 before J, because RM130 and RM134 ship one record and RM130 is the smaller half of it — build the two-axis table with ClinVar as its only authority, then add the second. G before J's section C, so --source pubmind inherits RM71's fixed worklist instead of reproducing the once-only behaviour. J's new checks inside D's audit, not after it, since they are emission sites like any other. A before or with J, because J's checks are what make RM126's warnings axis load-bearing; landing them first means retrofitting the release's declaration.

L is independent of every other lane and was built last, on 2026-08-31 — it adds two enricher modules and one licence row, touches no schema and no compiler path, and the only shared file it edits is draft.py, where it fixed a defect its own dogfooding exposed. It is in this table for the reason K is: a reader asking what was built inside 0.7.0 should find all of it in one place.

K is independent of every other lane and was built after them, on 2026-08-31 — it touches StudyRow and one function in compiler.py that no other lane goes near, and it is sequenced only by having arrived last. It is retrofitted into this plan rather than left out of it because a reader asking what was built inside 0.7.0 should find all of it in one table.

Shared-file hazards

  • compiler/src/just_dna_compiler/compiler.py is touched by B (the overlay apply and the reverse half), D (roughly 29 warning append sites) and E (_cross_check_literature). Serialize these three; do not run them as parallel worktrees. Suggested order B → D → E, so the warning-site audit happens once the overlay's own findings exist and E's new citation site is classified when it is added rather than afterwards.
  • enricher/src/just_dna_enricher/enrich.py is touched by B (writers and docstrings), C (the transaction), H (the new check) and J (the concordance check). B → C → H → J.
  • enricher/src/just_dna_enricher/clinvar_build.py is touched by J alone, but the change is shared property: _normalize_clin_sig stops being ClinVar's and becomes the workspace's one significance normalizer. Move it where both builders call it rather than importing across snapshot modules, and pin the two-token fix with a test that runs both sources' raw tokens through it — the defect is invisible from either side alone.
  • enricher/src/just_dna_enricher/clinical.py is touched by B2 and J, on the same dataclass. One lane, sequentially: B2 takes the authority out of the field name, J adds the second authority as a row.
  • schema/src/just_dna_format/normalize.py is I alone.
  • manifest.py is touched by A (nothing — the record table is its own module) and D (the new result fields). No conflict.

Standing requirements for every lane

  • A test per decision, not per file. P7's round trip for the overlay (compile → reverse → compile, byte-identical, with the overlay applied twice proving idempotency); the resume path for RM128 (a committed run equals an uninterrupted run); the interval union for RM126 (moved-and-moved-back still reads as moved, and a version against itself is empty).
  • Registry-iterating guards assert an equality over a walked set, never a floor. The overlay's covered-table set, the warning codes, the authority-key family and the significance map are all registries.
  • The concordance vocabularies get an arity test, not only a coverage test. Construct a three-authority and a five-authority case and assert the member set is unchanged — that is the property the two-axis split exists to hold, and a coverage test at N=2 cannot see it.
  • Every new warning phrase is pinned, and the catalogue in COMPILER.md gains it in the same commit.
  • No count is computed and discarded, and no check that cannot fail reports a zero.
  • Tri-state everywhere: unknown for an uncovered interval, unchecked for an unreachable source, and None is never False.
  • Docs move with the code, in the same commit. SCHEMAS for the overlay and the succession, COMPILER for the warnings channel and the citation site, ENRICHER for the transaction and the checks, MODULE_LIFECYCLE § 6.3 for what deleting a sidecar now costs (nothing), and CHANGELOG under the 0.7 heading.
  • The /create-module skill gains the overlay and the draft-panel --dry-run flag; its command tables rot silently, so re-run --help against them.

Provenance

Decided across five interview rounds with the maintainer on 2026-08-27/28, on the 0.7 branch, against the full text of CONSTITUTION.md and the complete roadmap entries for all twelve items. The fifth round was RM134, pulled in after the other eleven were decided and reviewed against them rather than on its own terms — which is how the RM130 collision and the normalizer defect were found.

Six things changed during the interview rather than being carried in from the entries, and each is recorded above where it applies: RM83 dissolved into RM124 once derived tables became pure build products; RM128's central ask dissolved the same way, into a transaction that keeps the promise instead of trading it; the metainfo artifact was not filed, because manifest.compilation.warnings already ships outside the digest; and the outranks overlap became a dated succession rather than either a merge or an unstated duplication. Then RM134 added two more: RM130's record shape was superseded before it was built, because one release holding both items exposed an authority baked into a field name; and a maintainer-set stress test at five authorities failed the drafted vocabulary, which is what produced the two-axis split and, through it, the finding that nothing may resolve a split at all.