Skip to content

RM index, items 100–199

One page of the RM index, which is paged by number: every item numbered 100–199, under the status or round heading it had when the index was one file. The rules for adding an entry, and the allocator that claims a number, are on the index page.

✅ The 0.6.1 round — what reading the code against the documents found

Filed and shipped 2026-08-18. A pass regenerated SCHEMAS/COMPILER/ENRICHER from the source alone, read the result against the shipped documents, and asked which of the two was wrong; eight times it was the code. Every one broke a rule this repo had already written down, and in four the file carrying the violation also carried the rule — so the durable half of each is a test, and in six of the nine that test walks a registry rather than a list. Five of the filings understated what was there and the entries in ROADMAP_HISTORY say so. No schema surface moved; all sixteen reference examples recompile byte-identical. See ROADMAP_HISTORY § 0.6.1.

  • RM100 — ✅ shipped in just-dna-enricher 0.6.1 on 2026-08-18. Five few-line defects, no shared root: python -m just_dna_enricher.cli loses three commands to a misplaced __main__ block (26 vs 23, measured); clinvar_build._sha256_file is defined twice and the shadowing one returns str | None into two non-optional annotations; enrich_gwas closes its client outside any try/finally and never reads mode; NCBI_API_KEY is read without load_env() against @credential-where-read, so a .env key is honoured or not by call order; and net.py documents nine @retry policies where the tree has twelve, behind a test that asserts a floor (>= 9) while walking seven of the ten modules carrying one. · from the code-first doc pass · also in ENRICHER, CHANGELOG
  • RM101 — ✅ shipped in just-dna-enricher 0.6.2 on 2026-08-18 (a partial cut; format and compiler stay 0.6.1). @client-exception-contract one layer up from RM97: five passes let a client's type out through a try/finally with no except, so except FrequencyEnrichmentError was silent for the failure it was written for — and our own frequencies command printed nothing where it promises FREQUENCIES FAILED. Repaired with an *Unavailable subclass per pass (after AcmgListUnavailable), which also splits the ClinGenError/GeneValidityError conflation of "could not fetch" with "your local CSV will not parse". RM97's coverage guard walked eight hand-written module names and missed identifiers, leaving OntologyClient leaking raw httpx for a release — @registry-completeness; both guards now walk the package. · from S37 (just-dna-registry) · also in ENRICHER, AGENT_NOTES, CHANGELOG

✅ The RM10/RM11 session — a downstream MCP surface restated what it could not generate

Filed and shipped 2026-08-20, all four from just-module-creator in one sitting. Their MCP tools had three answers that restated a schema fact instead of generating it; two of the three had no public symbol to generate from, which is the report. Additive and documentation only — nothing removed, nothing retyped, no authored surface moved.

  • ✅ RM117 — the observability half shipped 2026-08-31 in the uncut 0.7.0; outranks itself shipped 2026-08-20 and the SEVERITY half was closed 2026-08-21 (a checked verdict under authored control is something nothing else here does). Recast onto the OVERLAY, not outranks — RM135 filed that field for removal at the major, and growing observability on it is what RM135 warns against. The finding was already firing with the wrong words on it, and that is the durable part: clin_sig_concordance.csv holds contested subjects only and is rewritten whole, so a subject leaving it means the disagreement ended — and an overlay row answering it then reached nothing and drew the generic the subject may be mistyped, put to an author whose judgement had just been confirmed. So the work was stopping a wrong signal, not adding one; hence a code (overlay_answer_vindicated, actionable) rather than a rewording, and a test asserting the misleading line is GONE. An observation, never a verdict — the wording test greps for adjudicating words on a word boundary and caught a real one: the first draft said the conflict ended rather than that the correction is wrong, which grades the author's row while claiming not to. Scoped to the one table whose absence has a single reading; everything else takes RM137's split. The second signal is RM151, which shipped the same day: it needs the archive's value NOW against its value at record time, so it needed a fetch and an enricher home. · from S52 (just-module-creator) · also in CHANGELOG, COMPILER, ROADMAP_HISTORY § RM151
  • RM120 — ✅ shipped in just-dna-format + just-dna-compiler on 2026-08-20. The reporter retracted their own S11: "nothing establishes a human ever looked" names a missing attributor, not an illegitimate reader — the machine does read the article, so the refusal protected a fiction about who, and the column stayed empty for the only reader present. RM118 is the evidence: the rule produced 3,668 title-quotes with the check green. Our model already disagreed — Defaults.curator is literally ai-module-creator and Contribution.who reads "a name, handle, or model id". Shipped StudyRow.curator, the field VariantRow always had, on the table where the attestation lives; free text resolvable against authorship, never a machine_located: bool, which collapses agent found it, human confirmed it into one of two lies. ATTESTATION_BEARING unchanged. Wiring it found that @three-touch-points undercounts for _write_studies_csv: a fourth site, the row dict, where a missing key makes DictWriter write the header with an empty cell — the reversed spec re-validates and loses the value, and only the digest fixed-point catches it. · from S55 (just-module-creator) · also in CHANGELOG
  • RM118 — ✅ shipped in just-dna-enricher on 2026-08-20. A title appears in its own fulltext, so a provenance_quote copied from article metadata makes quotes_found pass while establishing nothing. Reproduced on our tree: four published modules, 3,668 quotes, exactly one distinct quote per PMID, each verbatim the title — a located passage varies row to row, one string per paper is a property of the article. Worse than the failure S11 was written to prevent, since the check agrees. The discriminator must be the metadata, not the string's shape: length cannot separate a 17-word title from a 17-word sentence, and a regex is as copyable as a quote. Shipped titles_as_quotes + a yellow CLI line, warning-only; normalisation is case/whitespace/trailing-period and nothing more, and it fires only when every quote for a citation is the title. A failing test corrected the report's premise — the title match misses on the trailing period against real JATS, which is punctuation rather than evidence, and both states are pinned. · from S54 (just-module-creator) · also in CHANGELOG
  • RM119 — ✅ shipped in just-dna-compiler + just-dna-format on 2026-08-20. Two halves. Staleness: literature.csv recorded quotes_authored=0 beside 69 real quotes in studies.csv (65 of them on the pmid reading zero) across four published modules — merge-not-clobber keeps the row the pass wrote before the quotes existed, and nothing compared two files in one directory joined on pmid. Shipped _check_quote_counter_is_current, naming both numbers; reads both citation sites through binning_citations, since walking bin_rows reaches DiplotypeRow which has no pmid — caught by the suite before shipping. The manifest: _literature_block's per-row guard is right and does not survive aggregation — sum() over all-null rows is 0, so the block published exactly the sentence its own docstring calls the most misleading thing it could say. Shipped Literature.quotes_unchecked rather than nullable counters, because a reader needs three states and int | None collapses two of them. · from S56 (just-module-creator) · also in CHANGELOG
  • RM116 — ✅ shipped in just-dna-compiler on 2026-08-20. A version-comparator tool needs the rows content_signature hashes; it returned the digest alone, so the mapping had to be rebuilt against private _TABLE_KINDS and _resolve_spec_defaults. Both halves reproduced on hfe_hemochromatosis, matching the reporter's digests to the character: renaming sources.csv → licensing.csv and editing a notice cell in it both leave content_signature unmoved (the licence layer is hashed by source_signature — a documentation defect), while the obvious build outside the function disagrees with it on a defaults-vs-cells pair (a correctness trap, and the fold's third line is the one whose omission looks fine). Shipped their candidate spec_tables(spec_dir) -> (tables, build), with content_signature becoming that plus the hash and no logic moving; their rejected alternative — exporting the two pieces separately — rejected for their reason, that three pieces must be assembled in one order. COMPILER.md now names the roster, the exclusion and the fold. · from S53 (just-module-creator) · also in CHANGELOG
  • RM115 — ✅ shipped in all three packages on 2026-08-20. RM113's question asked of the machine-produced tables, where the key existed only as a dict-key expression in the writing pass's body. key_fields already routed derived names through derived_model_for; the seven models simply declared no key. The reporter's approximation — the required members of the fact-field tuple — agreed on five of eight and was coarse on exactly the two tables where one subject carries several rows, dropping disease_id from gene_validity.csv and variation_id from clinical_assertions.csv; both reproduced before the fix. Two shapes the obvious design could not express: resolution.csv's key is a subject and not a uniqueness constraint (one rsID → several loci, so KEY_RULES gains subject), and gene_validity.csv's has two levels (assertion_id else the grain, so TableKey gains fallback, tagged so a grain cannot collide with an id). Shipped base.merge_key + seven _KEY_FIELDS declarations, and every pass now keys its existing map through it. The rewire exposed three lookup sites restating the key positionally, one of which would have refetched every cited article on every run. · from S51 (just-module-creator) · also in CHANGELOG
  • RM113 — ✅ shipped in just-dna-compiler + just-dna-format on 2026-08-20. natural_key answers per row with values; _TABLE_DUPE_KEYS was private and lambdas, so no consumer could obtain a kind's key columns — and the reporter's hand-kept string had gone stale, telling authors to key copynumbers.csv on modifier_cn, deprecated at 0.6. describe_table's docstring had promised the key since 0.5 and the dict never carried it, so this was an unimplemented sentence rather than a request. The obvious fix is a wrong answer: filtering _KEY_FIELDS through model_fields silently drops CopyNumberRow's property-valued modifier axis, so a derived member is mapped back to its authored column in the preferred spelling instead. Shipped hints.key_fields → TableKey(columns, rule, stamped) with rule a frozenset vocabulary (equality/overlap — binning groups, it does not dedup) plus the key block in describe_table. Structurally, eight models now declare _KEY_FIELDS and both dupe-key dicts derive from it through one _key_of, so the two readers cannot drift; pinned against natural_key over real rows. The guards found an asymmetry the design missed — variant_key is a stamped field on three models and a property on StudyRow. · from S48 (just-module-creator) · also in CHANGELOG
  • RM114 — ✅ shipped in just-dna-compiler on 2026-08-20. The studies.csv -> variants.csv companion pull was unconditional, so scaffolding a binning module wrote an empty variants.csv stub into a spec that compiles strict-green without one — RM47 is what made it wrong, a study row being legal with no variant identity so bins can be grounded through pmid. The constant's own comment already said alone; the condition was simply never applied. Shipped scaffold.companions_for(kinds), public because the reporter passes COMPANION_KINDS through and an internal-only fix would leave their surface giving the old answer. variants.csv -> studies.csv stays unconditional; the recognised set derives from _TABLE_KIND_CSVS + variants.csv, mirroring the compiler's composition rule, with sources.csv outside it. Reporter's blunter option (drop the direction) rejected — it loses the help in the one case the pair exists for, now pinned. · from S49 (just-module-creator) · also in CHANGELOG
  • RM112 — ✅ shipped in just-dna-compiler on 2026-08-20. draft.model_for is authored-only by design, so a tool asked what is in frequencies.csv or resolution.csv got "not an authored table of this format" — true and useless — and the only complete map was private compiler._FACT_TABLES. All four substitutes were checked and fail: FACT_CSVS is names-only and in another package, ARTIFACT_PARQUETS - LEAD_PARQUETS measures nine names against seven fact tables, and authoring_reference()["models"] is keyed by model name so it cannot answer the filename direction a tool caller holds. Shipped as hints.DERIVED_TABLE_MODELS + hints.derived_model_for, derived from _FACT_TABLES with set equality over the walked set (@registry-completeness) — publishing a hand-kept copy of the map would have been the defect wearing a public name. Both licence-table spellings are keys; verification.json excluded as the attestation document. describe_table's refusal deliberately unchanged. · from S47 (just-module-creator) · also in CHANGELOG

✅ The S57–S60 session — a dossier audit's four reports, from the same consumer

Filed 2026-08-20 by just-module-creator after auditing their own per-table dossiers and their own attestations. Two of the four are ours to fix, one is a documentation gap with a real edge, and one is a 0.7-sized design ask that turns out to answer RM83's blocking sub-question.

  • RM121 — ✅ shipped in just-dna-compiler + just-dna-format on 2026-08-20. stats.genes came from variants.csv alone, so a star-allele or copy-number module published genes: [] however many rows named a gene, and a registry gene index fed from it could not reach the module. The model had already answered the reporter's either/or: Stats says "derived from the spec", so this was an unimplemented sentence, not a scoping choice. Reproduced on cyp2c19_star_alleles: 1,332 rows naming CYP2C19, genes: []; seven of our sixteen examples have no variants.csv, and seven of the eight non-variant gene-bearing models make gene required. Ours rather than theirs because the only workaround is an invented variants.csv — a fiction, since studies.csv becomes required with it. Shipped module_stats beside variant_stats (a rename is major, S14's rule), with _GENE_BEARING_TABLE_KINDS derived from _TABLE_KINDS; derived sidecars excluded structurally. The fix recreated the RM44 defect it inherits from — the post-drop re-derive sat inside the variants.csv branch and pharm_variants.csv also drops and also carries a gene — moved after the loop and pinned. A patch: manifest.json is outside artifact.digest and stats outside content_signature, measured. · from S57 (just-module-creator) · also in CHANGELOG, SCHEMAS
  • RM122 — ⏳ parked on demand, moved to the minor-deferral file on 2026-08-21 (it was filed under open, release undecided when nothing about it was undecided — an item waiting on a caller belongs in the deferral file). S58's own ask shipped as documentation (the normative lookup paragraph in SCHEMAS, plus the honest sentence that the family is specified ahead of its consumers, plus the same warning in the authoring skill). This is the question they did not ask: whether the rule should also be a public function, so the first two consumers to implement it cannot disagree on the four sharp cases — the continuous shared endpoint, float32 comparison, unresolved versus no-match, and pleiotropy returning several rows. alleles.split_genotype is the precedent and it costs the format tier nothing. Parked because there is no consumer to fix the signature against: one row or one per trait, None or a three-state result — a leaf shipped against a hypothesis is one P3 keeps working forever. The settling event is specific: one consumer implementing the lookup against the paragraph, whose questions are the signature. Verified while answering: just-dna-lite reads the four tables only to count rows. · from S58 (just-module-creator) · also in SCHEMAS
  • RM123 — ✅ shipped in just-dna-compiler + just-dna-enricher on 2026-08-20. The reporter's generalisation is the item: a check that could not have failed should record why rather than record a zero. Three instances filed, two reproduced. PGx: RM73's per-leg tautology skip has worked since 0.6.0, but _function_check_record's answered branch built detail from the answered legs alone, so the mixed case the per-leg design exists for — CPIC tautological, PharmVar answering — published a clean two-authority comparison with no sign half of it was a copy against itself. The note lived in result.warnings, which is the run's stderr, and verification.json is what a later reader trusts. Fixed, and sorted by source in both branches because that file is a hashed input and legs fills in pass order. Hints: REDUNDANCY_BEARING is keyed on a bare column, so the vacuous-check reason printed on six (column, table) pairs whose checker cannot see the table — right advice, false reason. REDUNDANCY_BEARING_TABLES narrows the explanation only; the provider refusal stays column-keyed, and the six unscoped columns are checked claims (RM43 puts resolution on the positional kinds, RM47 makes a bin a second citation site) that a checker-name-string reading would have wrongly suppressed. Their third instance did not reproduce: enrich_gene_metrics has separated asked, source has nothing (not_found row) from never asked (no row, unconsulted) since RM98 in 0.6.1, against their installed 0.6.4. · from S59 (just-module-creator) · also in CHANGELOG, COMPILER
  • RM124 — ✅ shipped in all three packages on 2026-08-28, the 0.7 keystone. overrides.csv, keyed (table, subject, member, field) — one member column whose meaning the named table fixes, never a per-table key grammar. Operations update/insert/suppress; the wildcard member is group-scoped for update and refused for the other two. The covered set is SEVEN, every merge-not-clobber derived sidecar but licensing.csv — the proposal's "six" was a miscount that enumerated nothing. reverse emits the post-overlay table plus the overlay, so it applies twice and the fixed point is tested rather than assumed; that is what buys the round trip with no previous_value column. No operation reports its own no-op, because after a reverse all three no-ops are true of a healthy module and a warning there would make a module disagree with its own round trip. Merge-not-clobber's behaviour is unchanged — a re-run still gap-fills — and its cost is gone: rm plus a re-run is now free. Question 2 settled as a dated succession, RM135. The original entry is kept in history/ROADMAP_0_7 as the record of what was observed. · from S60 (just-module-creator) · also in history/ROADMAP_0_7 § RM83, CONSTITUTION (the cost amendment)

✅ The 2026-08-19 doc-audit patch round — six of the eight, fixed

Shipped 2026-08-20. Six of the eight findings RM104–RM111 were sized as patches and are here; RM108 and RM110 stayed open above, and the 2026-08-21 decision round settled both — RM108 takes the newest classification_date as current and marks rather than deletes, RM110 canonicalizes to pipe-joined / None-when-empty on both legs. Both then shipped on 2026-08-31 in the uncut 0.7.0. RM110's read on this line was wrong in an instructive way: it was filed as needing a canonical encoding when a test already pinned one on the live producer, so what it actually needed was a release — normalizing moves gene_metrics.signature, and a signature move is not patch work. Building it turned up a second correction to the same entry: "empty → null" was only half of it, since the 708 genuinely flagged rows are literals too. Five of the six were a derived value restated by hand somewhere else, which is the pattern worth carrying out of the batch rather than eight unrelated bugs.

  • ✅ RM104 — ✅ shipped in just-dna-enricher on 2026-08-20. reference was bound inside if wanted: and read unconditionally below, so the idempotent re-run — the path merge-not-clobber documents as supported — and any module with no variants.csv raised UnboundLocalError, outside the GeneMetricsEnrichmentError contract RM101 built for exactly that caller. One line; the test is the half that matters, since every existing merge test re-ran with wanted non-empty. @empty-work-is-a-path · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES
  • ✅ RM107 — ✅ shipped in just-dna-compiler on 2026-08-20. A duplicate (source, layer) row compiled green under --strict — no warning, a moved source_signature, and a pair free to carry opposite commercial_use in the one file the compile gate reads. The fix as filed would have done nothing: _TABLE_DUPE_KEYS is consulted only by _validate_table_kind, which ran over _TABLE_KINDS while sources.csv is a _FACT_TABLES member, so the red test came first and proved the wiring. Call site widened, map entry added, SourceRow given back by draft._CORE_DUPE_KEYS. @which-loop-calls-the-checker · from the 2026-08-19 doc audit · also in CHANGELOG, COMPILER, AGENT_NOTES
  • ✅ RM109 — ✅ shipped in just-dna-enricher on 2026-08-20. The merge key is (gene, dataset) and "already done" asked source.startswith("gnomad"), so an honest source="manual" correction did not suppress the fetch and the file came back with two rows under one key contradicting each other, with no compiler check to catch it. done now asks whether a row sits under a key this pass would write; scoping to the two route labels is what keeps a second authority's ClinGen row from suppressing the fetch. @suppression-from-merge-key · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES
  • ✅ RM106 — ✅ shipped in just-dna-compiler on 2026-08-20. _check_frequency_arithmetic ran in validate_spec (RM93) and again in the compile-side _frequency_checks with no filter, so the faf95 line reached the published manifest.compilation.warnings twice — 15 warnings, 14 distinct. Filtered on the message, the idiom _literature_checks eleven lines below already used. Second instance after RM94, so the test pins len(warnings) == len(set(warnings)) over the whole list. @no-rerun-with-counts · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES
  • ✅ RM105 — ✅ shipped in just-dna-enricher on 2026-08-20. LOGO_EXTENSIONS admits jpeg and discovery sorts, so the spelling the compiler prefers was the one upload._ALLOW_PATTERNS dropped, and the manifest attested bytes the published repo did not carry. Derived from LOGO_EXTENSIONS now, with a set-equality test (a floor passed). _collect_logo's pick order deliberately untouched. The process half: the skew was named twice — a CHANGELOG line and a deferring comment — and owned by no RMn for two releases. @publisher-allowlist-derived · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES
  • ✅ RM111 — ✅ shipped in just-dna-format + just-dna-compiler on 2026-08-20. Shipped strings said a registry overrides the authored license; the registry's publish path never writes it, and what the compiler actually does — _check_declared_license_agrees warns on a contradiction — is the opposite of overriding. Strings corrected, behaviour untouched, module.version named as the field that genuinely is stamped. Four sites, not three: the item's normalize.py:40 is the identity-keys note and says nothing about license; the real third and fourth were manifest.py's docstring and its own license field. @field-description-is-a-claim · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES

✅ The S61 lookup round — one report, and its twin the report did not reach

Filed 2026-08-21 by just-module-creator, found by running lookup_variant against real rsIDs rather than by reading it. Shipped the same day as a patch.

  • RM125 — ✅ shipped in just-dna-enricher on 2026-08-21. lookup_variant(rsid="rs4988235") returned the live coordinate 2:135851076 and a finding saying "position remains unset", in one payload. The ordinary three-valued failure: at the moment the cache link speaks, the answer is unknown — a live leg that has not run may still fill it — and the link asserted rather than withheld. An earlier fix had corrected the same line's other half (it stopped saying "not in Ensembl" and named the snapshot it searched), leaving it speaking for the rest of the run. Of the reporter's three candidate shapes, one loses the cache-warming signal they explicitly wanted kept and one is unimplementable at the emission site, which leaves theirs: the caller reconciles. Cheap, because enrich() discards both links' warnings into _, so lookup_variant is the only reader either sentence ever had. Each link now reports what it searched; lookup_variant says "{rsid}: position remains unset" once at the end, guarded on there being an rsID (a position-only lookup fills rsid_candidates, never loci). Widened past the report: the ClinVar twin, documented as signature-identical "one implementation, no drift", had neither correction and produced a second false claim in the same call. Held by a green suite because no test pinned either phrase and every fixture passed an empty cache dir, so the miss line was unreachable. Severity was already info — the level was right, the sentence was wrong. A patch: advisory findings are written nowhere. · from S61 (just-module-creator) · also in CHANGELOG, ENRICHER

The 2026-08-21 output-contract round — what a patch may change about a compiled artifact

One report from a new consumer, just-dna-registry, whose catalog sweep correctly found nothing to do while an indexed manifest field went stale underneath it. Probing it found the reporter had understated their own case, and that two of our documents size the same change differently.

  • RM126 — ✅ SHIPPED in 0.7 (2026-08-28: the record, needs_recompile, the roster, the sweep and the release gate) — owed, not parked, and now discharged: the charter now requires this channel (P3: a corrected derivation may ship in any release but never silently), so until it exists the Constitution names a surface that is not there. A consumer can ask is the stored input still legal? (validate_spec → ok) and was this compiled under an incompatible compiler? (versions → a patch is compatible), and neither is would recompiling produce different output? Answering it today means enriching to a scratch dir and recompiling, which is the operation rather than a triage for it. Reproduced wider than reported: all 16 reference_examples/ compiled under v0.6.1 and 0.6.6 from byte-identical spec inputs, an interval that is entirely patch releases — 16/16 changed a published manifest field, 10/16 moved artifact.digest, 0/16 moved content_signature. The digest movement is studies.parquet +257 bytes on each of the ten, which is RM120's curator column, first in v0.6.5 — so the parquet schema moved in a patch, which the reporter had not seen. Authored identity held throughout, which is why no existing surface can see any of it. Asked for as an interval-keyed declaration (compiled_under → current) with the axes separated and explicitly not a should_rebuild verdict — the same fact costs a registry an immutable PATCH and just-dna-lite a free rebuild, so the decision stays the consumer's. Three constraints: unknown-interval is a state, never an empty result (None is never False); content_signature gets its own axis because for a registry it is a permanent duplicate-content claim; and the map must be a measurement, not hand-kept — the sweep above is the guard's prototype. Open on the representation (interval tables are O(releases²) unless composed) and the tier. Narrowed 2026-08-24 by S65, from the consumer who shipped the recomputation half: a roster of which manifest fields are pure functions of the authored rows shrinks this item rather than growing it (spec_tables + module_stats already answer that whole class); the roster's boundary is conditional, since stats is re-derived post-drop only when the drop removed something, so a recomputation is permanently the pre-drop side for a module that lost the sole row naming a gene; convergence makes the interval shape correct rather than merely tidy, because the interval from a version to itself is empty and a field-keyed shape would mint a PATCH every run forever; and compilation.dropped_rows shipped 2026-08-24 to close the residue their variant_count guard could not see. Scoped for coexistence, at their request: we state what a release did, they check what a stored artifact says. · from S62 (just-dna-registry), narrowed by S65
  • RM127 — ✅ CLOSED 2026-08-21 — filed, rewritten and answered in one pass. The Constitution was amended the same day: P3 gained release class and artifact staleness are different axes and authored identity is not the sizing test, the charter gained a Rules only header item, and Principle 9 promoted the cost-by-layer pricing that had been stated only inside an amendment entry. Reasoning moved to CONSTITUTION_AMENDMENTS_HISTORY.md; the charter came out 11.5% smaller while gaining three rules. First filed as the release-class table and the practice disagree, indicting StudyRow.curator shipping in 0.6.5 — withdrawn, because curator is additive, no published module can carry it, and no stored value became wrong, so a patch is defensible and the table is merely strict. The original framing is preserved in ROADMAP_HISTORY because it is the tempting reading. The defect is RM121 alone and its change class does not exist in the taxonomy: an existing published field whose derivation was corrected, so the same spec yields a different value — neither additive nor removal/retype. Why it read as safe is the keeper: the sole test applied was does authored identity move, and stats sits outside content_signature by design, so that test cannot fail there — a tautology read as a pass (@tautology-zero), one level up from where we usually catch it. The structural edge: a field outside identity is one no digest, signature or revalidate can see move, so the cheapest changes have no detection channel — measured, six of sixteen examples changed a published indexed field with both hashes byte-identical. A corrected derivation is a bug fix, so deferring it to a minor means serving a wrong value meanwhile; the version number answers is the code compatible, never are your outputs stale, and the axes must be separated rather than reconciled — which dissolves the three original candidates. What is left is one charter question for the maintainer: does P3's sentence get amended to decouple release class from artifact staleness? · from S62 (just-dna-registry)
  • RM128 — ✅ SHIPPED in 0.7 (just-dna-enricher), all three asks. The corruption half shipped 2026-08-24: every sidecar writer goes through layout.atomic_writer / atomic_write_text, so a killed process leaves the previous table rather than a syntactically valid short one — nine writers routed where three were reported, walked by an AST guard rather than counted. The short file was the dangerous half because resolution.csv is read back and merged under a subject key (S51), and the three branches that deliberately write no row for an unanswerable subject make a truncated table indistinguishable from a module whose author resolved less — the reported incident had a zombie run overwrite a restored 330-row table with 162 rows, after which the module validated, closed and compiled green. The central ask dissolved: the choice looked like keep the promise that a refused strict run changes nothing, or recover the thirty minutes, and it is not a choice — the run becomes a transaction, staging each live answer beside the target and committing once at the gate, which keeps the promise absolutely and recovers the work. Same-directory staging is the correctness condition (a cross-partition shutil.move is copy-then-delete), what is staged is the raw answer so every derivation recomputes and a resume reproduces an uninterrupted run, and only positive answers are staged (@unreachable-not-absent). --keep-staging keeps them for debugging; write=False stages nothing and locks nothing, so @flag-means-same holds. The lock is flock on the spec directory, non-blocking, no lockfile — a lockfile left by exactly the kill this item is about would block every later run, and the staleness rule that would fix it is a clock (@hash-the-probe) — with a documented degradation for a platform or filesystem that will not take it, both branches reached by test. Progress is (done, total) over subjects, argued rather than guessed: the incident is an idle timeout so phases are ruled out, total must be known up front so links are ruled out, and subjects are the unit an author already has. enrich --rederive landed here too as RM83's residue, composing with the transaction rather than adding machinery. · from S66 (just-module-creator) · also in CHANGELOG, ENRICHER, MODULE_LIFECYCLE, PROPOSAL_0_7
  • RM129 — ✅ shipped in just-dna-format + just-dna-enricher on 2026-08-24. verification.json's producer sat at the document level and record_verification refills it on every write, so a merge that correctly kept an older run's record restamped its attribution — a 0.6.4 clinical_significance record came back attributed to 0.6.6. The reporter's argument is from the other fields: source, release and checked_at all describe one piece of work and sit on the record, and producer, naming who ran it, was the only one that did not. Shipped VerificationRecord.producer: str | None beside them, with the document-level field kept and re-described — it pairs with produced_at as what last wrote this file, and its old description ("Tool and version that put the checks") was the false claim itself, printed by describe/reference (@field-description-is-a-claim). Outside VERIFICATION_FACT_FIELDS on the reasoning that excluded checked_at, so no published verification.signature moved — tested, not assumed; None on an older record reads as not recorded rather than as a release, since defaulting it to the reading version manufactures the attribution the item is about. merge_records carries whole records, so the value travelled for free. The merge itself was not changed and the reporter's note that it was correct is what kept the repair off it. · from S71 (just-module-creator) · also in CHANGELOG, SCHEMAS
  • RM130 — ✅ SHIPPED in 0.7 (2026-08-28). The observability half shipped 2026-08-24 — clinical_significance writes a detail grouped on opposed, and _findings_warning reports a non-zero findings at validate/compile, where nothing read VerificationRecord.findings at all and twenty contested rows out of 141,616 reached only a consumer who went looking. The sidecar is the other half, and it shipped in RM134's shape rather than the one this entry first decided: clin_sig_concordance.csv keyed (variant_key, genotype) carrying authority_concordance (concordant/discordant/single/none/unchecked) and authored_position (matches_all/matches_some/matches_none/absent/unchecked), with each authority's own call in the paired clin_sig_authority_calls.csv keyed (variant_key, genotype, authority). The original shape — the authored value beside the source's — named its authority in a field (ClinSigConflict.clinvar), which would have cost a key change or a retype one item later when RM134 arrived; two items in one release is what caught it. Nothing resolves a split and confidence is never normalized across authorities. A conflict is a question and an overrides.csv row is the answer, so the record joined the overlay's covered set (seven → eight) while the detail table stayed out by name — an author answers the question and does not rewrite what an archive published. Warning-tier in both modes (@clinsig-never-escalates), actionable rather than carried, because the count is over the post-overlay rows. · from S70 (just-module-creator) · also in CHANGELOG, ENRICHER, SCHEMAS, COMPILER
  • RM131 — ✅ SHIPPED in 0.7 (2026-08-28: carried and warnings_summary beside an unchanged warnings, 69 codes of which 9 are carried, every emission site in three tiers classified, and RM124's suppressed rows reported by reason). Both halves in one release, because the audit is the cost and the sequencing existed to avoid doing it twice. All three result models carry warnings: list[str] with no code, no count and no actionability; a 190-row module returned ~14 kB, and strict=false changes what counts as an error, not how much prose the channel carries. The reporter's sharpest point is that we compute the answer and spend it on severity alone — _BLAME_TIER/_BLAME_ROW is whose limit this is and its comment says "blame decides severity and nothing else", with _closure_warning reaching the same distinction from the other end. S67 is the identical shape one level down and was fixed there. Why the offered minimal version is not free: warnings_summary: dict[str, int] is additive and harmless, but the code is permanent under P3/P6, so the first set shipped is the one consumers key on forever and it has to be derived across ~29 append sites and 16 returning helpers never written to be classified — the container is free and the vocabulary is the release. Three candidate derivations, none obviously right: from the pinned catalogue (honest but partial by construction, and a digest that silently omits is worse than none), from the emission site (complete but a refactor then renames a published key — P3's rename arriving by the back door), or from the check as a first-class argument (most work, most stable, precedent in VALID_VERIFICATION_CHECKS). The actionability half was designed first — it needs no vocabulary, blame already classifies two families, and the reporter's own fallback (a carried list beside warnings) invents nothing — and then shipped in the same release as the codes, since the audit is shared. No cap, no truncation, no verbosity flag, per the reporter and us. Two figures in this entry were wrong and were re-derived rather than trusted: the append-site count missed the findings/messages collectors, the two .extend sites into the schema tier, and the deprecated resolver in the enricher, whose warnings reach the same published channel. The str-subclass transport leaks a code at a pydantic field and at any reformat, which is why the two public entry points that continue a run read a classified list beside the result and three prefixing sites go through findings.restate. · from S68 (just-module-creator) · also in CHANGELOG, COMPILER, AGENT_NOTES
  • RM132 — ✅ SHIPPED in 0.7 (2026-08-28: PharmVariantRow.pmid, both literature cross-check sites, and a derived citing-kind roster the enricher reads through). A ClinPGx-drafted module carried 1,482 drug-response rows with nowhere to cite any of them — sixteen model fields, thirteen authored, none a PMID or DOI. The reporter asked which of three provenance models was intended; the tree already answered it one release ago (@rm47-bin-cites: the bin row cites, the citation table describes), under a rule worth stating generally — a row cites when its claim is finer-grained than studies.csv' key. studies.csv keys on (variant_key, pmid) so a study attaches to a variant, while pharm_variants.csv keys on (variant_key, drug, genotype, phenotype_category, annotation_id), so one study row would attach to every drug/genotype/phenotype for that variant — which the reporter worked out and declined to build on. So the column is missing rather than deliberately absent: evidence_level points at somebody else's grading of evidence, not at the evidence, and the licence row states redistribution terms rather than grounding. The fix is PharmVariantRow.pmid plus both literature cross-check sites in the same release (_cross_check_literature, enrich_literature), which is RM47's recorded lesson verbatim — a column without them makes every new citation read as a stale orphan. Enricher reaches the rows through public compiler symbols, never a second roster (RM40/RM41) — shipped as load_citing_rows/table_citations over _CITING_TABLE_KINDS, every _TABLE_KINDS model declaring a pmid, with an ast walk over the enricher's own source asserting no roster is kept there; load_binning_rows/binning_citations stay narrow. The open question is answered: provenance_quote does NOT follow — the row cites, studies.csv/literature.csv describe, which is what stops StudyRow's provenance column set migrating one column at a time; the consequence is that a pharm row contributes a denominator of zero to the quote-counter check rather than being skipped. literature_row_uncited's sentence gained the third site (the code did not change). · from S73 (just-module-creator) · also in CHANGELOG, SCHEMAS, COMPILER, ENRICHER
  • RM133 — ✅ SHIPPED in 0.7 (2026-08-28, as a registry-owned short_description key beside IDENTITY_AUTHORITY_KEYS, with a published ~120-char constant); the binding question it arrived with was answered and CLOSED, and the binding is deliberately untouched. Measured: editing module.description from 44 words to 11 moves no content_signature, no artifact.digest and no fact signature — and drops the closure, since manifest.inputs covers the raw bytes of module_spec.yaml. README.md, far longer, sits outside inputs and is freely amendable, so the shortest fixable prose in the system was the one that could not be fixed. The binding stays, and the reason is the partition rather than the cost: the ask was to split it along content_signature's line, which excludes name, version and namespace alongside title and colour — so a binding drawn there makes a closure transferable across a rename. That exclusion exists so a registry strip does not move content identity, which is right for a dedup key and exactly wrong for an attestation; the two cannot share a partition. Do not re-propose. The reporter's narrower six-field version lacks that attack but inherits the cost they named — hashing a parse, against RM82's deciding property that a binding change be a byte transform needing no loader, no parse, no schema knowledge. What unblocks it is registry-owned metadata, not a binding change: IDENTITY_AUTHORITY_KEYS is the standing precedent for values stored beside a module rather than inside it, so an amend_display that stores an override leaves manifest.inputs matching and the closure standing — the registry's endpoint is therefore not gated on us. Left to design: where a bounded short_description (~120 chars) lives so it lands amendable — not on ModuleInfo, where it would reproduce the defect. · from S64 (just-module-creator) · also in FAQ
  • RM134 — ✅ ALL FOUR SECTIONS SHIPPED in 0.7 (2026-08-28: the PubMind snapshot, pubmind build/publish, and one shared significance normalizer carrying both measured token fixes; 2026-08-30: draft-panel --source pubmind, pubmind in DRAFT_PROJECTIONS, and the hint's PubMind leg — the gene map is ClinVar's per-record attribution matched at the exact position, never a span, since PubMind names no gene) in the same release. PubMind (doi:10.1038/s41467-026-76834-4) mines variant–disease–pathogenicity associations out of 41.7 M abstracts and 5.4 M full texts with LLaMA-3.3-70B. It is a source, not a competitor — its own discussion calls it "a literature-grounded complement to human curated databases", and the Constitution's no-gene–disease-inference non-goal settles the rest: the thing it does best is the thing we are constitutionally not in. Probed rather than read: the API takes gene/MONDO/PMID only and returns aggregates, so the one per-variant channel is ANNOVAR's hg38_pubmind_db (909,224 rows, VCF-style coordinates despite the packaging, so no translation). 48 % of it is enumerated codon alternatives, not observed variants — decomposing leaves 342,209 joinable keys over 305,935 loci. Consolidation is on extracted text, so it has record identity where we have variant identity: 8.4 % of coordinates carry several PVIDs, and at HFE C282Y one of eight pairs rs1800562 with a chromosome 22 gene while calling it Benign — the shape _gene_locus_conflicts already catches. On our own corpus: 40.9 % of loci matched, 32.3 % of authored ALTs, 62 % verdict agreement. Designed in four sections on a 2026-08-28 direction, superseding an earlier draft that stopped at a report-only check and called the rest blocked: (A) pubmind build/publish mirroring clinvar_build.py, every normalization drop counted into release.json, publish refusing on the PharmVar precedent (@gated-source-caches); (B) a three-way module ↔ ClinVar ↔ PubMind check with seven Kleene outcomes, authorities_differ the one nothing today can report, reusing ClinSigConflict.opposed for severity and never escalating (@clinsig-never-escalates); (C) --source pubmind on draft-panel rather than a twin command, with pubmind in DRAFT_PROJECTIONS so a module drafted from PubMind cannot confirm itself (@draft-digest); (D) a hint that may not pre-fill the cell B cross-examines. "Authority" throughout means an authoritative annotation source à la ClinVar, never resolution.csv's authority (@source-vs-authority). The gate is one question: the ANNOVAR-shipped table publishes no data terms, unknown is not permissive (@no-named-licence), and the unblock action is to ask CHOP in writing — separately from RM27's undesigned redistribution axis, which gates publishing such a module rather than building the snapshot. ✅ Decided in PROPOSAL_0_7 on 2026-08-28 — BUILDS in 0.7, with eight corrections found by reviewing it against the eleven items already decided: the concordance record is shared with RM130 and supersedes RM130's shape (an authority in a field name costs a major once a second one arrives); one normalizer, not two, after fixing the whitespace gap that sends Uncertain significance and Conflicting to other; pubmind_sig → clin_sig; derivation's P5 question closes (the snapshot is never published, so the name is not a one-way door — audit --min-confidence and the DRAFT_PROJECTIONS key instead); the three-way check subsumes the two-way; RM71 lands before --source pubmind; withheld draft rows are named at draft time; and the new checks sit inside RM131's audit and are what forced RM126's warnings axis. A five-authority stress test failed the drafted vocabulary — one field carried two axes — so it split into authority_concordance (concordant/discordant/single/none/unchecked) and authored_position (matches_all/matches_some/matches_none/absent/unchecked), five members each at any N. Nothing resolves a split: the precedence list is recorded as methodology and computed with by nothing. · from competitor assessment 2026-08-28, extended by user direction the same day · detail in PUBMIND_ASSESSMENT
  • RM135 — 🟣 queued for 1.0, filed 2026-08-28 by PROPOSAL_0_7. RM124's overrides.csv and ProvenanceItem.outranks (S52) are one concept in two files — an authored value beating a source, with prose — split on authored-versus-derived, which is the line P5 says to draw before either grows a second field. Decided as a succession rather than a merge: both stand in 0.7, the duplication is stated, and the survivor is the simpler rule — the existence of an override in an authored table auto-beats the derived value, no separate declaration. A removal, hence major-only under P3. No deprecation warning in 0.7, and that is P3 rather than caution: the overlay covers derived tables while outranks covers an authored cell, so an author warned off it today has nowhere to go, and a warning nobody can clear is noise rather than notice — the warning ships with whichever minor extends the overlay to authored tables. · from PROPOSAL_0_7, 2026-08-28 · also in history/ROADMAP_0_7 § RM124
  • ✅ RM136 — shipped 2026-08-31 in the uncut 0.7.0 (just-dna-compiler loader made public + just-dna-enricher). The compiler applies overrides.csv before any check reads a row; the enricher re-read the raw derived file, so an author's correction kept being reported on every run, forever, with nothing saying it had been honoured one tier over. Read-only, at INPUT reads, per field — three choices, each with a refused alternative: the enricher never writes through the overlay (RM83); merge baselines stay RAW, because a pass that reads its own output writes it back and post-overlay rows would bake the correction in; and per-FIELD, so correcting a coordinate answers the coordinate check and leaves an unrelated finding standing — per-row was cheaper and would silence findings the author never looked at. No second apply_overrides, which is the entry's central refusal: compiler.load_overlay went public (the S74 shape) and licensing.overlaid_input_rows calls the real one, asserted against it by test. Answered is not agreed: the pair leaves disagreements, stays in subjects, and PairCheck.answered counts it — dropping it from the denominator would report a cleaner module than there is. One check consults it today (rsid↔coordinate, the finding an overlay row on resolution.csv can actually answer); wiring is deliberate per check, since which cells does this comparison read is a per-check fact and guessing it silences a finding nobody answered. · from the RM124 wave-1 audit · also in CHANGELOG, ENRICHER, INTEGRATION_0_7
  • ✅ RM137 — shipped 2026-08-31 in the uncut 0.7.0 (just-dna-format + just-dna-compiler). An overlay update on a row the compiler drops matched on lap 1 and warned on lap 2, so a module and its own round trip disagreed on manifest.compilation.warnings. "Count it over the overlay's own rows" needed one step the entry did not have: counting the update rows outright is a tautology, counting the ones that reached nothing is the lap-dependent original, and the stable quantity is a property of the TARGET — could an artifact of this module carry that row at all, which is is the pmid cited / can the module place that variant_key, both computable from data that survives the trip. The unreachable finding fires matched-or-not, and that asymmetry IS the fix — an earlier cut classified only the unmatched set and was silently lap-dependent again, caught by asserting equality BETWEEN LAPS rather than "lap 2 warns", which passes on the broken code. Neither reading is "a typo", correcting the entry: a mistyped pmid is also an uncited one, so a mistake lands in the unreachable bucket, and the reachable bucket means the narrower and more useful the table is short — re-run enrich. New code overlay_update_target_unreachable (actionable); overlay_update_unmatched reworded. Scoped to LOSSY_OVERLAY_TABLES (literature.csv, resolution.csv), asserted as a registry equality — the other six rebuild whole, so their warning is lap-stable already. Two build traps: the predicate shares cited_pmids with the drop rather than restating it, and it mirrors the drop's empty-cited guard or a module citing nothing gets a stable false positive. Classified late in both validate_spec and compile_module via defer_unmatched=True, because the inputs do not exist where the overlay is applied; hoisting the studies load was refused, since pre-flight warnings seed the compile's list. · from the RM124 wave-1 audit · also in CHANGELOG, COMPILER, INTEGRATION_0_7
  • ✅ RM138 — CLOSED 2026-08-31 with no code change, inside the uncut 0.7.0. carried holds full message text, growing compilation 1.84× across the corpus (1.96× on pathogenic_clinvar, the 113-warning module it was filed about). Not a defect — the shape was decided per item and both its properties hold; what the decision lacked was the size. Re-measured with the compression a real transport uses, and that settles it: 1.06× gzipped over the corpus, 1.13% worst case. carried is a verbatim subset of warnings, which is exactly what DEFLATE's back-references eliminate, and the whole with-carried payload gzips to 0.21× the uncompressed warnings-only one. The raw column reproduces the entry's own figure, which is what makes the second trustworthy. Decision: keep the encoding, recommend compression where the size lands (catalog, API, anything shipping manifests over a wire) — a deployment concern, not a schema one. The three cheaper encodings stay rejected and are now cheap to reject: indices break the subtraction and positionally couple two published fields, a code list answers a question warnings_summary already answers, and a per-message codes is a third shape that hands consumers a derivation where they read an answer. Closed inside 0.7 deliberately: a fourth encoding after 1.0 is a removal, hence major-only under P3. · from the RM131 review · also in CHANGELOG, INTEGRATION_0_7
  • ✅ RM146 — shipped 2026-08-31 in the uncut 0.7.0 (just-dna-format; additive — a marker on each field declaration, no column and no signature). A module authored on 0.6.6 hit a registry running 0.6.1 and got studies.csv line 2 [curator]: Extra inputs are not permitted — byte-identical to what a typo ([curatr]) produces, and the two want OPPOSITE actions. pydantic's message under extra="forbid" cannot carry the distinction because the information was not in the model. base.since("0.6.5") on every authored field, read by base.field_first_seen(model), composing with vocabulary() in one json_schema_extra; stamped_identity_field takes first_seen as a REQUIRED argument, since a stamped column is still one an older reader refuses. Backfill measured, not recalled — parsed per tag from the AST rather than imported, because old code need not import under a current Python: 414 fields across 31 models, 115 from 0.2.0, 81 from 0.4.0, 150 from 0.5.0, 160 from 0.6.0, 3 from 0.6.5 and 78 landing in 0.7.0. curator is why the answer is per (model, field): VariantRow at 0.2.0, its StudyRow twin at 0.6.5, so a name-keyed roster gives one answer for two facts — the wrong one for the module in the report. Guard is a set EQUALITY over the walked registry, plus a check that the registry is complete (RM96's shape) and that every declared version is a release that exists. Two things the build turned up: the mechanical pass wrapped nine ClassVar constants in Field(...) and the suite caught it, unwrapped by AST since the multi-line forms are invisible to a regex; and the entry's "402 fields" was written before 0.7's own additions, which is why the test asserts a floor on the total and equality on the coverage. · from S81 (just-dna-registry) · also in CHANGELOG, SCHEMAS, INTEGRATION_0_7

✅ The 0.7 build round — PROPOSAL_0_7.md decided, then built

The 2026-09-01 source-adoption round (RM163–RM168) lands in the same uncut 0.7.0 and is a second round rather than an addendum to the first — PROPOSAL_0_7_PT2 is its record. Its shipped items are indexed here, beside the twelve; RM164 parked and is filed under Deferred to a later minor since 2026-09-11, and RM171 — the spin-off it filed — shipped on 2026-09-03.

  • ✅ RM168 — ✅ SHIPPED 2026-09-01 in the uncut 0.7.0 (just-dna-enricher only). MANE stops being a sentence in CIVIC_IDENTITY_PROTOCOL § 3b and becomes a cache: MANE_TERMS, $JUST_DNA_MANE_CACHE, mane_build.py, a mane build sub-app, a cache status row. Three files in one pass — the summary (19,437 rows, 74 MANE Plus Clinical, and MANE_status is never collapsed because CDKN2A's two rows are the case the item exists for), changed_select_accessions (120 rows; Update_Affects_CDS Yes on 74 is the numbering-frame axis stated by the source, so the currency check is read one small file rather than diff two releases), and protein_coding_genes_not_in_mane (222 genes over a 7-member reason vocabulary derived from the file, so @unreachable-not-absent is served by the source and pending MANE review is a third state). release.json is copied from the 96-byte README_versions.txt, which states two releases no filename carries. Pinned by the versioned directory — current/ discovers a version, release_<v>/ pins it (@current-discovers-a-version-a-directory-pins, the new gotcha). Terms are NCBI's policy, not a licence: every gating axis None, no --use flag, and only NCBI's side read. The bound ships with it: MANE is the default, not the answer — it makes CDKN2A's problem visible and is silent on RUNX1's derived 27-residue offset. The one item of the six whose probe the build did not move. · from the 2026-09-01 source-adoption round · related RM159, RM153, RM152 · decided in PROPOSAL_0_7_PT2
  • ✅ RM163 — ✅ SHIPPED 2026-09-01 in the uncut 0.7.0 (just-dna-enricher + two verification-check members). A fourth registry in identifiers.py asks the PGS Catalog about every authored pgs_id, plus a pgs.py client, PGS_TERMS as a per-score licence floor, and pgs_catalog as currency's second probe. The verdict is read off the body: 200 + {} comes back for a never-assigned id and for a malformed one, so the status carries no existence information (@existence-not-identity). The absence message inverts @rsid-absent-two-readings — only ~35 % of the accession range is assigned, so a typo is the likely reading and withdrawal the rare one, named but not as equals. Drift is two fields, not the entry's four: the Catalog publishes no match_rate_floor and no research_tier, and a check with no source-side value cannot fail (@tautology-zero). The licence half is a correctness requirement: license is per score and varies, so the constant is a floor each score's own string overrides — and a module naming one academic-use-only score is now refused by the compile gate by name, where a flat constant would only have warned. Attests under two members, because currency and drift have different subjects and denominators. Four contradictions recorded, the sharpest being that overrides.csv cannot answer a pgs.csv finding at all — the overlay is derived-tables-only and this table is authored. · from the 2026-09-01 source-adoption round · related RM16, S86, RM153 · decided in PROPOSAL_0_7_PT2
  • ✅ RM165 — ✅ SHIPPED 2026-09-01 in the uncut 0.7.0 (just-dna-enricher + one verification-check member). STRchive (MIT, 82 loci) adopted by column: check-repeat-bands over the band table, draft-repeats over the identity half, as two commits. The split is the finding: the catalogue reproduces htt_repeat_expansion's first two bands exactly and gives FMR1 one 45–200 where the module has 45–54 and 55–200 — losing 55, the premutation threshold — so a provider would have written that erasure as the answer. pathogenic_max is emitted nowhere: 250 is the largest allele the literature reports, not a clinical bound, and as measure_max a 300-repeat allele would match no bin at all, silently, --strict included (@bin-grounding). Warns in both modes. Four contradictions, the first structural: the identity half is mostly uncarryable because RepeatAlleleRow has no column for coordinates, locus_structure, ref_copies or the disease ids — which is RM65/RM87 rather than a shortfall here — and a pre-existing crash in compiler.draft.append_partial_rows (any table whose header is narrower than its model, all four partial-row providers) was found and left for its own item. Named, not built: RM66's evidence (locus_structure on 23 of 82) and RM170's second instance (STRchive's evidence carries Disputed 3 / Refuted 1). gnomAD's TR release stayed out on category — one row per sample per locus. · from the 2026-09-01 source-adoption round · related RM65, RM66, RM87, RM164, RM170 · decided in PROPOSAL_0_7_PT2, which had held the provider to 0.8
  • ✅ RM167 — ✅ SHIPPED 2026-09-01 in the uncut 0.7.0 (just-dna-enricher + one verification-check member). litvar coverage asks LitVar2 per module locus and names the tier that answered — allele-resolved, position-only, absent, plus unchecked — because 328 and 3,945 are both true of rs429358 and only one is about the allele in the module. The fall-back to the position node is never silent, and the position-only residue is counted over the union of every allele node. Writes no row, the entry's pre-authorised outcome: a PMID list is not a table kind. Corpus measurement: 180 of 389 loci (46.3 %) have a CAID node; 14,168 papers sit on a position node no allele node claims, 6,700 of them APOE's. Three of the proposal's numbers did not reproduce — HFE has one gene node and 298 text mentions, the 423-locus join is 389, and the stated id grammar is contradicted by its own example. The bound ships with the pass: against the two CIViC legacy insertions it returns no node for any of the four candidate alleles, because PubTator3 has abstracts and the alleles are in a paywalled table — it answers which papers discuss an identified allele, never which allele a name meant. Terms are NCBI's policy, so every gating axis None. Also repaired clingen_allele._parse, which discarded a one-sided allele whenever an rs-number arrived. · from the 2026-09-01 source-adoption round · related RM134, RM153 · decided in PROPOSAL_0_7_PT2, which had reversed twice and proposed 0.8
  • ✅ RM166 — ✅ SHIPPED 2026-09-01 in the uncut 0.7.0, split: the cross-check built, the licence half closed measured. A drugLabels.zip builder beside clinpgx_build (same cache, payload-read LICENSE.txt, its own CREATED_*.txt and therefore its own release.json — relationships.zip was once a year newer than clinicalAnnotations.zip), plus a cross-check joining at two tiers with the tier distinguishable in the finding, because a gene-level agreement and an allele-level one are not the same claim. It is five regulators, not one — FDA, Health Canada, EMA, Swissmedic, PMDA — so the number of authorities is a parameter and the surface is named for the labels, never for an agency (ClinSigConflict's mistake). A blank Testing Level withholds: a third of the file states none, and reading that as No Clinical PGx would manufacture a negative regulatory claim. Warns in both modes. The licence half closes on measurement: the ClinPGx route is the same CC BY-SA + no-sale gate, and the FDA's own table is 126 rows of HTML with no stated terms at all — so lane diversification wants its own entry, choosing candidates for their terms first. Noticed, not built: ClinPGx publishes ≥12 archives and the builder reads one; clinicalVariants.zip is pharm_variants.csv territory with a comma-combining type. · from the 2026-09-01 source-adoption round · related RM134 § B, RM29b · decided in PROPOSAL_0_7_PT2, which proposed 0.8 — largest and least urgent is an argument about order, not release
  • ✅ RM176 — ✅ SHIPPED 2026-09-02 in the uncut 0.7.0, owner enricher, severity high. Asked of every cache but Ensembl: is there a common rebuild endpoint, and does each lane have download + build + upload? No, and the three gaps were one defect — the roster was a four-tuple list in cli.py, so nothing compared it to the set of things it was a list of. Three lanes were not in it (acmg, strchive, drug_labels), so cache status reported nine caches on a machine that has twelve; the same three had no resolver, so each check took an explicit path and the flagless route fell through — ACMG's to scraping NCBI's v3.2 page while the snapshot holds v3.3, reporting a correctly authored row as wrong; three had the licence to publish and no way to (CIViC CC0, STRchive MIT, the drug labels CC BY-SA), which the old roster's own comment already called a gap rather than a refusal. What shipped: caches.CACHE_LANES, a registry walked by a test against the *_build modules on disk, carrying each lane's three stages and the reason as a field for every stage it lacks; three resolvers wired into the flagless branch of their own checks (the tests assert the call, not the resolver — @ensure-must-be-called); strchive publish and clinpgx publish-labels plus ensure_civic/strchive/drug_labels_snapshot; and cache rebuild, one endpoint over eleven builders calling the same download_*/build_* the per-lane commands call. Two premises generalized: plan_reference_snapshot takes a payload filename from its caller (two snapshots hold no parquet at all), and the publisher's allowlist is derived from the plan rather than restated — the drift that lost citations/ and LICENSE.txt a release each (@publisher-allowlist-derived). The outcome is three-valued: ACMG needs an Elsevier workbook, PharmVar a personal key, CIViC a pinned release, Ensembl is built elsewhere — not run, never failed, so a nightly rebuild does not alarm on four lanes behaving as their licences intend. Builds into <base>/<lane>/, never in place. Two defects the suite caught: publish-labels had landed after the __main__ guard, and cache status composed f"{name} build" — right for ten lanes, naming two commands that do not exist. Builder deps were already [dev]-only; nothing moved. Out of scope by decision: the PGS/PRS parquets under just-dna-seq are just-prs's Dagster pipeline. · from the maintainer's 2026-09-02 question · also in CHANGELOG, ENRICHER § The caches
  • ✅ RM175 — ✅ SHIPPED 2026-09-02 in the uncut 0.7.0, owner enricher, severity high. PharmGKB renamed the table on 2025-07-29 — clinical annotations are now called summary annotations — and the archive followed. clinicalAnnotations.zip was last written to S3 2025-07-05, 24 days before the rename post, is on no downloads page, and the API still answers it 200 through a 303 to the frozen object; clinpgx_build.DEFAULT_CLINPGX_URL named it, so every annotations.parquet this lane ever built rested on the database as it stood 14 months ago. What shipped: the default URL, two member names and the id column moved to summaryAnnotations.zip / summary_annotations.tsv / summary_ann_alleles.tsv / Summary Annotation ID, with no vocabulary member, no model field and no parquet column changed — the other fourteen columns and Phenotype Category's values are identical. And a guard, which is the item: require_current_archive reads the member names first and refuses the retired spelling by name, quoting the rename, its date and the URL to use instead — three arms, three diagnoses — because a retired filename that still 200s parses fine and yields a plausible parquet (@specific-rejection). Both spellings live in one table the reader takes its names from, and RETIRED_ARCHIVE is returned by nothing, so no path through the module can read a 2025 archive; no both-vintage compatibility layer, deliberately. The data moved: 5,179 shared ids with 7 gone / 11 new, 8 changing Level of Evidence, 40 Drug(s), 14 Score, 68 Level Modifiers, every URL rehosting — at the parquet's own grain 16,087 → 16,117 rows over 5,190 annotations, 30 rows changing evidence_level. The digest moves, and a drafted module can see a level change under it. The docstring's 4,618 of 5,113 is gone from all five live files, restated as a relationship rather than swapped for a fresh count. Fixture is now assets/clinpgx_annotations_slice/, cut verbatim from the 2026-08-05 archive, with the retired-vintage zip built from the same rows. Left unbuilt on purpose: the three currency canaries the entry listed (URL audit, S3 Last-Modified in release.json, sibling-age comparison) — each is its own design, and nothing here would notice summaryAnnotations.zip itself going quiet. And no-JS fetches of clinpgx.org are no evidence at all: every HTML route serves the same Javascript Is Disabled! shell, so the downloads listing took a browser. · from the maintainer's 2026-09-02 investigation · supersedes RM173 · also in CHANGELOG, ENRICHER § Pass 6, probes/CLINPGX_ARCHIVES
  • ✅ RM184 — ✅ SHIPPED 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; no behaviour change). CacheLane.env_var, the variable that overrides a lane's location — the one attribute the RM176 registry still left consumers hand-keeping, fourteen names in three places. The literals moved out of the resolvers into locations.<LANE>_CACHE_VAR beside CACHE_BASE_VAR, so field and behaviour are one string; pinned per lane on behaviour and by an equality over every JUST_DNA_* the module declares. str, not optional — no lane lacks one, and inventing the state was refused on RM87's argument. · from S89 · related RM176, RM182 · also in CHANGELOG, ENRICHER § The caches, INTEGRATION_0_7 § 2.4
  • ✅ RM183 — ✅ SHIPPED 2026-09-03 in the uncut 0.7.0 (just-dna-format only; no schema, parquet or manifest change). needs_recompile(None, current) — and "", and whitespace — answers the unknown arm (every axis None, complete=False, compiled_under=None, span=(None, current)) instead of AttributeError from a .strip(), because Compilation.compiler_version is str | None and a registry walking manifests it did not produce will meet one. Absent is not malformed: a present, unreadable stamp ("0.7", "v0.7.0", a trailing note) still raises, now quoting the whole stamp rather than its last token. Two fields widened to str | None. · from S88 · related RM126 · also in CHANGELOG, SCHEMAS § the release record, FAQ, INTEGRATION_0_7 § 2.8
  • ✅ RM180 — ✅ SHIPPED 2026-09-03 in the uncut 0.7.0 (just-dna-format + a compiler comment and two tests; no parquet column, no vocabulary member). OverrideRow.reason/decided_by/decided_at are outside content_signature — rewording a reason is a patch, as a README caveat is (S25) — and the six value cells stay inside. Not exclude=True: the stamped-column idiom emptied reason in every model_dump() writer (draft._authored_dump, the enricher's overlay tests), so the mechanism is a field marker OUTSIDE_CONTENT_IDENTITY read by integrity.content_signature alone, with content_identity_exclusions walked over _ALL_MODELS by test. Parquet, manifest.inputs, the verification binding, reverse and artifact.digest all still carry the prose. Decided before the cut because no published module carries an overlay and the window was the release. curator/method on variants.csv stay inside, and that asymmetry is stated rather than repaired. · from S87 · spawned RM181 · also in CHANGELOG, SCHEMAS § the authored overlay, FAQ, INTEGRATION_0_7 § 2.1
  • ✅ RM153 — ✅ SHIPPED in 0.7 (2026-08-31: clingen_allele.ClingenAlleleClient, an identity_derivation="caid" snapshot row, Picard-style indel anchoring, CLINGEN_ALLELE_REGISTRY_TERMS; no schema change). RM152's residue, and it answers its two questions in opposite directions. The CAID pass is taken: the ClinGen Allele Registry serves an rs-number and a GRCh38 coordinate, no key, 102 probe requests with zero failures, and it takes recovery over the direction set from 138/290 (48%) to 237/290 (82%) — 52 variants via an rs-number, 12 via a coordinate, and 35 one-sided indels anchored VCF/Picard-style, which is all 35 that previously read no_identity. The rs-number is preferred because ClinGen supplies it and Ensembl verifies it — two authorities, so the check is real, which is exactly what a lifted coordinate lacks. It runs at draft time, never in civic build, because a build that fetched would forfeit the offline byte-reproducibility the dated input exists to give. Liftover is refused, reopened at the maintainer's instruction and closed on the measurement (CIVIC_UNRESOLVED): its ceiling is 13 evidence rows on 9 variants and the honest recovery is at most one, because 3 are gene-level assertions no genotype satisfies, 5 are imprecise by the source's own HGVS (the registry refuses to parse them at all), and variant 2099 lifts exactly to a different allele than its own name describes. pyliftover agrees with Ensembl on all 18 endpoints so buys no accuracy; Picard was not run because 8 of 9 have no REF/ALT to feed it. The registry's terms are unestablished — reg.clinicalgenome.org/site/terms answers 200 with a Genboree broken-link page, so every axis is None; ClinGen's CC0 covers the gene-curation surface, not this one (@probe-names-the-table). Left open at the cut: 53 variants carry no identifier at all, and 5 of the 9 can never be reached by any identity pass. Re-measured 2026-09-01 and the first half no longer holds — all 53 were put through a four-tier identity procedure and 34 resolve from the fragments CIViC publishes in the variant's own name; 33 of them were adopted the same day as RM159, taking coverage to 270/290 (93.1%). The item's own addendum carries the correction and CIVIC_UNRESOLVED the residue class by class. The 5 stand. · from RM152's build, 2026-08-31 · related RM48, RM68, RM15, RM152
  • ✅ RM152 — ✅ SHIPPED in 0.7 (2026-08-31: civic build, draft-panel --source civic, CIVIC_TERMS; no schema change of any kind). Filed carrying no release class because both adoptions S84 proposed had been refuted by measurement, and it acquired one only when the probe it named was finally run — the refutations held and a third route nobody had proposed turned out to be buildable. The probe corrected the item's own number in both directions: PREDISPOSITION × DOES_NOT_SUPPORT reads as 1 contested subject by profile and 3 by variant (a two-variant profile's refuting row reaches both members), two of the four are lone refutations that contested does not describe, and genuine risk-vs-protective opposition is 0 at every scope and every status basis. Widening past the germline filter adds 0 contested; the assertions table is structurally incapable of the axis, since AssertionSignificance is a different 16-member enum without PREDISPOSITION or PROTECTIVENESS, so filtering by them is a type error rather than an empty result. Two findings nobody sought: the API defaults both connections to NON_REJECTED, so the 11,518 everything quotes has an undeclared denominator (ACCEPTED is 4,904 and the bulk TSV is accepted-only at 4,903) — and the coordinates are GRCh37 or absent, never GRCh38, which the snapshot survives by reading the rsID and GRCh38 accession CIViC publishes beside them rather than lifting anything (@old-assembly-vs-shift, RM48). The direction-axis concordance apparatus was refused, which was the open question the item carried: opposition is 0, and nothing else in the enricher fills direction at all. The drafter was un-refused on the axis, because its stated defect — rows with an empty significance column — is true of clin_sig (812 NA) and false of direction (0 of 1,458). Dogfooding it found a defect in shared code: append_partial_rows built its covered-set from partials[0].match_on and compared each row against its own, so a mixed-arity batch re-added rows every lap. Measurements in CIVIC_SURVEY. · from S84 (just-module-creator, 2026-08-31) · related RM130, RM134, RM150, RM48, RM153
  • RM139 — ✅ SHIPPED in 0.7 (2026-08-31: ReleaseRecord.unmeasured + the direction split in sweep.gate_findings). Filed 2026-08-30 by running RM126's gate for real at the 0.7.0 cut, and decided the day after. gate_findings read a module compiled on one side only as a compile failed; RM70 put an optional column on pharm_variants.csv, cyp2c9_warfarin_grch37 uses it, and 0.6.6 refuses that spec under extra="forbid" — so nothing failed and no before state exists to compare against. The two directions are facts about different releases: missing from AFTER is a regression in the release being cut and stays fatal (now carrying the compiler's own errors, which build_outputs had been discarding), while missing from BEFORE fails until the published record's new unmeasured list names it. Not the per-module escape hatch the original entry refused, because the check is an equality over the measured set: it cannot cover a module measured on both sides (reported as a note), cannot cover one this release broke, and as_record fills it from the measurement rather than leaving it to memory. The other three refusals stand. Re-measured end to end against 0.6.6 installed: 15 measured, 1 unmeasured, gate exit 0 where the cut had needed a human. · from the 0.7.0 cut · also in CHANGELOG, COMPILER, SCHEMAS
  • RM140 — ✅ SHIPPED in 0.7 (2026-08-31: StudyRow.statistical_test, plus an analysis-aware duplicate_study_citation). Filed and built the same day from a consumer's reproducibility benchmark: two agents, identical prompts, one overlapping row, two different p-values — and neither was a misreading, because the paper reports an allelic Fisher's exact test (OR 1.4, p 0.36) and a univariate logistic regression (OR 1.42, p 0.75) of the same variant. One run's row paired the second's effect_size with the first's p_value. Everything was green, quote verification included, because the quote grounds the significance verdict and contains no statistic to witness. study_design describes the study; nothing described the analysis, so a correct row and a mispaired one were byte-indistinguishable and no check could be written — the facts it would compare were unrecorded. One optional free-form column, no vocabulary and no gate (the reporter's own argument: the gate is unwritable before the column, and shipping both would make every published row retroactively incomplete). The one behaviour change is that both stated and different now suppresses the duplicate-citation warning — Kleene, not a != b, since an absent analysis is unknown and cannot establish distinctness. _KEY_FIELDS is deliberately not widened: it drives hints.key_fields and the published key.columns, and re-keying a shipped authored table is major-only. · from CONSUMER_SUGGESTIONS_HISTORY § S75 · also in CHANGELOG, COMPILER, SCHEMAS, PROPOSAL_0_7 (dated addendum)
  • RM141 — ✅ SHIPPED in 0.7 (2026-08-31: resolution.unresolved_subjects + the coverage check in the pre-flight). The third break of the validate/compile parity rule, hiding behind the rule's own exemption: what stays compile-only is a check reading resolved rows, and whether the injected table can place a row is arithmetic over bytes the pre-flight has already loaded. So a spec whose resolution.csv covered some of its variants passed validate --strict clean and was refused by compile --strict immediately after. The predicate is now shared rather than restated, and the strict error is the compile's verbatim — asserted by equality, since a pre-flight refusing in its own words still sends the author hunting. Two of the compile's distinctions kept: nobody-asked is not asked-and-absent, and --no-resolve silences it the way it silences the fill. A double-report was found doing it — the pre-flight and resolution both reach the finding, measured at 24 warnings for 12 subjects, now de-duplicated on the message. The reported mechanism did not reproduce: enrich gap-fills a short table, on 0.7 and on 0.6.6 both. The reporter then corrected their own account — the file was never truncated (203 sorted rows, clean final newline, the 62 absent rsIDs scattered rather than a tail), so it is a complete write of an incomplete resolution set, which is what the code does: an unaskable subject is written as no row at all. RM128's atomic write would not have prevented it, and the closure is this item by the right route — reading the table against the spec is indifferent to why a row is absent. Measured: no reference example moves a digest, signature or warning. · from CONSUMER_SUGGESTIONS_HISTORY § S76 · also in CHANGELOG, COMPILER
  • RM142 — ✅ SHIPPED in 0.7 (2026-08-31: merge_sources_file behind if covered: in clingen.py). A single-variant SIRT6 module: the dosage pass reported missing: [SIRT6], wrote no gene_metrics.csv row, and wrote a ClinGen licence row anyway. Two costs — a false statement in a published artifact, since licensing.csv travels to the registry and is read as this module uses this source; and it fires declared_license_disagrees, sending an author to adjudicate a conflict that does not exist (two agents measured doing exactly that). The compiler cannot catch it: _source_checks exempts the annotation layer from its orphan warning by design (RM46), so only the pass knows whether it contributed. The fix is the rule the rest of the family already follows — gene_metrics/frequencies/assertions/gene_validity all derive the source set from the rows they wrote, and were checked: neither sibling writes a licensing file when it covers nothing. Keyed on covered, not out (history) and not not missing (which would drop a real obligation from any module with one uncurated gene beside a curated one) — both directions and the second lap have tests. · from CONSUMER_SUGGESTIONS_HISTORY § S77 · also in CHANGELOG, ENRICHER
  • RM143 — ✅ SHIPPED in 0.7 (2026-08-31: build_disagreement_error, wired into both validate_spec and compile_module). A GRCh37 coordinate pasted into a GRCh38 module: enrich --strict refuses with a good diagnosis, enrich best-effort reports it and writes the table, and compile --strict then built the artifact silently — internally consistent and about a locus 5.6 Mb away. Two of the reporter's three asks were already shipped and the reply says so: their (2) (re-run the check in the compiler) is refutable on the data — resolution.csv holds one coordinate, not both, so there is nothing to compare — and their (3) (make the compile warn) is verification_findings_recorded, which shipped in this same release and is absent from the 0.6.6 they measured. What was missing was the last step of their (1): the record existed and no severity attached to it. The strict line does not move: strict still means reproducible, never right, and genome_build_agreement is the exception on internal-consistency grounds — the rows contradict the module's own declared genome_build. Every other recorded finding still only warns, pinned over four checks including reference_allele, which produces this diagnosis's input. Three non-behaviours have tests: no attestation is silent, findings=0 is a clean bill, and a skipped record is unknown (--offline writes one). · from CONSUMER_SUGGESTIONS_HISTORY § S78 · also in CHANGELOG, COMPILER, FAQ
  • RM144 — ✅ SHIPPED in 0.7 (2026-08-31: the count and the agreeing rows in _check_declared_license_agrees). The check filtered to the rows whose licence differs and rendered that remainder as the whole set, so a two-source module declaring the licence one row holds exactly printed sources report ['CC-BY-4.0'] — the agreeing row invisible in the sentence complaining about agreement. Two problems with different repairs read identically: your declaration is unsupported versus not universal, the second being the ordinary mixed-licence shape where the most restrictive term binds and the declaration is already right. Cost two agents a full re-adjudication that found nothing wrong, and survived RM142's fix. The count leads, the unsupported case gets its own sentence, and the denominator counts rows not distinct licences — a licence-less row and a non-annotation layer both stay outside it. Suppressing on any match was refused (the reporter's own argument: declaring the least restrictive of several is exactly what to warn about). declares license still leads, and the non-escalation is re-pinned. · from CONSUMER_SUGGESTIONS_HISTORY § S79 · also in CHANGELOG, COMPILER
  • RM145 — ✅ SHIPPED in 0.7 (2026-08-31: the standing in VariantRow.state's description). One of: risk, protective, neutral, significant, alt, ref — six peers, while derive.py calls alt/ref the retired descriptors and maps both to direction=unknown. A consumer passing our descriptions through verbatim offered an agent six equal choices and it picked alt for a heterozygote; the reporter had to read derive.py in their .venv to author one cell. Measured: 377 risk, 4 neutral, and zero uses of the other three across the sixteen examples. Three groups, not the two asked for — state is the P5 anti-pattern the charter names by hand, so the split is by which axis a value was on: significant is a significance claim stat_significance owns, not a dead value, and grouping it with alt/ref would say it means nothing. Each group names its successor, since a standing with no destination is a warning nobody can clear. Removal refused and not requested — major-only, and the effective_* aliases derive from these. · from CONSUMER_SUGGESTIONS_HISTORY § S80 · also in CHANGELOG, SCHEMAS
  • RM147 — ✅ SHIPPED in 0.7 (2026-08-31: LiteratureRow's docstring + a test; no behaviour changed). An agent read five literature services by hand and recorded it as five licensing.csv rows at layer=literature; the reporter removed them correctly on our own rules (RM46 — a literature source's terms are per article; RM142 — a pass contributing no row records no source) and then asked where the record of looking goes, since after removal there was no trace anywhere. The home already existed: an uncited literature.csv row is kept in the CSV and dropped from the artifact with literature_row_uncited (RM79), shipped for a deleted citation and the exact shape for the opposite case. It is structured, checked, cannot make a licence claim — which is what made the original rows wrong — and is about the paper, which is what was consulted; a service is only how the author reached it. Their logs/ reading was close and the typed row is better; a new layer member was refused on their own argument, and worse than they said since VALID_SOURCE_LAYERS is a wire vocabulary. · from CONSUMER_SUGGESTIONS_HISTORY § S82 · also in CHANGELOG, SCHEMAS
  • RM148 — ✅ SHIPPED in 0.7 (2026-08-31: the reading in VariantRow.direction's description). Two runs of one prompt, same model and paper, wrote risk/suggestive against unknown/not_significant for one variant on one body of evidence (p ≈ 0.073, OR 3.58, CI 0.96–13.4, 28.4% power) — both green, both defensible against a description that named the members and said only orthogonal to state. Not a vocabulary gap: direction is the sign of the estimate, stat_significance is how far to lean on it, and the orthogonality is the answer to whether a sign you cannot lean on is still a sign. The state they wanted a member for is already the pair direction=risk + stat_significance=not_significant, asserted by a test that authors their row. unknown is now bounded — no sign to record, never a sign you may not act on — because absorbing the second is what made the two runs equal. A new member refused as a second spelling of the pair: P5 overloading arriving as a synonym, and permanent under P3. · from CONSUMER_SUGGESTIONS_HISTORY § S83 · also in CHANGELOG, SCHEMAS
  • RM149 — 🔵 open — a first corpus is DRAFTED and unreviewed as of 2026-09-13, second pass same day (features/, 13 files, 228 scenarios, guarded by schema/tests/test_feature_corpus.py); a minor, release undecided; asked by the maintainer 2026-08-31. Express our described scenarios as Gherkin: freeform prose in expected-behaviour descriptions is producing ambiguities faster than it resolves them. Evidence is this repo's own week — S80 (six vocabulary members printed as peers, an agent picked a retired one), S83 (two runs of one prompt wrote risk and unknown for one variant, both green and both defensible against the description), S79 (a warning read as unsupported when it meant not universal). Three fixes to prose in a week, none a code defect. The gap is that a scenario exists as a docstring, a test name and a paragraph in COMPILER.md, which can drift from each other and from the code. Four questions block a start: which side is the source of truth (Gherkin from tests is documentation that cannot drift; tests from Gherkin makes every existing test a migration — opposite projects, same output), what is in scope of ~140 warning codes, where it lives and therefore whether it publishes under P3, and what a second dialect costs the next contributor. Two dated addenda of 2026-09-13 are that measurement: two of the four questions are answered by the ask (code → Gherkin, and the root rather than docs/, which keeps the P3 question open), scope is decided by three walked registries, and the first drafting pass produced four findings — a DRIFT between COMPILER.md's severity table and its own mishap matrix, the mode ladder being two mechanisms rather than one, plus RM242 and RM243, both filed and fixed. The body's ~140 warning codes is itself wrong: there are 73. What found things was deriving a scenario from an emission site, not reading prose harder. · from the maintainer, during the 2026-08-31 consumer round
  • ✅ RM151 — shipped 2026-08-31 in the uncut 0.7.0, the same day it was filed, as RM117's other half. An overrides.csv row answering a contested clin_sig is written about a PARTICULAR disagreement; if the archive later says something else, the reason describes a disagreement that is no longer on record. The baseline is the previous run's clin_sig_authority_calls.csv — it records each authority's clin_sig, verbatim clin_sig_raw and dataset — and it is the ONLY table in this format that keeps a prior value, so the finding NAMES its table and nothing promises this for frequencies.csv or resolution.csv (@probe-names-the-table). The ordering is the feature: the commit rewrites that file, so the comparison is computed in the staging phase and an AST guard pins the read above the write, demonstrated failing on a swapped source copy before it was kept. A move is observable exactly once, which is the honest shape for an observation — persisting it needs the overlay row BOUND to the value it justifies, an authored-surface change and the binding RM117's three objections all turned on missing. Three states, and withheld (no_prior_record / unchecked_now) never reads as unchanged. A move our own normalizer made is reported apart from the archive's (same verbatim raw token). Any overlay row naming the subject counts — the opposite rule from RM136's per-field overlay_answers, because this RAISES a finding where that one SILENCES one. Wording pinned by a word-boundary grep for adjudicating words. · from S52 (just-module-creator), via RM117 · also in CHANGELOG, ENRICHER, ROADMAP_HISTORY § RM117
  • RM150 — ✅ SHIPPED in 0.7 (2026-08-31: contested added to VALID_DIRECTIONS). Filed to ROADMAP_0_8 the same morning and taken into 0.7 by the maintainer — no sense postponing this — with its shape unchanged. The residue RM148 did not take: RM148 removed one of S83's three shades by reassignment (an unestablished sign is still a sign, so it is direction=<sign> + stat_significance=not_significant), and that reasoning holds. It did not reach the other two, and RM148's own description said unknown meant "not assessed, or the sources conflict" without letting a consumer tell which. Those are an absence and a finding, not one thing, and no pairing of direction with stat_significance can say two sources disagree about the sign — which is why the member is earned here where it was refused there. unknown keeps its original meaning: re-pointing a shipped member is a retype in all but name (P3), so this is an addition beside it. contested is the workspace's existing word (clin_sig_concordance.csv, clin_sig_concordance_contested). The trap was the first edit, not a follow-up: trimmed_state() reads _DIRECTION_TO_STATE with .get(direction, "neutral") — a default, not a raise — so trimmed_state("contested") already returned neutral before the member existed, exactly as a string that is not a direction does. Adding the member alone would have shipped a wrong legacy state with nothing failing. The map entry went in first and the guard is a registry-iterating equality over the walked set, because an output assertion passes on the unfixed code and proves nothing. _STATE_TO_DIRECTION deliberately gains nothing (no legacy state means contested) and stat_significance gains nothing (a disputed sign is not a disputed strength). · from S83 (just-module-creator) · also in CHANGELOG, ROADMAP_HISTORY, AGENT_NOTES

⏳ Open, no release decided — ROADMAP.md § Active items

Four items as of 2026-08-21 (RM103's manifest half, RM108 and RM110 here, plus RM117's observability half indexed under its own session below), and the heading was kept while it was empty. Three of the four — RM110, RM103's manifest half and RM108 — shipped on 2026-08-31 in the uncut 0.7.0; RM117 is the one still open, and the rows below carry their own status. The count in this sentence is dated on purpose: it says what the decision round left, not what is open now, and the ✅/⏳ on each row is what answers that. It held six until the 2026-08-21 decision round, whose whole content was that every one of them was a decision nobody had made: RM102 closed outright (→ ROADMAP_HISTORY), RM122 parked on demand (→ the deferral file), RM117 narrowed to its free half, and RM103 split, with the refusal moved to ROADMAP § The 1.0 cleanup. What is left is four minors whose shape is settled and whose release nobody has argued — a different thing from what this heading used to hold, and worth the distinction: two of the six were never design-blocked, one by a test that already pinned the answer and one by a record that held no incident. Items whose release nobody has argued yet belong here rather than in a ROADMAP_0_8/ROADMAP_1_0 file. Nothing is open under this heading as of 2026-09-11 — RM164 was the last, and its section moved to the 0.8 file the day someone asked which 150+ items were open without a release; the heading stays because the allocator writes a new reservation under it. And the section exists because there was nowhere to index one — every other heading names a release or a terminal state, so RM88 and RM89 would have been filed into a document this index could not point at, which is the exact failure this file was written to make impossible. Both have now left it: RM89 the next day, closed by the consumer answer it was waiting on, and RM88 in 0.6.1 once the policy it was really blocked on was decided. Both stay indexed under their release below, because an item that shipped is exactly what this index exists to still be able to find — and the heading stays for the next item filed before its release is argued.

Five of the rows below arrived at once, out of a batch of six: RM163–RM168, the 2026-09-01 source-adoption round. RM168 shipped the same day and is indexed under the 0.7 build round above. They answer one question — what else should we adopt as enrichment sources, besides CIViC and PubMind — and they were unusual for this heading in that none of them had been probed when they were filed. Each stated its measured half (a table kind with no provider, a snapshot's actual contents, a number from a probe already run) separately from its candidate half, and none asserted a licence.

All six were probed the same day, and decided with the maintainer that evening — the round is PROPOSAL_0_7_PT2, which is a record rather than a plan and wins over the entries below. Five build inside the uncut 0.7.0 and RM164 parks on a measured negative, having spun off RM171, which shipped two days later. The keeper from the round: five of the six entries said something their own probe contradicted, and four of the six verdicts the proposal drafted were then overturned in the maintainer pass — so an unprobed entry is a question, and a probed one is still only a proposal.

  • ✅ RM199 — shipped 2026-09-10, owner enricher, severity medium. release.json is now the last thing a publish sends, on every path: it tells a puller which release it holds, so landing it before the bytes leaves a snapshot that reads as provisioned and is not. The ordering lives in SNAPSHOT_ROOT_FILENAMES so a --dry-run prints it in send order. Above 5 GB the payload goes through upload_large_folder — an atomicity choice, not a speed one: upload_folder is one commit with no resumption, fine at megabytes and wrong at 32 GB. A declared retirement on that path is refused, because RM186 promises one commit for the arrival and the departure and the large uploader cannot give one. Also fixed the payload-only branch of plan_reference_snapshot, which still had its own hardcoded root-file pair. · from RM198 · related RM186, RM191, RM198
  • ✅ RM198 — shipped 2026-09-10, owner enricher, severity medium. The AVI lane wired into cache status/pull/prepare/upload once RM195 made it publishable — pullable without being buildable, since this tier may not fetch the gated 88.5 GB source but may serve the 34 GB re-encoding. Wiring it found the publisher's root-file list was a hardcoded pair, so avi_knots.parquet — a root-level sibling of data/, and the only way to reconstruct the PHRED the artifact does not store — would have been dropped silently, leaving a snapshot that looks complete and cannot rank a score. Third instance of @publisher-allowlist-derived; the names are now locations.SNAPSHOT_ROOT_FILENAMES and the publisher walks them. · from RM195 · related RM191, RM195
  • ✅ RM197 — shipped 2026-09-10, owner maintainer, severity low. Wide by position: chrom, pos, ref, alt0, alt1, alt2 and no stored alt — 3.371 B/row against 3.882, 29.7 GB rather than 34.2, taken because RM198 publishes the lane and transfer size binds where disk did not. Which base each column means is {A,C,G,T} − ref ascending, so nothing travels beside the data — proved over all 8,812,917,339 rows (rows == 3 × 2,937,639,113 positions, sorted, strictly ascending ALTs, none equal to ref; zero violations, 178 s after two OOM-killed attempts using group_by/n_unique). Not a schema break: to_long() recovers long rows and the RM193 join is unchanged. The builder refuses a locus that breaks the property rather than padding a null. One-column-per-base measured worse (3.677). · from RM191 · related RM191, RM193, RM198
  • ✅ RM196 — shipped 2026-09-10, owner maintainer, severity medium. pip install just-dna-enricher[atlas] installed two packages and could not import the client: the protos lived in docs/ and the bindings were git-ignored, so neither reached a wheel. Now the repository carries the pin, not the copy — a commit id plus a sha256 per file, fetched by atlas_protos.fetch_protos() and generated by enricher/hatch_build.py at build time, git-ignored and deliberately not build-ignored so sdist and wheel carry both. Backends are per package, so only the enricher moved to hatchling. hatch-protobuf cannot do it — measured: no import rewriting, and that rewrite is what stops the bindings shadowing the real alphagenome wheel. Verified from a clean venv with grpcio+protobuf alone. · from RM192 · related RM192
  • ✅ RM195 — resolved 2026-09-10, owner maintainer, severity medium. The Additional Terms define a Permissive Use class and grant it commercial use, then delegate membership to a sign-in-gated page that serves navigation chrome to curl; nothing in docs/vendor/ named an artifact, so AVI shipped commercial_use=None — warn, never gate. The maintainer saved the page: it carries its content as embedded JSON, and classifies AVI SNV scores as Permissive, for commercial and non-commercial use, with merged splicing and feature importance non-commercial only. commercial_use=True, pinned as alphagenome_download_page.html.gz + extraction, and the test asserts against those bytes rather than a constant. redistribution stays None — the page classifies use, and whether an HF publish is an 'open source release' under prohibition 1 is a legal reading nothing settles. · from PROPOSAL_0_7_PT4 · related RM191
  • ✅ RM194 — shipped 2026-09-11 in the uncut 0.7.0 as alphagenome expression, a recording pass, with both span forms and the interval winning, owner enricher, severity low. Its blocker is gone and was misdiagnosed: the interval RPC failed because Interval.strand has no zero member — the proto3 default STRAND_UNSPECIFIED is rejected as a bare INVALID_ARGUMENT — not because of the field mask (optional, measured) or 32 bp chunking (a parallelism strategy, not a requirement). score_interval shipped with RM192. Also found: the server returns a next_page_token on an exactly-full final page and following it 400s (AIP-158 violated; 1,000 bp is fine, 1,024 bp is not). Still open because a gene costs ~50 min to query end-to-end and because which source supplies the gene span is a design decision. Distal scores run ~10× lower, so distance must be recorded beside the score, and the lane needs the second source name alphagenome_atlas. · from PROPOSAL_0_7_PT4 · related RM192
  • ✅ RM193 — shipped 2026-09-10 in the uncut 0.7.0, owner enricher, severity medium. alphagenome check + the variant_impact_agreement member. Mostly offline by design: without a threshold there is no question the local snapshot cannot answer, and with one the knot table names the candidate set from 466 KB before any request. Refuses an unbounded refinement — the whole column is 272 days and ~92 M RPCs — with the refusal costing zero calls. Four no-answer reasons kept pairwise apart, none of them a zero, and a transport failure deliberately produces no finding. Emits its own member rather than a second reference_allele (@one-registrys-outage-may-not-speak-for-another). Caught a silent-wrong-answer bug: VariantRow stores 22, the artifact ships chr22, and the join matched nothing while looking like an uncovered region. · from PROPOSAL_0_7_PT4 · related RM192, RM191
  • ✅ RM192 — shipped 2026-09-10 in the uncut 0.7.0, owner enricher, severity medium. uv add alphagenome costs 550 MB and 47 packages (measured 2026-09-11; the entry shipped quoting 255 MB and 81, neither of which had been run) against a tier whose whole runtime list is httpx/tenacity/huggingface-hub, and six of the twenty declared deps are never imported on any scoring path. The .proto sources are Apache-2.0, so grpcio + protobuf reach every Atlas RPC — measured at 19 MB and +2 packages in a clean venv, with scores decoding through struct.unpack from the standard library. Shipped as atlas_client.py + atlas_protos.py + just-dna-enricher atlas generate under a new [atlas] extra, the alphagenome extra deleted, and 25 tests inside testpaths (all green, live leg included). An AST walk pins the import floor to {grpc, just_dna_enricher}, which is what makes the size claim a property of the code. protobuf is runtime; grpcio-tools is the build-only one. · from PROPOSAL_0_7_PT4 · related RM191, RM193, RM194, RM196
  • ✅ RM191 — shipped 2026-09-10 in the uncut 0.7.0, owner enricher, severity medium. AlphaGenome's AVI artifact as an operator-built cache lane (alphagenome build --input, no default URL — the source is behind an eligibility gate). raw_score stored as Int32×10⁵, checked lossless on every row rather than sampled; PHRED not stored, since it is an exact within-corpus rank and a 466 KB, 41,474-row knot table carries the curve, the per-value ambiguity interval and decidable threshold safety instead — genome-wide, exactly one knot straddles any integer threshold 1–50 (0.00076, 676,356 rows, PHRED 2.99961–3.00027). Losslessness is about the decimal: raw_score_e5 / 1e5 disagrees with the printed value on 53% of rows. Its size finding was wrong first and corrected: 43.0 GB was the builder's own sink_parquet fragmentation, and the artifact is 34.2 GB — see RM197. · from PROPOSAL_0_7_PT4 · related RM192, RM193, RM195, RM197 · also in probes/ALPHAGENOME_ATLAS
  • ✖ RM189, RM190 — reserved and released without an item being written, both by the allocator itself: .claude/rm-next.py had no --help and no unknown-flag branch, so a typed flag fell through to the reserve path and claimed a number. The numbers are spent rather than free — ids are never reused, so a later reference to RM189 or RM190 can only mean this — and the tool now refuses an unrecognized flag instead of allocating on it.
  • ✅ RM186 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; LayoutShift + cache prune + the per-lane glob registry — no schema change). The maintainer corrected a premise of this repo's own: @snapshot-accumulates had been read as never delete, but a HF dataset repo is git-backed, so a delete is a commit and a superseded revision still resolves — three were read off the hub while auditing. The risk was never lost bytes; it is that a retired file goes on answering 200 (CLINPGX_ARCHIVES) and that a sweep removes what nobody looked at. So exactly two routes, neither a side effect of a publish: a declared retirement, where the change that retires a file carries the migration — new absent and old present over the remote, so it fires once per repo and rides the upload's own commit as delete_patterns — and cache prune, which names only a data/ file the lane's glob excludes or a declared retirement, never README.md/.gitattributes/release.json/LICENSE.txt/sidecars, and without --yes prints and stops. The one declared entry is already past its own condition (ClinVar carries both spellings, so prune owns the 159 MB remnant) and stays literal: loosening it would make every publish a prune. Second reader of the per-lane globs made them SNAPSHOT_FILE_GLOBS, walked by test with STRchive the one enumerated exclusion — no data/, so n/a rather than clean. Measured read-only against all nine repos; nothing deleted, publishing stays the maintainer's. · from the 2026-09-03 published-artifact audit · related RM185, RM178 · also in CHANGELOG, AGENT_NOTES @a-publish-may-not-orphan-the-bytes-it-stops-describing
  • ✅ RM187 — ✅ shipped 2026-09-03 in the uncut 0.7.0, owner enricher, severity high (just-dna-enricher only; no schema, no parquet column, no CLI flag). A real failure, not an audit: NCBI closed the connection 180,927,542 bytes into a 193,427,450-byte ClinVar VCF during RM179's republish, and download_clinvar_vcf raised httpx.RemoteProtocolError while _rebuild_clinvar catches ClinVarBuildError — so the lane could not report built=False and the traceback escaped rebuild_lane, which in a full cache rebuild aborts every lane after the flaky one. Four of eleven builder downloads leaked (clinvar_build ×2, constraint_build, clinpgx_build); all eleven had no retry, on the largest requests the tier makes, while every live client has had attempt_floor since RM42. The bodies were fifteen identical lines copied eleven times, so the repair is one net.stream_to_file rather than four patches — eleven public signatures unchanged, each downloader building its own return type from a StreamedFile. Atomic, translated, retried on TransportError only (a 404 is the same 404 four times over), and restarted from zero per attempt because an appending retry yields a real digest over nonsense. constraint_build had no error type at all, which is why its download had nothing to translate into; the three new *Unavailable types are subclasses so a caller catching the build error still catches them. Third appearance of @client-exception-contract after RM97 (clients) and RM101 (passes), and the guard walks the package by AST, twice — no download_* may open a stream, and httpx.stream appears nowhere outside net.py — because RM101's own guard hand-kept eight module names and missed identifiers. Two defects the sweep found: pubmind_build was the one handler of eleven leaving its .part behind, and four computed a sha256 only to log it. Not done: no resume, so a retry re-fetches from byte zero — Accept-Ranges support is unmeasured and a resumed digest is a different design. · from the 2026-09-03 republish incident · related RM179, RM97, RM101, RM42 · also in CHANGELOG, ENRICHER, AGENT_NOTES
  • ✅ RM185 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; one publisher guard wired into three publish paths — no schema change). The general form of RM179, decided with the maintainer: a snapshot's release.json describes every half the artifact carries, so a publish carrying one half replaces the whole description while the publisher's add-never-delete leaves the other half as bytes nothing describes. OrphanedSidecarError refuses it, its own type for the reason PublishCollisionError is one. It reads the remote tree, not the remote release.json — by the second bad publish the block was already gone while the sidecar was still there, so a guard interrogating the description would have waved it through exactly as the first one went through; the bytes are what a puller gets. Scoped to publishes carrying release.json (one carrying none overwrites no provenance) and passing on a repo nobody has published to (a first publish is not an orphan). The dry run runs it too, since a rehearsal that skips what the real thing refuses on is @publisher-allowlist-derived's defect again. The repo it was written against was republished the same day, so the guard is verified against both the found state and the repaired one. · from RM179's deferred half · related RM179, RM186 · also in CHANGELOG, AGENT_NOTES @a-publish-may-not-orphan-the-bytes-it-stops-describing
  • ✅ RM182 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; one field on CacheLane and the reporter that reads it — no schema change). cache status labelled a lane with release.json's dataset; eleven builders write it and clinvar_build writes clinvar_file_date instead, so the one lane that refreshes weekly was the only one printing a blank release while every slower lane named its own. release_label joins build_command as a field for the reason that one exists — a convention holding for eleven of twelve is not a convention, and the reporter composing one is what produced the blank. ClinVar's is clinvar_dataset_label, the function clinvar_draft writes and clinical.tautology_reason recomputes, shared rather than mirrored; the override set is asserted as an equality over the registry. Refused: adding dataset to the builder's release.json, which repairs nothing already on disk and makes a second writer of a label that function owns. · from the 2026-09-03 published-artifact audit · related RM176, RM179 · also in CHANGELOG, AGENT_NOTES @a-lane-with-two-halves-publishes-the-provenance-of-one
  • ✅ RM179 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; one rebuild adapter — no schema change, no CLI flag). A ClinVar snapshot is two halves on two cadences — the VCF and var_citations.txt — which is why build_citations merges a citations block into release.json rather than writing its own file. _rebuild_clinvar built the VCF half only, and release.json is written by that half, so every cache rebuild clinvar --publish uploaded a citations-free provenance over a repo whose sidecar it had not replaced: the published snapshot on 2026-09-03 held records from 2026-08-29 beside citations from 2026-06-27 and said nothing, the block having been there at revision 8f5c5720 and gone after the 2026-09-02 publish. The publisher adds and never deletes, so the sidecar outlived its own description. Now both halves or neither: a failed citations half returns built=False with no out_dir, and --publish uploads only on built is True. Costs a fully offline --source rebuild of this lane, stated rather than hidden. Refused: a second --source grammar for one lane's second input. Not repaired here — the publish-side guard (refuse to overwrite a remote release.json describing sidecars the plan does not carry) is a policy call, and the live repo still needs a citations rebuild, which is outbound. · from the 2026-09-03 published-artifact audit · related RM176, RM178 · also in CHANGELOG, AGENT_NOTES @the-reporter-cannot-compose-a-lanes-label
  • ✅ RM178 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; one download loop, one normalization, one archive reader — no schema change). HfFileSystem.get opens the destination before it resolves the remote path, so _provision_snapshot's optional-file try left a 0-byte LICENSE.txt in every cache pulled from a repo that publishes none — four of the nine published snapshots — and a re-pull truncated a good local copy from 36 bytes to 0. Both verified against the live hub. An empty licence is not a smaller absence: absence withholds (license_sha256 null, with a warning) while an empty file pins sha256:e3b0c442…b855, the hash of the empty string. Fixed at both ends — the fetch stages through .part, and blank normalizes to None at the sink (SourceTerms.row) rather than in four callers. Refused: deleting phantom files from an operator's cache, for the same reason a foreign parquet is reported and not removed. · from the 2026-09-03 published-artifact audit · related RM176, S44 · also in CHANGELOG, AGENT_NOTES @a-failed-fetch-is-not-a-no-op
  • ✅ RM177 — shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; fourteen --out defaults and locations.repro_out, no schema change). Nothing a command generates goes in the repository root had been the rule since civic reproduce was corrected for it, and it was prose: nine builders (civic, clinvar, pubmind, gnomad_constraint, mane, strchive, acmg_sf, mitomap, mitomap_miss) defaulted --out to a bare relative name that landed beside pyproject.toml, and four (clinpgx build, clinpgx build-labels, cpic build, pharmvar build) required --out with no default, which is how the docs came to say --out ./clinpgx. Every default is now repro_out("<lane>") → data/repro/<lane>/, cache rebuild keeps data/caches/ as a named constant, and an AST walk over cli.py refuses any --out default written as a literal — sixteen found, one input-shaped exemption (clinvar citations) enumerated as an equality. civic reproduce moved to data/repro/civic_reproduce. Refused: a .gitignore line per lane, the repair that produced the nine repeats. · from the ENRICHER reference's own --out ./clinpgx example · related RM176 · also in CHANGELOG, ENRICHER, AGENT_NOTES @a-default-spelled-per-command-is-a-rule-in-prose
  • ✅ RM174 — the stamp half shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher, two parquet columns + one release.json field; no model, no authored surface). CIViC EID 8721 is molecular profile 5278 (VHL S183L AND VHL D126N) and the builder stamped molecular_profile_id from the variant's single-variant profile, so the parquet said twice that a two-variant claim was a single-variant one. evidence_molecular_profile_id / evidence_molecular_profile_name now ride beside the join key on every row — the key must stay the variant's own profile or the row does not join — and a composite is the inequality of the two ids, derived not stored, counted in release.json as composite_profile_rows. The name is null on a TSV row because that file publishes none. The shape underneath is not repaired and is not repairable here: 8721's own description says heterozygous compound mutation, so the two variants are in trans and HaplotypeRow is cis — there is no brick for that claim, and it is RM28's, parked. This item is where RM28's first counted corpus entry came from (1,964 CIViC profiles, 209 multi-variant — AND 141 / OR 72 / NOT 1, nested). Repair 1 refused on the record: it drops the row like the TSV path and deletes two of RM170's three subjects. · from the RM170 probe · also in CHANGELOG, ENRICHER, PROPOSAL_0_7_PT3
  • ✖ RM172 — reserved and released without an item being written. The number is spent rather than free: ids are never reused, so a later reference to RM172 can only mean this.
  • ✅ RM171 — ✅ SHIPPED 2026-09-03 in the uncut 0.7.0 (just-dna-enricher only; two cache lanes, a parents field on CacheLane, a draft source, a SourceTerms row, five abbreviations in the shared normalizer — nothing removed, promoted or retyped). The entry's binary was the wrong question: sixteen new expert-panel calls, or does ClinVar already have this treats "16" as a fact about MITOMAP when it is a fact about one join against one ClinVar vintage. What shipped derives the increment on every rebuild instead. Two parents, one derived child — mitomap (the published pg_dump, six tables kept, both variant tables because the repo's only mtDNA module draws from rtmutation and neither of its variants is in mmutation) and mitomap_miss, whose acquire stage is both parents on disk, whose build is an exact (start, ref, alt) join on chrMT with no position-level fallback, and whose release.json pins both parents so a ClinVar rebuild without a child rebuild is detectable. A parent that is absent is built=None naming it — never False (another lane's absence filed as this one breaking) and never an empty miss (the strongest possible claim, from a comparison that never ran). Four buckets, one drafts: photocopy / rated miss / unrated miss, plus unmintable as a fourth because the question cannot be asked of a right-anchored : deletion without an rCRS base Principle 2 forbids fetching. The bracket is a normalization and the confirmation token never is — MITOMAP's legend says in as many words that Reported/Cfrm is not an assignment of pathogenicity — so the five VCEP abbreviations became keys in the one shared normalizer. [VUS*] is withheld and the withhold had to be upstream: normalize_clin_sig's own default is other, a definite member, so an unmapped token falling through would have become a confident call. The rejoin moved the number: against the 2026-06-27 ClinVar (3,104 chrMT alleles) the rated miss is 6, not 16 — all 16 reproduce, but 13 are : deletions the design itself puts in the unmintable count, and all 6 survivors key on an indel neither side left-aligns, which the lane publishes rather than hides. nlmid walked whole: 6,372 PMIDs / 397 empty / 1 Ovid id whose first eight characters are digits, so the predicate requires the whole cell. genotype is stubbed and the reason is not the contig's: sole_expressible_genotype fills ALT on chrMT for ClinVar, whose record is a claim about an allele, while MITOMAP's is a claim about a literature corpus — three of the six drafted rows are reported only heteroplasmically — so a MITOMAP-drafted module cannot compile until a human writes those cells. Five departures from the build plan, each in the ROADMAP_HISTORY entry and in a dated addendum on PROPOSAL_0_7_PT3. · from RM164's probe, 2026-09-01 · related RM164, RM176 · designed in rm171_diff_strategy · also in CHANGELOG, ENRICHER
  • ✅ RM162 — shipped 2026-09-01 in the uncut 0.7.0 (tooling only; no package, no schema change). Sn has had an allocator since the triage loop was built because the id is written into a document; RMn had none, so the number was read by grepping "the highest in use" and the gap between reading and writing is where another session reads. The incident is reproduced, not hypothesised: on 2026-09-01 two sessions in this working tree filed different work as RM159 a minute apart (git 741ec59, which renumbered one to RM161). .claude/rm-next.py scans every docs/**/*.md and reserves the next number in the same locked write — reading the maximum outside the lock and appending inside it is the same race with a smaller window. The lock is on docs/, never on RM_TOC.md: a lockfile left by the kill it guards against would block every later run (@flock-not-a-lockfile), and flock binds an inode, so a rename-over leaves the holder locking an unlinked file while a second process acquires immediately — measured in a sandbox before the tool was written. A reservation is a visible 🔷 index row, not a side-car the index cannot see; --release leaves a ✖ tombstone because ids are never reused — the first cut deleted the row, the number went invisible to the scan, and a released RM10 was handed straight back out. Pinned by a test that runs eight allocators at once and runs the same eight with flock neutered to watch them collide. · from the 2026-09-01 collision · also in CHANGELOG, CONSUMER_TRIAGE_LOOP, AGENT_NOTES @an-index-is-not-an-allocator
  • ✅ RM103 — the manifest half shipped 2026-08-31 in the uncut 0.7.0 (just-dna-format + just-dna-compiler); the refusal half stays on ROADMAP § The 1.0 cleanup. ModuleInfo(version="abc").version is "0.0.0": normalize_version strips every non-digit, finds none, pads to three zeros — deliberate since RM17, but 0.0.0 is a legal SemVer and a plausible pre-release, so an unreadable string reached manifest.identity.version as a confident claim. Fixed additively: Identity.version_coerced_from publishes the authored string beside the coerced one, None when nothing was rewritten. The RM17 coercion is untouched and the test parametrizes every digit-bearing case so a later change cannot undo it while looking like this item. No sentinel exists — every three-number string is somebody's real version. The build turned up a second-order defect and fixed it in the same commit: reverse_module took its version from the caller, who holds the coerced string, so lap 2 had nothing to coerce and the new field went absent — RM137's exact shape, in the release that files RM137. Reverse now recovers version_coerced_from from the artifact's own manifest and re-emits the pre-coercion string, asserted as a fixed point over two laps and demonstrated failing on the naive version first. That also repairs a quieter loss: reverse used to drop module.version entirely unless a caller passed one. Note the compiler has always warned naming both values, in compile and validate alike — the silence was the model's and the artifact's. · from S42 (just-dna-lite) · also in SCHEMAS, CHANGELOG RM104–RM111 — the 2026-08-19 doc-audit batch. Eight code findings turned up by validating just-module-creator's 24 table references against this repo's code, all re-checked against the tree at format/compiler 0.6.1 + enricher 0.6.4. Six shipped on 2026-08-20 and are indexed in their own section below; the two left here needed a decision first, and both were decided on 2026-08-21. The doc half of that audit is tracked in the interim handoffs, not here.

  • ✅ RM160 — the provenance half shipped 2026-09-03 in the uncut 0.7.0 (just-dna-enricher client + command + check, two optional studies.csv columns, one VALID_VERIFICATION_CHECKS member). civic build reads dated files; the wider basis RM169 adopted is a VCF, and a VCF record needs a POS — so submitted evidence on a variant with no GRCh37 coordinate is published on the GraphQL API and nowhere else. Ten records' citations were unreachable, including variant 1955, whose only reachable evidence for the numbering convention its identity turns on is EID 9969 (PMID 12202531, Dollfus 2002, free full text), SUBMITTED. Its coverage half shipped as RM169 — the wider corpus was a dated, pinnable VCF all along. Built as shape 3, decided 2026-09-02: read SUBMITTED at enrich time, so civic build/civic reproduce keep their byte-reproducibility contract and the published snapshot does not grow. Shape 1 (hash a capture) was available and not taken; shape 2 was dissolved by RM169. civic citations <spec> drafts into studies.csv — never literature.csv, which is derived from those PMIDs and would be dropped uncited — and three routes reach a variant id because the first two miss the class: the snapshot join, the curated name-identity table, and --variant-id for a record neither can place (1955 is not in the identity table). status rides as confidence/confidence_unit, unconverted; the pin (fetched_at, dataset) is on the (civic, literature) SourceRow, because annotation is civic_draft's slot and a route is not a source. evidence_status_currency re-asks from enrich and reports a status accepted or rejected since, or a citation added since — warns in both modes, escalates in neither, and stays apart from dataset_currency. Rejected evidence withheld and counted; a paper whose live items disagree has its confidence withheld; --offline records skipped/offline, never ran, findings=0. One authored column pair more than PROPOSAL_0_7_PT3 priced it at — corrected there as a dated addendum. · from the 2026-09-01 residue round · related RM169, RM152, RM159, RM153 · also in CHANGELOG, ENRICHER, SCHEMAS

  • ✅ RM170 — shipped 2026-09-02 in the uncut 0.7.0 (just-dna-enricher, plus one VALID_VERIFICATION_CHECKS member; no model, no parquet column, no authored surface, and no new VALID_WARNING_CODES — that set is the compiler's, guarded to codes a compiler check builds, so the two finding codes stay this pass's own). Three VHL variants carry a claim and a published refutation; the snapshot kept both rows and nothing downstream said so, so a risk row could be authored over one with every gate green. Not contested: that counts risk vs protective, and a Does Not Support row enters no camp (CIVIC_DIRECTION_MAP → None), so contested_variants is correctly 0 on every basis and blind to this. Two surfaces: draft-panel --source civic names the variants it wrote a direction for that the snapshot also rebuts (before, it only counted the refuting rows it withheld — a different fact), and enrich folds in published_refutation whenever a CIViC snapshot resolves, so a hand-author meets the same sign. Two finding codes, because two sentences: refutation_beside_claim and refutation_without_claim, carried in the record's detail and restated at compile as verification_findings_recorded. Warns in both modes, escalates in neither, repairs nothing — a refutation withholds a claim rather than establishing its opposite. The finding keys on the refuting evidence id and fans out, so EID 8721 (the combination genotype VHL S183L AND VHL D126N, written as two single-variant rows) is one refutation over two subjects whichever way RM174 goes. The record states its basis on every run including the empty one: every assert-and-refute pair in CIViC rests on submitted content, so on the accepted basis the class is empty by construction and a bare findings: 0 would read as clear water. comparison_plan gained an authored selector rather than a second copy. · from the RM169 widening · design record probes/rm170_kleene.md · measurements probes/CONTRADICTION_CORPORA.md · also in CHANGELOG, ENRICHER
  • ✅ RM169 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher; no schema change). RM160 held that reading CIViC's unreviewed majority meant reading the API and forfeiting byte-reproducibility. False, and the check was one HTTP request: CIViC publishes <date>-civic_accepted_and_submitted.vcf in the same dated release directory the builder already reads. civic build --submitted / civic reproduce --submitted take it: 507 rows on 270 variants → 1,149 on 397, refget cross-check 57 → 129 coordinates, 0 mismatches, rebuild still byte-identical. Every row gains evidence_status (CIViC's own word, unconverted) and release.json gains status_basis/status_counts/vcf_evidence/unjoinable_submitted. The VCF is not the input: a VCF record needs a POS, so it drops every variant with no GRCh37 coordinate — 52 of the 54 it drops are the unresolvable_identity class RM159 resolved by name. TSVs stay primary. VariantSummaries.tsv is accepted-only too, so 112 of the 127 new variants have no row there (the other 15 join the TSV normally); their identity is read from the same CSQ entry through the same parsers (57 CAID · 40 rsID · 14 GRCh38 accession · 1 both) and stamped identity_derivation="vcf_csq" — the member names the file, not the route. Nothing is placed from the VCF's GRCh37 POS (RM48 stands). Also enumerated the whole release surface: 7 TSVs + 2 VCFs, and GeneSummaries.tsv is byte-identical to FeatureSummaries.tsv. · from RM160 · related RM160 (its provenance half stays open), RM159, RM48, RM152
  • ✅ RM159 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher; no schema change). civic build placed a row only from CIViC's identifier columns and dropped 53 variants as unresolvable_identity; for most of them the identity was published all along in the variant's own name (N150fs (c.448delA), IVS2+1G>A, D1709N). 33 of the 34 that resolved are adopted as the shipped constant civic_identities.CIVIC_NAME_IDENTITIES, emitted as identity_derivation="curated_name" — coverage 237/290 → 270/290 variants (81.7% → 93.1%), rows 474/533 → 507/533, unresolvable_identity 59 → 26. Shipped as data rather than as a draft-time lookup because 4 of the 33 needed a judgement no lookup makes (788's legacy IVS2 converts structurally to the wrong exon and both readings are real registered alleles 9 kb apart; 2459 pairs a missense label with a synonymous cDNA change; 804 names a consequence over an intronic allele; 2196's rs-number is position-level) — so the build stays offline and byte-reproducible and the procedure lives in CIVIC_IDENTITY_PROTOCOL. Excluded: TP53 R72P, whose identity is the reference allele (g.7676154G=, the name has ref and alt inverted) and ref == alt is not a variant row. Each row is keyed to the exact name it was read from and lands in one of four counted states in release.json (applied/superseded/renamed/absent, summing to the table), so a curated answer cannot outlive the record it answered and a supersession is a free currency signal. civic reproduce reads 57 of 57 coordinates against the GRCh38 reference via refget, 0 mismatches, up from 24. · from the 2026-09-01 residue round · related RM153, RM48, RM152
  • ✅ RM161 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-format). The pre-build gate run for the cut exited 1: gene_validity.superseded_count and identity.version_coerced_from "moved and the release record does not list it". Both are real manifest additions from the 2026-08-31 batch carrying a DeclaredChange written the day they landed, and neither reached manifest_fields — the declarations went in at 06:52 and 07:04, after the measurement the list came from. The shape is the record's own construction: as_record writes the measured half with declared empty on purpose, which is what makes the gate work and also what lets a later item leave the measured list behind; and the gate needs the previous release installed, so a checkout sees nothing. Guarded by an asymmetry: a declared addition must be listed (a field that did not exist before moves wherever its block appears), while a declared correction may be unmeasurable on the corpus — gene_validity.classifications and gene_metrics.signature stay declared and unlisted, and the gate already notes the reverse case. Test walks RELEASE_RECORDS and fails on the pre-fix tree naming exactly the two. Evidence sentence unchanged; after the fix the gate reads "release record for 0.7.0 covers the measurement", exit 0. · from the 0.7.0 pre-build gates · also in CHANGELOG, COMPILER, AGENT_NOTES @a-record-written-in-two-passes-drifts-between-them
  • ✅ RM158 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher). gwas._module_subjects built its (rsid, variant_key) list from variants.csv while five authored models carry rsid, so a module whose rsIDs live in haplotypes.csv or pharm_variants.csv got no associations and no line saying none had been asked for — a haplotypes.csv row for CYP2C19*2 returned []. The third instance in one sweep, and the one where the fix was already written: enrich.Subject and its collector exist for this exact question (RM43 — resolution read variants.csv alone until a PGx module enriched to an empty resolution.csv), and this pass, written afterwards, restated the narrow loop instead of calling it. _collect_subjects/_Subject are now public, since a private name is what kept the second caller from finding the first. studies.csv is deliberately not a subject — a study row references a variant the module already carries. Corpus measured before and after: pathogenic_clinvar 301, hboc_palb2 16, mt_heteroplasmy 2, grch37_build 0, identical lists. No column, no vocabulary member, no signature moves. · from the RM155 sweep · also in CHANGELOG, ENRICHER, AGENT_NOTES @roster-is-as-wide-as-the-tables-it-reads
  • ✅ RM157 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher). gene_metrics.module_genes built its list from variants.csv alone while nine authored models declare gene. It is not a report but the scope of three passes — constraint metrics, gene validity, ClinGen dosage — so a module whose genes live in its PGx tables had all three quietly do nothing: no rows, no findings, no line saying a question had not been put. Measured on this repo's own corpus: cyp2c19_star_alleles, apoe_epsilon, cyp2c9_warfarin_grch37 and hfe_compound_het returned [] while naming six real symbols. The workspace already held two answers to one question (pgx._module_genes reads two PGx tables), and RM104 had patched the symptom — an UnboundLocalError on "any module with no variants.csv" — without asking why the list was empty. Now derived from the same registry walk as the identifier roster, and refusing on an unparseable table rather than narrowing (IdentifierRoster.read_errors carries the loader's message so gene_validity's diagnosis is unchanged). pgx._GENE_TABLES stays two tables — it decides whether the star-allele cross-check applies, a different question. No schema change and no ordering change (rows sort by (gene, dataset)). · from the RM155 sweep · also in CHANGELOG, ENRICHER, AGENT_NOTES @roster-is-as-wide-as-the-tables-it-reads
  • ✅ RM156 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher). RM155 widened the identifier rosters to nine tables per column; two gates in front of them were still keyed on variants.csv alone, so on the module shape the widening was most for, the wide roster was never reached — check_identifiers(spec_dir=) raised variants.csv is invalid: ... not found, and the command returned "no variants.csv — nothing to check" one call earlier and hid it. Reproduced on this repo's own corpus: cyp2c19_star_alleles, apoe_epsilon, cyp2c9_warfarin_grch37 and hfe_compound_het carry no variants.csv and name six real gene symbols between them; all four exited 0 having asked nothing. The guard is now the roster rather than a filename (nothing to check = no id-bearing table read, and only then no attestation — the half of the old comment that was right); an absent variants.csv is no rows, a present-and-unparseable one still raises. A third vacuous pass beneath them: _gene_locus_conflicts returned compared=0 with no reason, the ran(0, 0) its own attestation forbids. No column, no vocabulary member, no signature moves. · from the RM155 sweep · also in CHANGELOG, ENRICHER, AGENT_NOTES @roster-is-as-wide-as-the-tables-it-reads
  • ✅ RM155 — shipped 2026-09-01 in the uncut 0.7.0 (just-dna-enricher). check_identifiers built its trait and gene rosters from variants.csv while eleven authored models declare one of those columns (StudyRow has carried trait_efo_id since 0.3), so a module carrying its ids in studies.csv reported nothing checked and nothing flagged — a retired CURIE could ship with every gate green. The unreadable 0 is the item: it said declares no trait and traits are in a table nobody read at once, which is @unreachable-not-absent at a finer grain, so widening alone would have left the hole. Both halves shipped — roster derived from DRAFTABLE (nine tables per column, asserted as an equality over the walked _ALL_MODELS, never a literal) plus *_tables_read/*_tables_not_read on the report and a CLI count that names its denominator. MeasureBinRow correctly absent (abstract base, four concrete subclasses each their own entry); the three derived models with these columns excluded on purpose — a stale id there is the source's currency, not the author's. A third instance one level up: report.clean is all() over a possibly-empty set, so the CLI printed a green pass having asked nothing. No column, no vocabulary member, no signature moves. · from S86 (just-module-creator) · also in CHANGELOG, ENRICHER, AGENT_NOTES @roster-is-as-wide-as-the-tables-it-reads
  • ✅ RM154 — shipped 2026-08-31 in the uncut 0.7.0 (all three packages). An rsID the source HAS, whose every locus the allele-aware filter rejected, was written as status: not_found — this source has no record of your rsID — byte-identically to a genuine absence. Reported over five subjects of a 64-variant module authored from an hg19 supplementary: the paper spells the submitted strand, so its G/A meets GRCh38's C/T. The fourth state in RM98's family (asked-and-failed / nobody-asked / no-position-and-silent-about-why / answered-and-rejected). Both obvious repairs refused: a new VALID_RESOLUTION_STATUS member is a wire change, and deleting the row moves resolution_signature — variant_key/rsid are fact fields, status is provenance and is not, measured rather than reasoned. So the row stays and only the reason moves: EnrichmentResult.allele_mismatches carries AlleleMismatch(rsid, genotype, loci, offered, strand_flip). A second defect in the quoted sentence: hosting_verdict's two False arms shared the event-length arm's reason, so a strand flip was diagnosed as "The event sizes differ" about two 1 bp substitutions — contradiction_reason is now undecided_reason's twin, with a test asserting the arms' reasons are pairwise distinct. No column, no vocabulary member, no signature moves. · from S85 (just-module-creator) · also in CHANGELOG, ENRICHER, AGENT_NOTES @answered-is-not-absent
  • ✅ RM108 — shipped 2026-08-31 in the uncut 0.7.0 (all three packages). ClinGen's assertion_id embeds the curation timestamp (CGGV:assertion_…-2019-08-18T160312.829Z), so a re-curated assertion arrives under a different id, misses _merge_key, and is appended beside the row it replaces — manifest.gene_validity.classifications then published ["definitive", "refuted"] with nothing saying which stood. Decided: newest classification_date is current, nothing is deleted (S45's answer on a weaker signal; the date orders and does nothing else). The entry's marker COLUMN did not survive the build, and the reason is the durable part: the row that must be marked is the one already in the file, and merge-not-clobber forbids the pass editing it — so the marker would be right on every run except the one that created the ambiguity. A boolean fails that way; superseded_by fails that way plus three more (GenCC rows may have no assertion_id to point at, a twice-superseded row needs an immediate-vs-current rule, and a pointer LOCATES rather than asserts). So nothing is stored — classify_currency derives it in the format tier for both callers (@derived-not-stored), no column changed, gene_validity.signature does not move, and no existing module recompiles anew. Grouping is (gene, disease_id, moi, submitter) — the grain minus dataset, since a re-curation is by definition a later release. Two edges withhold: a tie on the date, or any group member stating none, leaves no row current and none superseded (a tie-break on assertion_id was refused — an id carries no chronology). Warns in both modes in both tiers and escalates in neither; both codes carried. New: manifest.gene_validity.superseded_count, gated on a round-trip assertion. Two things the build turned up: it is the first fact-table check to run on both sides, so the message doubled and doubled warnings_summary with it — the fact-handler loop now dedupes on the message like every other both-sides check; and the existing merge test re-ran an identical export, which cannot see this defect at all, so a two-export fixture is part of the fix. · from the 2026-08-19 doc audit · also in CHANGELOG, SCHEMAS, AGENT_NOTES
  • ✅ RM110 — shipped 2026-08-31 in the uncut 0.7.0 (just-dna-format + just-dna-enricher). The snapshot leg kept gnomAD's JSON array literal ([] is not in _NULLS) while the live API pipe-joined. Re-probed against the published v4.1 parquet before any code moved, and worse than first filed: not one of the 18,111 rows is null or empty, so if row.constraint_flags: was true for 100% of them — 17,403 carry [] and the 708 genuinely flagged rows are literals too (14 distinct shapes, ["outlier_mis","outlier_syn"] among them), so splitting on | yielded one bogus token. Only 3.9% of genes are actually flagged. "Empty → null" was therefore only half the fix: the non-empty cells needed parsing. Decided as pipe-joined-when-non-empty / None-when-empty on both legs — which test_gnomad.py already pinned on the live producer, so nothing was undecided and the item needed a release, not a decision. The normalizer landed in FORMAT, not the enricher: gene_metrics.normalize_constraint_flags as a mode="before" field validator, because the published snapshot is immutable and every gene_metrics.csv already written from it (this repo's hboc_palb2 included) carries [] on disk — a producer-side fix leaves those contradicting the column and hashing apart from a live fetch. Three call sites share the one function: live route, lookup_snapshot (the published snapshot), constraint_build (future ones). Cost measured at one corpus row, whose gene_metrics.signature and artifact.digest move — hence a minor with a CHANGELOG line. · from the 2026-08-19 doc audit · also in CHANGELOG, AGENT_NOTES @one-normalizer-two-spellings

⏳ Deferred to a later minor — ROADMAP_0_8.md

  • RM164 — 🟡 DECIDED 2026-09-01 — PARKS to 0.8 on a measured negative, and stays open; the section moved into the 0.8 file on 2026-09-11, ten days after the decision that sent it there. Reopen it with a source, never with an argument; the mmutation spin-off is RM171. Measured over _TABLE_KINDS and the enricher's providers: all nine kinds are in DRAFTABLE by construction, but a provider exists for four — haplotypes/allele_function/diplotypes (pgx_draft ← CPIC), pharm_variants (clinpgx_draft), variants (clinvar_draft, civic_draft, pubmind_draft). heteroplasmy.csv, repeat_alleles.csv, copynumbers.csv, pgs.csv and activity_phenotype.csv have none, and no pass cross-checks them; enrich() resolves heteroplasmy rows (the third table that can ask, and the one keying with alts) but resolution is not a source. The corpus behind the kind is one module — reference_examples/mt_heteroplasmy, two MT-TL1 variants of one gene, hand-authored — which is @probe-uniform-corpus exactly. MITOMAP is the candidate and two things must be established before that is a plan, neither of them by recall: (1) its terms, which are not CC0 and are asserted nowhere here — RM153 is the standing reminder that a terms page can answer 200 with something that is not terms (@no-named-licence); (2) whether it carries the axis at all, since HeteroplasmyRow binds a level band per (gene, reference_sequence, tissue, variant_key) and a per-variant pathogenicity table with no tissue and no threshold fills the identity columns and none of the binding ones. A negative closes the item, which is why it is named in advance. · from the 2026-09-01 source-adoption round · related RM165 · probed and drafted in PROPOSAL_0_7_PT2 (2026-09-01, proposed PARKS to 0.8 — drafted, not decided; the axis negative is now measured against MITOMAP's own pg_dump, and the source is reachable after all)
  • RM181 — digest by domain. A byte digest moving beside intact signatures says something changed, not what; provenance has no shift tracker of its own. Filed 2026-09-03 from the maintainer's S87 decision, for the 0.8 review of what the hash family covers. · related RM180, RM126
  • RM188 — the competitor survey, scoped into 0.8 by the maintainer on 2026-09-03: run Calwbio's and genomi's pipelines on real input, obtain their reports, read the logic back out of them, and re-fold what survives into module mechanics — a table kind, a bounded rule, a vocabulary member, or a named gap in USE_CASES. One probe document per competitor in the PUBMIND_ASSESSMENT shape; an inference found there becomes the table that would permit it, never the inference. Both surveys filed 2026-09-13 — probes/GENOMI_SURVEY.md and probes/CLAWBIO_SURVEY.md; neither competitor is a format, and what each routes is in its own § plan. Round 2's roster was searched for rather than named, same day: eight GitHub query shapes, 22 repositories skimmed against one question (curated rows a human committed, or a fetch at query time), and nothing in the 2026 crop above ten stars. Four get a survey — OakVar (a real module format with a store; a module there is a plugin, so a typed column contract plus requires, wrapped around arbitrary code), SNPedia via snappy's 106,603-entry extraction, BioMCP (630★, genomi's class at scale), Exomiser (the one established tool shipping annotation as a versioned bundle). Five axes were already found by the skim alone, each sighted independently 2–3×: array/chip callability per variant (we have nothing — requires_callable asks about the consumer's own VCF), per-variant ancestry transferability, an effect modified by a non-genetic factor, source disagreement as a recorded verdict (clin_sig_concordance's shape on tables with no verdict), and a salience separate from clinical severity. The agent-skill class closed for one directory read (awesome-bio-agent-skills, 178★, ClawBio's six skills inside it). · related RM134, RM16, RM28, PUBMIND_ASSESSMENT · also in probes/GENOMI_SURVEY, probes/CLAWBIO_SURVEY

✖ Closed as an item

  • RM173 — ✖ CLOSED 2026-09-02, not shipped, superseded by RM175. It measured a 13-month gap between clinicalAnnotations.zip and clinicalVariants.zip and read it as two live surfaces refreshing out of lockstep. CLINPGX_ARCHIVES established the 15-column table was renamed to summaryAnnotations.zip at the PharmGKB→ClinPGx transition: the old name is a frozen S3 object the API still answers 200, on no downloads page, and the builder's default. Right about the number, wrong about what it meant — the currency signal it ended on is not a thing to publish but a lane reading a dead file. Its own earlier premise (that clinicalVariants.zip is a third source worth adopting) had already been replaced the same day by the 96.3% join; the surviving type finding — one vocabulary, two separators — carries into RM175 unchanged. · from RM166's build · closed into RM175
  • RM102 — ✖ closed 2026-08-21 as a decision not to act, after the half of it that was a real defect had already shipped. The enricher's load_dotenv writes a whole .env into os.environ from library paths, and override=False skips a variable that is present, so deleting one is what lets the file supply it. The bug half — load_dotenv_file=False reaching none of the six resolvers, each computing its default directory as an argument through an unconditionally-loading _cache_dir — shipped in 0.6.3 with a subprocess test per resolver and a registry walk over both helper families. The rest closed on the record rather than on an argument: one incident, and it is not one — S39's reporter lost about an hour to a test failing with the wrong message; the credential was their own, in their own process, from their own .env, and nothing crossed a boundary. Against that, both candidate repairs cost a full minor and the better one is silent for every caller who never passed the parameter (S14's shape). ENRICHER § cache locations documents the behaviour, which was the reporter's own fallback ask. Reopen on something worse than a lost hour — a credential reaching a subprocess, a crash report, any boundary at all. · from S39 (just-module-creator) · also in ENRICHER, CHANGELOG