The consumer-suggestion triage loop¶
How to run CONSUMER_SUGGESTIONS.md as a conversation rather than an inbox: a watcher notices when a consumer has finished writing, an agent triages what is new, and every item gets a maintainer reply written back into the document itself. The document is the transcript, and it is also the state — there is no queue, no database and no external ledger.
A generalized, self-contained copy of this loop is published as a gist —
https://gist.github.com/winternewt/54b94bda01812be937b892146d1bb254 — with the three scripts
parameterized (INBOX/HISTORY/PREFIX) and every repo-specific reference stripped. If you change the
pattern here (the algorithm, a script's contract, a gotcha), update it there too; if you change
something only true of this repo, do not.
Sync in from the gist at the start of a pass, and adopt what has come back¶
The gist is where fixes from other repositories arrive, so it is read as well as written. Other
trees run this loop, and a defect in the pattern is usually met there first — the **Status-preamble
collision in §6 was found while adopting the loop into a second repository, not here. That makes the
published copy an inbound channel, and a fix sitting in it unadopted is the same failure this whole
document is about: something answered somewhere nobody looks. What stays one-way is the content —
the gist never reads this repo's items, only its machinery.
Check the digest first, and pull the files only when it has moved. Fetching both scripts and diffing them every pass is almost always work for nothing: the gist changes rarely, and the diff is noisy enough (below) that reading one is not free either. The gist's own revision id is a digest of the whole thing and is public, so the check costs a single unauthenticated request:
GIST=54b94bda01812be937b892146d1bb254
curl -sf "https://api.github.com/gists/$GIST" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["history"][0]["version"])'
In sync through gist revision 96b6953362bb4854e05db04334086c43c9eeb170, ours, pushed
2026-09-30. It carries the Claude Code arming outward: wake-on-inbox.sh (the uncapped socket
wrapper with its armed canary) and arm-on-prompt.py (the prompt hook, which ignores the watcher's
own [watcher posts), plus a §2 paragraph and two README rows. The existing scripts went across
byte-identical, verified against the pushed version. The unauthenticated API read above was stale
after that push: for a while it still named the previous head and file list, while
gh api /gists/<id> already showed the new one. So verify an outward push through gh api, or by
fetching pinned to the returned version, never by the anonymous check. Before it,
b063b65ef777cf06002e6fa14953d6503692d1f1, inbound, adopted 2026-09-25: the watcher became a newest-wins singleton per watched file (exit 3 when superseded), the
arming moved to a one-shot background task (§1), and a fence-blindness note joined §6. The two Python
scripts and item-next.py were byte-identical to the previous revision, so the fingerprint gate below
was satisfied trivially. Before it, f53d7e181ca89fcd26d5aa9cda365e9d93699345, ours, pushed 2026-09-13,
in sync by construction. It carries one thing outward — the allocator's argument surface, below
— and touches no other file: the three scripts were verified byte-identical against the previous
revision after the write. Its predecessor e81d9a18755ea215144518489c52e5c5669b0b25 (2026-09-01, ours)
was the baseline before it, and carried two things outward: the tracked-item allocator as the
generic item-next.py (§1 here, §2 there) and the duplicate-id entry that had been owed since
2026-08-31. The three existing scripts went across byte-identical, verified against the previous
revision — this push adds a file and extends two documents, and touches nothing else. It supersedes
fd7c3b5e462fcc27c8f597b24f466738c1b42991 (2026-08-21, ours, which only added the CDN note) and
32967e3699c0… before it, the substantive one of that pair; before those fffbcc65653b…
(2026-08-21), which was a genuine inbound adoption, the first since the channel was built and the one
that paid for it. Earlier still: a74366fa6d72… (2026-08-21, ours), bd793a8ce98b… (2026-08-18) and
ab7e2a89c48d… (2026-08-16), the baseline the digest check was first pinned to.
A published gh gotcha found by this push, and it exits 0. gh gist edit <id> -f <name> does not
read stdin on every version — on 2.4.0 it selects a file for interactive editing, so
gh gist edit … -f README.md < README.md reported success and changed nothing; only the API call
gh api -X PATCH /gists/<id> --input <payload.json> actually wrote. Verify a push by re-reading the
version-pinned raw URL and comparing bytes, never by the command's exit code — --add for a new file
did work, which is what made the partial success look complete.
The scripts were byte-identical across the inbound revision — everything that arrived was prose, so
the auto-adopt gate below was satisfied trivially: with no fingerprint() change there was nothing to
restamp.
Same sha back means nothing has arrived and the sync-in is finished for that pass; a different one means pull the files and read the diff. Update this line, with its date, whenever an adoption or a push lands — it is the baseline, and a stale one re-diffs work already taken, which after a push means re-reading your own writing as though a stranger had sent it.
history[0] is the newest entry, checked and not recalled — its committed_at equals the gist's
updated_at. Getting that orientation backwards is the failure worth guarding, because it does not
announce itself: the check would read unchanged forever and the inbound channel would look healthy
while nothing came through it. If api.github.com is ever unreachable, the fallback is a conditional
GET against the raw URLs carrying a stored ETag, where a 304 is the same "unchanged" answer. Build
one or the other, never both.
Pin the raw fetch to the revision the digest check just returned. The bare
<gist>/raw/<name> URL is CDN-cached and goes stale for a good while after a write — measured on
2026-08-21, immediately after our own push: the API reported the new revision while every unversioned
raw URL still served the previous one, and Cache-Control: no-cache did not shift it.
<gist>/raw/<version>/<name> was correct straight away. This matters because the two halves disagree
in the worst direction: the digest says something arrived, the files say nothing changed, and the
natural reading is that the digest check is broken. It is the history[0] hazard one layer down —
an inbound channel that looks healthy while nothing comes through it.
The local names differ for the watcher only (watch-inbox.sh → .claude/watch-suggestions.sh):
GIST=54b94bda01812be937b892146d1bb254
V=$(curl -sf "https://api.github.com/gists/$GIST" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["history"][0]["version"])')
G=https://gist.githubusercontent.com/winternewt/$GIST/raw/$V
for f in triage-state.py triage-archive.py; do curl -sfL "$G/$f" -o "/tmp/gist-$f"; done
diff -u /tmp/gist-triage-state.py .claude/triage-state.py # noisy: most of it is parameterization
Read the diff for logic, not for prose. Nearly every line differs because the gist is
parameterized and names no repo document, so a raw diff buries the two or three lines that matter.
Compare the code with docstrings and comments stripped, and adopt only what is a behaviour change:
in the 2026-08-17 sync those were RULE_RE (a trailing horizontal rule was being hashed as if the
consumer had written it) and the archiver's closing line, which still said "add each one's row to the
index table" — wording §4 retired when the table became a contents list.
Auto-adopt is the default, and it has exactly one gate: does adopting move a fingerprint? A fix to
the machinery is presumed wanted, since it was found by someone running the same loop and there is no
local reason to differ. But a change to fingerprint() re-scores every marker already stamped, and
those sections then read revised — our own adoption impersonating a consumer revision, which is the
one signal the ledger exists to carry. So run the ledger over both documents immediately after,
and if nothing moved you are done.
If something moved, prove the cause before restamping, and the proof is cheap. Do not reach for
git archaeology: if the ledger read all-current immediately before the adoption, then the old
function still matched every recorded marker, so the whole delta is the function and no prose changed.
That check is one command and it is stronger than a diff. Restamp to the new values and say so in the
commit. --backfill will not do it — it only touches unmarked-reply, deliberately, because silently
restamping a revised section is how a genuine re-triage signal gets erased. Adopting RULE_RE moved
four (S2, S6, S7, S12 — the four reports ending in a ---) and they were restamped on that proof.
What must not be adopted, since the gist is the generic copy and this tree is not: the
INBOX/HISTORY/PREFIX parameterization, the FEEDBACK.md defaults, and any pointer to
TRIAGE_LOOP.md, which is the gist's name for this file. Adopting those silently repoints the tools at
documents that do not exist here.
Push the other way too, and check it in the same pass — the sync is bidirectional even though each
direction is a separate act. The standing debt was discharged on 2026-08-18: the branch-pause in
watch-suggestions.sh (§1) and the digest check above both reached the gist in revision bd793a8c…,
genericized, along with the correction to §2's "why not git" that the branch-pause forced — the
published copy still called the loop must not commit a fatal reason for the design, which stopped
being true here when §5's permit was granted, and a watcher that pauses because the loop commits cannot
sit in a document saying it does not. Updating the gist is a publish and stays the user's to authorize;
it was authorized for that push and is not a standing permission.
One item was owed outward on 2026-08-21 and was pushed the same day: stamp the arrival fingerprint, not the value the ledger prints after you write the reply (§6). That is pattern material rather than local tuning — the gist never carried the wrong sentence, it simply never said which sha to write, so a reader who reasonably guesses "run the ledger and copy what it prints" walks into the trap on every reply longer than a paragraph, which is most of them. It went into the gist's Step 3, where the stamping actually happens, and the baseline above moved with it.
Nothing is owed outward as of 2026-09-01. The duplicate-id entry in §6 — an id under two
top-level headings archives as one section, and the fingerprint check reports it as a mutation after
writing — was owed from 2026-08-31 and went across in the push above, into the published §5's
second-repo subsection. It is pattern material rather than local tuning: section_span and the
verification are both in the published copy, so the same refusal happens in any tree running this loop,
and the entry is a hand step plus a why not fixed in the tools. The allocator went with it.
The one item owed since 2026-09-04 was pushed on 2026-09-13, authorized, and nothing is owed
outward as of it. .claude/rm-next.py's reserve path was also its default
path, so any flag it did not recognise fell through to allocate(): --help spent RM189 that way and
a typo spent RM190 while the first was being repaired, each leaving a tombstone row because ids are
never reused. Fixed here in 0d73268 with KNOWN_FLAGS, a real --help, exit 2 on anything else, and
--note's value excluded from the flag scan. The published copy still has it, checked rather than
assumed: item-next.py at the baseline revision ends in a bare fall-through to allocate() with no
flag validation anywhere above it, so item-next.py --help reserves an id in any tree running the
generic loop. It is pattern material by the same test as the allocator itself — the defect is in the
shape a tool whose no-flag path mutates, not in anything local
(@an-index-is-not-an-allocator).
Reproduced against the published copy before porting, which is the only way to tell a generic defect
from a local one. On a scratch index holding one RM1, the gist's own item-next.py --help printed
no help, exited 0, and left a 🔷 reserved row for RM2 behind. The port carries KNOWN_FLAGS, a real
--help printing the module docstring, exit 2 naming the unknown flags, and --note's value excluded
from the flag scan so a note may open with a dash; the usage block gains a --help line and a
paragraph saying why the refusal is load-bearing rather than tidy. Both arms were run on that same
fixture after porting — help prints and the index is untouched, an unknown flag exits 2 and the index
is untouched, and --dry-run, --note and a plain allocation still behave. Verified by re-reading
the version-pinned raw URL and comparing bytes, never by the command's exit code, and the other three
files were diffed against the previous revision to show the write touched nothing else. Nothing else since e81d9a18… is owed: the line-length gate and the
runbook's ruff format bullet are this repository's CI, and the §4 counter's third failure below is an
entry about a section the gist deliberately does not carry.
Three items went outward on 2026-08-21, in revision 32967e36…, with the CDN note following in
fd7c3b5e…, and nothing was owed as of it.
The gist had the furniture entry for prose moving in — a footer swallowed by the last section — and
no flush-left # gotcha at all, in either half; it now carries the joined version, since
triage-archive.py shares the boundary scan and the same line that truncates a reading truncates a
move. Also pushed: the fence-aware span logic with its refusal, and a corpus() fix for --next,
which read a fixed inbox/history pair while the same document describes splitting the history file —
on a fixture with a split-off half the published copy returned an id five short of the true one, each
of them already answered. All three were reproduced against the gist's own scripts in a scratch repo
before being pushed, which is the only way to tell a generic defect from a local one.
The allocator crossed, but its placement here did not. .claude/rm-next.py is published as
item-next.py, parameterized on INDEX/SCAN_DIR/ITEM_PREFIX/ANCHOR — this tree's RM_TOC.md,
docs/ and RM are all defaults over there, and the published copy names none of them. What stayed
local is the reservation anchor being ## ⏳ Open, no release decided, which is this index's own
heading; the generic default is the first ## heading, since a generic copy cannot know one.
§4 and §5 are deliberately not generalized, decided 2026-08-21 — do not re-raise this. The published copy has never carried a thresholds section or an unattended-permit section, and it should not: the numbers are this tree's tuning and the permit is this tree's grant, so a generic copy stating either would be publishing a local decision as though it were the pattern. The practical consequence is that the stale-counter correction of 2026-08-20 (count the idiom, not a fixed version number) has nothing to be stale in over there, and a future pass that notices the absence should read it as intended rather than as a debt.
The live document holds only what is unanswered. An item moves to CONSUMER_SUGGESTIONS_HISTORY.md once its reply is written, which is the same split as ROADMAP.md / ROADMAP_HISTORY.md and exists for the same reason: an inbox only grows, and the eleven unanswered entries this loop was built for were invisible inside 6,000 words of answered ones. So an empty live file means nothing is owed — which is a property worth having and is destroyed the moment answered items are left in place.
The history file has itself been split once, and the loop is unaffected. On 2026-08-17 the items
the 0.5 line answered — S1–S24, S27 and S28 — moved to
history/CONSUMER_SUGGESTIONS_HISTORY_PRE_0_6.md,
byte-for-byte and with every fingerprint intact. triage-archive.py still archives into
CONSUMER_SUGGESTIONS_HISTORY.md, which is the file it should keep writing to; triage-state.py
globs for the history files rather than naming them, so --next counts ids across every half. The
contents list stays whole in the live history file — splitting an index is how an item stops being
findable, and that is what RM_TOC.md exists to remember.
This exists because the notes arrive faster than anyone re-reads them. Eleven of the seventeen entries in that file had no reply of any kind when the loop was built, and six of those are not recorded anywhere else either, so the only way to find out whether an item had been considered was to read 6,000 words and guess.
1. Setup¶
Four scripts, all in .claude/, none packaged. The first three are the Sn loop; the fourth allocates
the RMn the loop's route (a) files, and is the only one that writes outside the two consumer
documents:
.claude/watch-suggestions.sh |
debounced watcher: one line of stdout when the file stops changing. The only one that is really bash |
.claude/triage-state.py |
the ledger: which sections are new, revised, or already answered. Takes a path, so it reads the history file too; --next prints the next unclaimed Sn. Also prints STRUCTURE lines when a fenced block breaks the section boundaries |
.claude/triage-archive.py |
moves answered sections into the history file and verifies each fingerprint survived the move. Refuses outright on a STRUCTURE finding, since a span that ends at the wrong line moves the wrong bytes and the fingerprint check cannot see it |
.claude/rm-next.py |
the RMn allocator: scans docs/ for every number, and reserves the next one by appending a 🔷 reserved row to RM_TOC.md under an advisory flock. --dry-run reads without claiming, --list shows standing reservations, --release RMn drops one |
Run the three Python ones, never bash them — ./.claude/triage-state.py or
python3 .claude/triage-state.py. They carried a .sh extension until 2026-08-16 and the mismatch had
a cost; §6 has it.
RM_TOC.md is an index, and an index is not an allocator¶
This is a reproduced incident. Sn has had --next since the loop was built, because the id is
written into a document. RMn had nothing: the number was read by grepping "the highest in use", and
the gap between reading it and writing the entry is exactly where another session reads. On 2026-09-01
two sessions sharing this working tree filed different work as RM159 a minute apart and the tree
carried two RM159 entries pointing at different items — git 741ec59 renumbered one to RM161, choosing
the one that was cheaper to move because the other pair was contiguous and already referenced from
three probe documents.
Grepping cannot fix it; the claim has to be an atomic write. So rm-next.py does both halves in
one critical section: scan, then append the reservation, with the whole window under the lock. Reading
the maximum outside the lock and appending inside it is the same race with a smaller window.
The lock is on docs/, the directory — never on RM_TOC.md itself, and the second reason is
measured rather than assumed. A lockfile left behind by exactly the kill it guards against would block
every later run (@flock-not-a-lockfile). And flock binds an inode: an editor or an atomic writer
that renames a new file over RM_TOC.md leaves the holder locking an unlinked inode while a second
process locks the new file and acquires immediately — verified in a sandbox before the tool was
written. The directory is stable across that, and it is the idiom
just_dna_enricher.transaction.spec_lock already uses for enrich.
A reservation is a visible index row, not a side-car. 🔷 reserved sits under the open-items
heading; replace it with the item's real row when you write the entry. A number claimed and abandoned
is then visible rather than silently burned, and a state file the index cannot see is how a number
goes missing — the failure RM_TOC.md exists to prevent. --release RMn leaves a ✖ tombstone rather
than deleting the row, because ids are never reused: the first cut deleted it, the number became
invisible to the scan, and a released RM10 was handed straight back out. Same rule as Sn, and the
tombstone must contain the number literally, since a scan is all that reads it.
The lock is pinned by schema/tests/test_rm_allocator.py, which runs eight allocators at once and
runs the same eight with flock neutered to show they collide. A guard nobody has watched fail is a
guess.
Since 2026-09-30 a prompt arms it: name this runbook in a prompt, in any wording. The committed
UserPromptSubmit hook .claude/hooks/arm_triage_watcher.py (registered in .claude/settings.json)
fires on any prompt citing CONSUMER_TRIAGE_LOOP.md. The citation is the only invariant, because the
triage seat is the session pointed at this runbook, so no other seat arms a watcher. A prompt carrying
the watcher's own [watcher prefix never arms it: its event line cites this runbook too, and a
machine-written line re-arming the machine that wrote it is a loop to rule out, not merely bound. The hook starts
.claude/wake-on-suggestions.sh detached. That wrapper runs the watcher outside the harness, which
since Claude Code 2.1.285 stops every background task it owns after its timeout (30 minutes by
default, 2 hours at most). It wakes the session by posting a user message to the session's own
inbox socket ($CLAUDE_CODE_MESSAGING_SOCKET). It keeps watching after each event, so nothing needs
re-arming. It survives /clear, since the process and the socket stay the same, and it exits when it
is superseded or when the session ends and the socket disappears.
Two things to check on every arming. First, the hook adds an [triage watcher] armed on … line to
the turn. Second, the wrapper's first post, [watcher …] armed, arrives during that same turn. That
second one is the canary for the one undocumented thing this path rests on: the socket's message
format ({"type":"user","message":{…}}, taken from Claude Code's own help text; the
cross-session messaging docs specify only
the auth line). If the canary does not arrive, the format changed. Say so, and use the fallback below.
A hook arms it, never the agent, because auto mode's classifier refuses to let a session write or
launch something that wakes itself, and a hook is not classified. Stay out of bypass mode: a session
that bypasses prompts holds every message the watcher posts ("did not attest its permission mode")
until approved. The setting that lifts that hold, crossSessionInbound: "accept", is ignored in repo
and local settings and opens the inbox to every session on the machine. Never forge the sender-mode
field to get past the hold: it is the safety gate itself. schema/tests/test_triage_watcher.py
runs the whole path against a fake inbox socket.
The fallback is the one-shot background task, which wakes the agent on a real event and on
nothing else. Pass timeout: 7200000, and re-arm it when it stops. Bash with
run_in_background: true, command:
coproc W { exec .claude/watch-suggestions.sh; }; p=$W_PID
while IFS= read -r line <&"${W[0]}"; do
case "$line" in *"nothing pending"*|*paus*|*resum*) continue;; esac
echo "$line"; break
done
kill "$p" 2>/dev/null; wait "$p"
Not the Monitor tool any more, decided 2026-09-25. It once took persistent: true; it now caps
every watch at 30 minutes, so a persistent watcher becomes a timer that wakes the agent at each expiry
to be re-armed and spends tokens on an empty inbox. A background task had no expiry until 2.1.285. Save $W_PID
before the loop — bash may unset it once the coprocess exits — and arm again after each event. Adopted
from the gist (revision b063b65e…), where it was found in another tree first.
Arm it on every run, and don't check first. The watcher is a newest-wins singleton per watched file:
a new start records itself in a pidfile under $XDG_RUNTIME_DIR, keyed on the watched file's resolved
path, and stops the previous owner after checking that pid's command line; every poll it re-reads the
pidfile and exits if it is no longer the owner. So arming twice leaves one watcher. Checking first is
what goes wrong: the harness's task list shows only what the current session spawned, so re-arming on
"no tasks found" left two, and every settle was reported twice. A sibling repository running its own
watch-suggestions.sh is a different key and is never touched.
The task's exit status says how it ended, because the wrapper ends in wait on the watcher:
| status | meaning | action |
|---|---|---|
| 0, with a line | an event | triage, then arm again |
| 3, no output | superseded: a newer arming took over | none, the newer watcher is live |
| 0, no output | stopped on purpose (a TERM while it still owned the pidfile) | arm again if still wanted |
| anything else | the watcher failed | read the output, fix, arm again |
The harness reports the superseded case as failed with exit code 3; that is expected. Verified on a scratch inbox before adoption: a second arming made the first exit 3, a TERM to the owner exited 0 and removed the pidfile, and an append fired the event line.
TaskStop cancels it. It reacts only while the
session is open and the REPL is idle. Editing the script does not reach a running monitor — bash
reads a script incrementally — so TaskStop and re-arm after changing it. Nothing needs installing — inotify-tools, entr, fswatch and
python watchdog are all absent from this machine, and stat polling is enough at this cadence.
Hooks cannot do this job. Claude Code hooks fire on the agent's own lifecycle (PreToolUse,
PostToolUse, SessionStart, Stop); a consumer editing a file triggers none of them. The trigger has
to be a process.
The cooldown is sized for an agent author¶
Consumers write these notes through an agent, so the shape is a burst — five edits in a minute — with
gaps wherever the agent stops to read or probe something. A 60-second timer fires in the middle of such
a run and triages half a note. COOLDOWN defaults to 150s, POLL to 10s, so an event lands 150–160
seconds after the last write. Consecutive saves inside the cooldown collapse into one event, because
each mtime bump restarts the timer.
Raise it if a consumer's runs are slower. The only cost of waiting is latency, and the cost of firing early is a reply to a half-written item.
First run¶
The watcher seeds itself from the current mtime with the dirty flag clear, so it never fires for a change that predates it. Run the ledger once at startup to pick up the standing backlog:
.claude/triage-state.py # every section and its verdict
.claude/triage-state.py --pending # just the ones needing work
Point it at the history file too, every time — an empty inbox is not an all-clear. That run is the
lint described in §6, and it is the only thing standing between a mis-archive and a permanently lost
item: a well-formed archived section reads current, so anything else means a marker went in wrong or
did not survive the move. On 2026-08-17 the inbox was empty and this turned up two, S35 revised and
S36 unmarked-reply, both from the pass that had archived them.
2. How state is derived¶
The document is the ledger. A triaged section carries a marker inside its reply, holding a fingerprint of the consumer's own text:
The fingerprint covers the section body with every **Status paragraph and every marker removed, so
it describes what the consumer wrote and never what we replied. Lines are right-stripped and blank runs
collapsed, so trailing whitespace and reflowing do not count as a change. Four verdicts follow:
| verdict | meaning | action |
|---|---|---|
new |
no reply, no marker | triage it |
revised |
marker present, fingerprint moved | the consumer edited an answered item — re-triage |
unmarked-reply |
answered before the ledger existed | --backfill stamps the marker — only for replies older than the ledger, never one you just wrote (§6) |
current |
marker matches | nothing to do |
Why not git¶
One fatal reason, and one that expired. A consumer may well commit their own addition, at which point
git diff HEAD is empty and the loop sees nothing at all — that one is unchanged and is on its own
enough. The original second reason was that the loop must not commit, so a HEAD baseline would never
advance; the §5 permit retired that premise on 2026-08-11, and it is recorded here rather than quietly
deleted because a design defended by two reasons is worth re-checking when one of them goes. The design
survives it: the surviving reason is fatal by itself, and the ledger's properties below were never
consequences of the commit rule. git diff and git log -p stay in the loop for reading what changed,
as context. Correctness never depends on them.
The in-document ledger has properties no side-car state has: it works on an uncommitted tree, survives anyone's commits, travels with the repo, is legible to a human scanning for the backlog, and cannot drift out of sync with the replies it describes.
Self-firing is not a loop¶
Writing a reply bumps the mtime, so the watcher fires again. That run finds nothing pending — the fingerprint excludes the reply — and no-ops. Expect the second notification; it is the mechanism working.
3. The algorithm¶
Step 0 — establish what already shipped¶
Do this before reproducing, and certainly before designing. new in the ledger means no reply in
the document, never no work done. On the first full run, two of eleven items were already fixed —
S1 by 0.4.1 plus RM17, S2 by a block in authoring_reference() whose code comment names S2 — and a
third's preferred option had shipped in 0.5.2 from a different report (S14). Answering those as though
they were open would have designed a feature that existed.
Cheap and mechanical: grep the item's symbols in the source, then docs/CHANGELOG.md,
ROADMAP_HISTORY.md and RM_TOC*.md for its subject; git log -S "<a phrase from the fix>" finds when a
guard landed. For S1 this collapsed a feature request into one missing error message.
Step 0b — reproduce before classifying¶
Compare the claim against the code, not only against the docs, because the docs are often the thing
that is wrong. This step is the only thing separating a real defect from bucket (b), and it cuts
both ways in this very file: S6's chrY half did not reproduce against a real SRY row, while S13's
warning-string marker reproduced end to end. Scope the probe to the table you actually looked at — an
unscoped negative finding becomes a permanent false constraint.
Probe the behaviour, not only the sentence — the probe is where the adjacent defect turns up. S16
asked whether unknown files in a spec directory are tolerated; building the probe meant putting several
files there, one of which was varaints.csv, and that revealed a mistyped table name being dropped from
a green compile — a defect nobody had reported. S7's answer needed three compiles and a second probe of
merge_sources_csv to establish that a rebuild cannot move fetched_at unless the sidecar is deleted,
which is the fact the whole item turns on.
A bucket-(b) verdict is not the cheap outcome. Three of the first eleven were non-issues and each cost the most probing, because "nothing is wrong here" has to be shown, and a reply that cannot show it is worthless.
Step 1 — charter legality, first-hand¶
Read CONSTITUTION.md yourself. Never delegate this to a subagent: a summary of a charter drops the qualifier the decision turned on, and this step decides whether a repair is legal at all. It is the one part of the loop that cannot be automated away.
Legality sizes the release; severity only orders the queue inside it. This is the correction most worth carrying:
| change | release | why |
|---|---|---|
| new optional column, table, or manifest field | minor | additive; an unset optional column is omitted from content_signature. A manifest field was never in artifact.digest at all, so it is cheaper still |
| pure legibility — a warning, a count, an error message, a doc | patch | moves no authored identity. A better diagnosis on a path that already failed changes no verdict |
| a new flag, parameter or alias beside the old one | minor | additive; P3 keeps a superseded name as a working alias |
| removal, promotion to required, retyping — including a rename | major | breaks a reader or invalidates published data. A rename is a removal plus an addition, so the addition being legal does not make the rename legal (S14) |
So a severe finding whose fix is a new optional field is still minor (S13/RM44), and a trivial one whose
fix is a retype is still major. Severity decides what gets done first, never what version it lands in.
"It moves artifact.digest" is not on its own a reason to defer: P4 scopes byte-reproducibility to a
fixed compiler_version, and the authored identity does not move.
Where the amendment does not reach. The 2026-08-11 amendment is about columns and tables. Do not
stretch it to a published function signature or CLI flag — S14's rename stays major for that reason,
while adding a differently-named alias would have been minor. And a change to a dataclass a consumer
reads (Finding.row) is best made additively for the same reason a column is: S18's fix adds line
rather than redefining row, because a consumer already compensating for the old meaning would break
silently.
Round-trip is the trap. P7 can make the obvious repair illegal outright — RM43's option 1 does not merely
move a digest, it moves content_signature, because reverse_module re-emits a filled coordinate as an
authored one. A repair that fails P7 routes to (d), not (a).
Step 2 — route¶
| verdict | lands in | must contain | |
|---|---|---|---|
| a | real, repairable, legal | ROADMAP.md as ## RMn plus an RM_TOC.md entry |
**Severity** … · **Status** open — 0.x · **Owner** … · **Motivating case** (Sn in CONSUMER_SUGGESTIONS_HISTORY.md) |
| b | non-issue | the reply only | what was probed and did not reproduce. Never a bare "works as intended" |
| c | documentation defect | the doc, fixed in the same pass | the reply naming the file changed |
| d | real, no acceptable repair | ROADMAP.md, Status open only |
the paragraph saying why each candidate repair is wrong |
An item filed as a documentation gap usually has a code half — look for it. All three entries a consumer grouped under "Documentation gaps" (S15–S17) ended up with code: a lock, a near-miss guard, and a specific error message. The reporter is describing where they got stuck, which is a fact about the docs; what stuck them is often a surface that could have told them.
A wrong consumer conclusion is a place to look for our own defect. S7's author read a moved
artifact.digest as a moved content identity — and SCHEMAS.md's hash table was calling the digest "the
version's immutable content identity", against the charter. Bucket (b) on the report with bucket
(c) underneath it is a common pairing, and the (c) half is the one that stops the next person filing
the same item.
Bucket (d) is the one that earns its keep, and the one an unattended agent does worst. It is the
fix it / surface it line from CLAUDE.md: surface anything whose obvious repair is
itself a design decision, and say why each candidate fails. RM33's paragraph is the model, because one
of its two obvious fixes is charter-illegal.
If (a) ships in the same pass rather than being filed: code plus a test, a
CHANGELOG.md entry, and the item moves to
ROADMAP_HISTORY.md with its rationale and a **Residuals** paragraph
(RELEASE_CYCLE.md § Closing an item, enforced by
test_closure_residuals.py). A remainder is filed as its own RMn whose entry names this one, or
closed with won't fix — <reason>. It never goes to a parked bullet or to the consumer's reply alone:
RM31 did both, and the defect came back as S117, S120 and S121. RM_TOC.md is updated either way — it
is the single complete index, and an item missing from it is how RM33 became unfindable.
Step 3 — write the reply¶
**Status —** is the existing house idiom in this file; do not invent a second one. It goes first in
the section, immediately after the heading, and says four things: the verdict, where it landed (an
RMn link, a doc, or a shipped version), what was actually reproduced, and what the consumer should do
now.
## S14 — `resolve_with_ensembl=False` reads as "skip Ensembl" …
**Status — accepted; suggestion (1) shipped in compiler 0.5.4.** Reproduced: a spec with a
committed `resolution.csv` compiles green with every coordinate null under `--no-resolve`.
The warning now names the unread table and its row count. Suggestion (2) is filed as
[RM45](ROADMAP_HISTORY.md#rm45--…) — splitting the flag is a CLI break we want to make once.
<!-- triaged: 0.5.4 · sha 43031f8f63b3 -->
Append, never edit. The consumer's prose is evidence and stays byte-for-byte; the 0.5.2 block says so in the file itself — "the notes below are left as written — they are the report, not the resolution." Same rule as drafting: it appends, it never mutates.
One reply may cover several sections, as the 0.5.2 block does for S3–S6. The ledger understands that shape and marks each covered section individually, since one paragraph cannot carry four fingerprints.
Stamp the fingerprint the section carried before the reply existed, and do not copy the value
the ledger prints once your reply is in place and the marker is not yet. With no marker to stop at,
reply_end() falls back to the single-paragraph rule, so paragraphs two onward of your own reply leak
into the hash. That contaminated value is not an edge case — it is the normal state of every section at
the moment you go to stamp it, and any reply worth writing is longer than one paragraph.
Do not reach for --backfill to dodge that. It computes the same contaminated value, and it puts the
marker somewhere that makes the mistake permanent and silent — §6 has the reproduction. It is for replies
older than the ledger, and only those.
The placeholder recipe needs nothing remembered, and it is the reliable way to get the value. Stamp twelve zeros at the end of the reply as you write it:
Then run the ledger. With a marker present the reply is excluded whole, so the revised line prints the
true fingerprint — sha <real> (was 000000000000) — and you paste that back and re-run to confirm
current. Do not rely on having kept the value from when the section read new: it is the same number,
but a batch that answers two items after both arrived never printed either of them alone. That is
exactly the shape of the S63/S64 pair, which arrived four minutes apart.
Step 4 — move the answered item to the history file¶
It cuts each section — heading, consumer prose, reply and marker — out of
CONSUMER_SUGGESTIONS.md, appends it to
CONSUMER_SUGGESTIONS_HISTORY.md under its group heading, and then
verifies the move: every section's fingerprint is compared before and after, and the write is
rejected if one changed. Do this by hand only if the tool cannot (it prints what it would move under
--dry-run).
- Byte-for-byte, and checked rather than intended. The one property that matters is that the consumer's prose is unchanged, and it is easy to break by reflowing a line while pasting. A changed sha means the prose was touched — that is why the tool refuses rather than reports.
- Add the contents line, and keep it to one line. The tool does not generate it: naming what an item
was and how it ended is editorial. The format is
- **Sn** <what it was> — <status> (<RM if any>), under 80 characters, because it is a contents list and not a second copy of the reply — the detail lives in the section's**Status —**paragraph, the one place it cannot drift out of step with the answer. (The first version of this index was a four-column table with a paragraph per row; it duplicated every reply and was already the thing most likely to rot.) The list still matters for the reasonRM_TOC.mddoes — an item missing from it is howRM33became unfindable, and inbound links elsewhere in the repo are file-level (S5 in CONSUMER_SUGGESTIONS.md), so a reader following one lands on the live file and needs a pointer onward. - A group's dateline is repeated in both files when its items split — while S8 was still open, the "adopting 0.5.2" dateline sat in each. Repeating four lines of context beats moving a preamble away from an item it still introduces.
- A block reply travels with the items it answers. The 0.5.2 block reply sits in the just-dna-lite preamble and answers S3–S6, so that whole group moved together.
- Archive in one pass at the end if you are working a batch. Sections are appended in the order
given, so a group whose items are archived in two batches ends up with its heading twice; the fix is
mechanical (move the later sections under the first heading) but avoidable. Check
grep -n '^# 'on the history file when you are done, and give a section its own heading if it was appended under one it does not belong to — S18 arrived after S17 and would otherwise read as a "documentation gap". An item filed under no group heading needs one written by hand, and the tool now says so instead of guessing: it printsNo group heading travelled with Snand you add a#line naming who reported it and when. That notice exists because the silent version of this went wrong twice — see the title-is-not-a-group gotcha in §6.
Step 5 — hygiene¶
- Serial, one item at a time, and claim the
RMnwith the allocator —.claude/rm-next.py, which scans and reserves in one locked write. Never keep the number in your head across a long pass, and never read it off RM_TOC.md by eye: this bullet said to do exactly that until 2026-09-01, and it is how two sessions in one tree both filed RM159 (§1 has the incident). The index answers where does RM47 live; it cannot answer which number is mine, because reading it claims nothing. The number this bullet used to name is deliberately gone — a counter written into prose goes stale the day after, which is@counted-prose-needs-a-fixed-fieldone document over. - Commit as you go, one commit per item — see the grant in §5 for what that covers and where it
stops. Stage explicit paths; never
git add -A. (This bullet read do not commit until 2026-08-11, when the loop got a permit of its own; the global default followed on 2026-08-21 and now says the same thing everywhere.) - Say what was skipped. If an item is left untriaged, leave it
newrather than writing a placeholder reply. An empty verdict is honest; a hedged one is not. - Run the whole suite after each fix, and
ruff checkplusruff format <files>before you finish — since 2026-09-11 CI gates onruff format --check .as well, so an unformatted commit fails the job. Markdown is outside the gate, deliberately: the formatter rewrites fenced Python inside.md, and a consumer's code fence in an archivedSnmust stay exactly as written or the fingerprint moves. Six code fixes in one pass touched all three packages; the suite went 1382 → 1410 and stayed green throughout, which is the only reason a batch that size is safe to leave in the tree. - A new item can arrive mid-pass. S18 was filed while this pass was running and the watcher picked it
up on its next fire. Take it in the same pass if the context is warm, or leave it
new— but do not let it silently miss the CHANGELOG entry the rest of the batch gets. - Write the CHANGELOG entry for the batch, not per item. One dated heading covering the pass, with
the two-of-eleven-already-shipped fact in it: a future reader needs to know the loop's first run found
answered work sitting unanswered, or they will assume
newmeant untouched.
4. Thresholds — when to call the user¶
The loop's output is roadmap items and patch-level fixes, and it will produce both indefinitely without ever deciding to build or release anything. Triage answers a consumer; it does not schedule the work or cut the version. Those two are the user's call, so the loop has to ask rather than keep accumulating silently. Since 2026-09-25 only the last of them asks anything: the item count is a lamp the handoff reports, and the ceiling says stop, so do not collapse them into one "the loop halts" rule. All are counted off the tree, never remembered:
grep -c 'Status\*\* open — \*\*a minor' docs/ROADMAP.md # sizeable items, awaiting a minor
grep -c 'Status\*\* open — \*\*a patch' docs/ROADMAP.md # patch-level fixes
grep -h '^version' */pyproject.toml # versus the top CHANGELOG heading
Count the idiom, not a fixed version number — and the idiom is a fixed field ROADMAP.md protects.
Those greps read the two forms ROADMAP.md actually
uses — **Status** open — **a minor, release undecided** and **a patch** — because the first version
of this block counted **0.6** and went on counting it after 0.6 shipped. It then returned 0 however
full the queue was, which is the worst way for a counter to fail: a stale one reads as an all-clear, and
the whole point of counting off the tree is that nothing here is remembered. It then failed a
second time when the idiom itself drifted, which is why
ROADMAP.md § Active items now states that the release-class phrase is a
fixed field written verbatim and kept on one line. §6 has the incident; the short version is that a
counter reading prose is only as stable as the sentence it reads, so the sentence has to know it is
being read.
Twenty or more sizeable open items is a lamp, not a start pistol (since 2026-09-25). Report the
count in the pass handoff and do not ask about it. Whether a minor is worth starting is decided in an
interview with the user, on the items themselves, and RELEASE_CYCLE.md is the
sequence that follows. Until 2026-09-25 this paragraph said development of the next minor should
START and told the loop to raise an AskUserQuestion; the user retired that because an item count
says nothing about whether the items are worth building. One thing from the old reading still holds:
the count is never a scope freeze or a stop-filing rule. A minor keeps taking additive items right
up until it is cut. Re-read the set before reporting it: an item that duplicates another, or that
never had a reproduced case under it, is not grounded and should be merged or demoted rather than
counted.
Thirty sizeable open items is the ceiling — there, stop filing and block. A backlog that size means the lamp has been lit and unanswered for long enough that the queue is no longer being managed by anyone, and a thirty-first item buys nothing: nobody reads that far, and an unread item is indistinguishable from an unfiled one. Say what you would have filed, in the reply to the consumer, and block rather than adding to a list that has stopped being a plan.
Patch-level fixes are not counted toward a cut (since 2026-09-25). Every non-blocked patch item
goes into the next patch batch on main by default, and a patch carries however many fixes the latest
wave of usage produced; see RELEASE_CYCLE.md. A minor-class fix is never built
on main: file it, and build it only on the open minor branch (0.8 today) when the loop runs there
with the maintainer's grant; main holds patch scope only, so a patch can be cut from it at any
moment. This paragraph read around ten
accumulated patch-level fixes, publish time until then. Publishing is the user's domain and always
an ask; bumping, tagging and building a dist are inside §5's grant for a patch, which is that
section's own worked example.
But the grant is scoped to a patch, and a batch's release class is set by its most additive item. This sentence read never bump a version, tag, or publish; ask until 2026-08-24, which was written before §5 existed and then contradicted it — found by running the loop, on a pass whose twelve items were mostly warnings and error messages and which therefore looked like a patch. Three of them added an optional field or a public function, so the batch cut as a minor. The rule that resolves both halves: legality sizes the release (§3 Step 1), so count the class of the most additive item rather than the mood of the batch — and a minor is an ask, because whether a minor starts and when it is cut are the user's decisions (RELEASE_CYCLE.md).
Fewer is fine when something is critical — a wrong published number, a false claim in a printed contract, anything a consumer could act on and be harmed by. That is a judgement call and it is the agent's to make; do not sit on one because a counter reads four.
The numbers will drift, and they are a trigger rather than a law. This is an active testing phase, so they were picked to be roughly right and are expected to move — the two item counts were 10/20 until 2026-08-20, raised to 20/30 because a queue of eleven grounded items was nowhere near needing a planning decision and the loop was spending the trigger too early. On 2026-09-25 the twenty stopped being a trigger at all and became a lamp, and the patch count was dropped (the two paragraphs above). Update them here when they move again — and note that the count is deliberately of sizeable items, since a batch of one-line legibility fixes is not the thing that needs a planning decision.
5. What the loop agent may do unattended¶
A standing grant, given 2026-08-11, back when it overrode the global "only commit when asked"
default for this loop and nothing else. Since 2026-08-21 the global preference is commit-as-you-go with
tagging delegated, so the committing half is no longer special — what stays specific to this loop is
doing it unattended, across a long pass, including the version bump and dist build that a §4
threshold fires. It exists because a batch that sits uncommitted for a whole pass is one /clear away
from being unattributable, and because the §4 thresholds are worthless if the agent cannot act on
them.
Granted. Commit as you go — one commit per item, the loop's existing serial rule, with the Sn/RMn
in the message. Bump a version, tag it, and clean and build dists when a §4 threshold fires; bumping is
inside the grant by implication, since a 0.5.4 tag cannot exist while pyproject.toml reads 0.5.3.
Bump and build only the packages whose code moved — a partial cut is normal here, and schema has
sat at 0.5.0 across two compiler releases. When schema is next cut it takes the aligned number
rather than the next one in its own sequence (decided 2026-08-11: one version across the workspace beats
a dense per-package count, so a 0.5.4 compiler names a 0.5.4 schema).
uv sync and relock before tagging — the common gotcha, and it is silent. uv.lock is tracked and
records each workspace member's own version, so a bump that stops at pyproject.toml leaves the lock
naming the release you just left. Tag that and the tag captures a tree that cannot reproduce itself,
while uv run <cmd> keeps resolving through a stale wrapper. The order is bump → uv sync → commit
the lockfile with the bump → tag, so the tag includes the relock rather than trailing it. Check before
tagging, never after:
grep -A1 'name = "just-dna-\(format\|compiler\|enricher\)"' uv.lock # versus */pyproject.toml
git status --short uv.lock # must be clean at tag time
Not granted, and not by omission. git push, uv publish, tag pushes, releases — everything
outward-facing stays the user's. Building a dist is not publishing one; the artifact sits in dist/
until a human sends it. Wipe dist/ before building, because uv publish uploads everything it
finds there and a stale wheel left by this loop becomes the user's footgun at their publish step.
No history rewriting: linear commits only. No amend, no rebase, no reset, no merge commits, no
force-push, and no git stash in any form — that one has already swept an entire session's uncommitted
work here. Stage explicit paths; never git add -A, which is how a .env swap file with live tokens
once got committed.
If you corner yourself, confess with an AskUserQuestion. This is the other half of the rewriting
ban rather than a separate rule: amend and reset are exactly the tools you would reach for to make a bad
commit disappear, so a mistake stays visible by construction. A commit that should not have been made, a
tag on the wrong commit, a build from a dirty tree — say so, say what state the tree is actually in, and
offer the fix-forward options. A wrong commit followed by a correcting commit is a legible history; a
rewritten one is a lie the user cannot audit.
6. Gotchas found while building this¶
Each of these was a bug in the loop, not a hypothetical:
- A reply can live outside the section it answers. The 0.5.2 block sits under an
#heading and answers S3, S4 and S6 by name, so a naive presence test reads four answered sections as new.block_replies()reads that convention instead of forcing a new one. - The marker must not be hashed. Marking a block-replied section puts a standalone comment in its
body; when the fingerprint covered it, the section read
revisedfrom the instant it was marked. The marker is stripped wherever it appears, not only inside a reply. - A reply ends at its marker, not at the first blank line — found on the loop's own first run.
consumer_text()skipped the**Statusparagraph, so a reply of several paragraphs (the normal size for one that says what was probed, where it landed, and why a candidate repair was rejected) leaked paragraphs two onward into the fingerprint, and writing a reply reported the sectionrevisedimmediately. That is the same self-firing failure the marker exclusion exists to stop, arriving by a different route.reply_end()now runs to the marker; with no marker ahead it falls back to the single paragraph, which is what keepsunmarked-replysections from all hashing to the same empty text. The fix is checkable rather than plausible: S1 and S2 hash back to the exact fingerprints they carried before their replies were written. - Splitting a wrapped paragraph is a substantive change, and correctly reports as
revised. Only trailing whitespace and blank-run length are normalized away. S9appears twice as a heading —## S9at the 0.5-era notes and### S9as a follow-up nested inside S13's section. The ledger keys on top-level## Snand folds###into the parent, so that follow-up currently counts toward S13's fingerprint. Harmless, but it is why the unit is the top-level section.- The event line needs a cap. With a 17-item backlog the notification listed every one; it now shows
eight and
+N more. - A document's own title is not a group heading, and
triage-archive.pythought it was. It took the last#heading before a section as that section's group; for an item filed under no group — the normal shape once the split made the inbox empty, since a consumer appending one report writes no group heading — the last#before it is the live file's own# Consumer suggestions, whose span runs to the next##and therefore swallows the whole inbox preamble. So archiving appended a copy of "this file is the inbox, so an empty one means nothing is owed" into the history file, twice, before anyone read the result. The verification could not see it: fingerprints cover the consumer's prose alone, so the move reported "every fingerprint intact" while the file grew duplicate front matter. Fixed by taking the first#in a document as its title and a group as any later one; a section with no group now prints a notice telling you to add one, rather than the tool inventing a name, on the same reasoning that keeps the index row hand-written. Reproduced both ways in a sandbox before and after.
It does not disturb the preceding section, and the first repair of this claimed it did. The
injected heading is separated by one blank line, and fingerprint() ends in .strip(), so the
section above keeps its hash. Worth stating because the tempting story — "archiving S25 shifted S24" —
was written into this file and one commit message before being checked, and it explains a symptom that
has a different cause (below). The lesson is the loop's own: establish it, then write it.
- A marker can be stamped with a sha that never matched its section, and the ledger can only report
it, not explain it. S24 read revised with prose that git shows byte-identical since the pass that
archived it — verified by running the ledger against CONSUMER_SUGGESTIONS_HISTORY.md at that very
commit, where it already disagreed. So the marker was written wrong in its own pass, most likely by
stamping before a last edit to the consumer's text, and no later operation is implicated. Restamped
to the computed value on 2026-08-12 after establishing that much. --backfill deliberately will not
do this — it only touches unmarked-reply, because silently restamping a revised section is exactly
how a genuine re-triage signal would be erased — so it is a hand edit, and it needs the prose-unchanged
check first. If you cannot show the prose is unchanged, the verdict is honest and you re-triage.
- A Python script named .sh gets run as bash sooner or later, and import is an ImageMagick
binary. Both tools shipped as triage-state.sh / triage-archive.sh with a
#!/usr/bin/env python3 shebang, which is correct only for a caller who executes them. Someone —
a person, or an agent going by the extension — eventually types bash .claude/triage-state.sh. Bash
ignores the shebang, reads the docstring as commands, and reaches import hashlib, where
/usr/bin/import is ImageMagick's screen-capture tool and takes its argument as an output
filename. The repository root grew four empty files named hashlib, pathlib, re and sys — the
script's own imports, in order. triage-archive writes a subprocess instead, which is how you tell
which of the two was run. Renamed to .py on 2026-08-16.
Two things make this worth recording rather than just fixing. It is silent: import creates the
output file and then fails on its security policy, so nothing announces a write, and the litter is
0 bytes with plausible names — a stray re in a project root reads as a vendored module, not as
debris. And the extension was the entire invitation: nobody types bash foo.py. So the fix is
the rename; a guard would be defending the wrong door. Established before being written down —
bash -c 'import sys' in an empty directory creates sys, and running the real script under bash
reproduces the full set.
The same reasoning removed every bare-path invocation: triage-archive.py shells out to the ledger
through sys.executable and watch-suggestions.sh through $PYTHON, so neither the exec bit nor
the shebang is load-bearing anywhere any more.
- An id under two top-level headings archives as one section, and the fingerprint check reports it
as a mutation. Found on 2026-08-31, when a consumer filed S76 as a withdrawal section plus the
original text below it, both at ## S76. section_span resolves an id to the first matching
heading, so triage-archive.py S76 moved the withdrawal and left the evidence orphaned in the live
file — and the before/after check then compared two different sections and refused with the prose
was not moved verbatim, after writing. That refusal is the guard working rather than a second
bug: it cannot tell a duplicate id from a mutated section, and reporting is the safe direction.
Two things to do, and the first is the whole fix. Give the second section a distinguishing
heading — ## S76 (original text) — … is what the reporter used and it is the right shape, since
the ledger keys on ## Sn and folds ### into the parent (the S9 entry above is the same
mechanism). Then move it by hand after the tool run and verify it against
git show HEAD:docs/CONSUMER_SUGGESTIONS.md rather than by eye, which is the check the tool would
have done. Put it directly under the section it is evidence for, inside that group heading.
Not fixed in the tools, deliberately: making section_span return every match would make
triage-archive.py S76 move two sections for one argument, which is a worse surprise than a refusal
— and the shape is rare enough that a documented hand step beats a flag nobody remembers. What the
refusal cost was one careful minute; what a silent partial move costs is an orphaned half-report,
which is the S62 failure this file already carries.
- The archiver verifies the move, not the verdict. It will archive a section the ledger still calls
newwithout complaint: it checks that the prose arrived byte-for-byte, not that anyone answered it. The lint is the ledger itself pointed at the history file — a well-formed archived section readscurrentthere, so anything readingneworunmarked-replywas either archived unanswered or lost its marker in transit. Two seconds, and it is the only thing between a silent mis-archive and a permanently lost item. - A preamble line beginning
**Statusis read as a block reply, and marks every id it names answered.**Status:** intake for field notes — S1 and S2 are openis an entirely ordinary thing to write at the top of an inbox and collides with the reply idiom exactly;--backfillthen stamps both untriaged sectionscurrent. The loop's one job, inverted, by a line of prose nobody would look at twice. The block-reply rule is right — a release note under a#heading really does answer items by name — so the fix is on the writing side: do not open a preamble line with**Statusunless it is a real reply; a blockquote (> **Status:** …) is enough. Found while adopting the loop into a second repository rather than here, which is the argument for keeping the published copy in sync. STATUS_REandMARKER_REare still fence-blind, deliberately. The boundary tests consultfenced_lines, but reply and marker detection do not. So a consumer who quotes a**Statusline or a<!-- triaged: … -->marker inside a code block (while reporting a problem with this loop, say) has a section the ledger treats as answered, or even ascurrent. Left unfixed because the boundary case is the one that loses data and this one has not happened yet. If it does, the fix is oneif i in fencedinhas_reply,consumer_text,stored_shaandblock_replies, using the scan already there. Written down so nobody reads the tools are fence-aware as covering more than it does. Adopted from the gist, 2026-09-25.- A link guard and "the prose is evidence" collide, and the guard must give way.
test_doc_links.pyrequires every relative link in every markdown file to resolve, and it exists because an item moving live→history breaks every pointer at it. Consumers have now started citingRMnitems by link inside their reports — so when RM89 shipped, S35's quoted prose held a link the guard called dead, and the pass that archived the next item quietly retargeted it. That edit moved S35's fingerprint, and the ledger duly reported the sectionrevised: our own edit wearing a consumer revision's clothes, which is precisely the signal the marker exclusion exists to keep clean. Never edit a report to satisfy a tool. The link was not even stale in the sense the guard means — it records where RM89 lived on the day it was written, and the reply directly above it carries the current pointer. Fixed on the guard's side:_verbatim_linesexempts everything below a section's marker in both consumer documents, our replies stay checked because they sit above it, and a test asserts a genuinely dead anchor still survives down there. The prose was restored to what the consumer wrote and the fingerprint came back to the exact value it had carried. -
An off-switch spelled
${VAR:-default}is not an off-switch, and the watcher'sBRANCHknob was one.${BRANCH:-main}treats an explicitly empty value as unset, soBRANCH=— the one thing anybody would type to turn the pause off — silently restoredmainand enabled it instead. Found while porting the branch-pause to the gist on 2026-08-18, by running the off-switch in a sandbox rather than reading it, and fixed in both copies to${BRANCH-main}. The two spellings behave identically for every value except the empty one, which is exactly the value no test reaches by accident; the general form is that a knob's disabling value is its own case and needs its own probe. -
A
#at column 1 inside a fenced code block ends the section, and the ledger cannot tell.BOUNDARY_REis^#{1,2}and knows nothing about fences, so a python comment written flush left in a reply —# StudyRow, 0.6above a field declaration, which is the natural way to label a snippet — truncates the section there. Everything after it, including the marker, is outside the body the ledger reads, sostored_shareturnsNoneand a freshly-stamped section reportsunmarked-replyforever. It looks exactly like the git-sha failure below and has a different cause; what distinguishes them is thatstored_shaon the hand-sliced section finds the marker while the tool's ownsections()does not. Found writing S55's reply.
This was a writing-side rule until 2026-08-21, and that was the wrong call — it is fixed in the
tools now. The original reasoning was that teaching the splitter about fences means teaching it
about indented fences, tildes and nested blocks, and that the failure is rare and self-announcing
once you know the shape. Both halves of that turned out to be false, which is why the entry below
it exists: the shape occurred twice more, and it announced nothing either time — it split S55 and
then S62 across two documents while every check reported green. The deciding argument is that the
rule had no owner for the case that matters. A flush-left # in our reply we can indent; one
in a consumer's report we may not touch, and that is the standing rule this loop will not break. A
writing-side fix that cannot be applied to half its cases is not a fix.
fenced_lines() in the ledger now tracks `` and ~~~ openers indented up to three spaces, closing
on a run of the same character at least as long with no info string, and every span in both tools
is derived from it —sections,block_replies,backfill,section_span,group_spanandcurrent_group.triage-archive.pyimports them rather than keeping its own copy, because the two
tools disagreeing about a boundary is exactly how S62 got cut in half.fence_findings()` is the
backstop for what the tracker still cannot parse, and the archiver refuses rather than warning.
- The §4 counter has now failed three times, by three different routes, and all three were
silent. The first
time it counted a literal
**0.6**and went on counting it after 0.6 shipped. The fix was to count the idiom instead — and on 2026-08-21 the idiom moved: a decision round rewrote all four open items' status lines to lead with what had just been decided (**shape decided 2026-08-21**,**narrowed … to the observability half**), pushing the release-class phrase out of the slot the grep reads and, in one case, wrapping it across a line so no line-based grep could match it whole. The count went six to zero with four grounded items in the file, which is the worst direction for this failure: a zero reads as an all-clear, and the whole reason §4 counts off the tree is that nothing here is remembered. Found by reading the tree rather than by the counter, because the counter cannot report its own blindness.
The general shape is that a counter reading prose needs a fixed field, and the field has to be
named as load-bearing where it is written, not only where it is read. Both failures happened
because the phrase lived in ROADMAP.md while the only statement that it mattered lived here, in
the file the person editing a status line is not reading. So the rule now sits in
ROADMAP.md § Active items as well: the release-class token is verbatim,
on one line, and anything else the status wants to say goes after it. That is
@warning-text-is-api applied to our own documents — a phrase a tool greps is an API whether or
not the tool is a consumer's.
A smaller trap sits inside the fix: the paragraph that documents the phrase reproduced it, and the counter obligingly counted the documentation, reading five for four items. Describe the field without spelling it contiguously.
The third route was the simplest one and it defeated the repair: an item filed without the token
at all. RM235 went into ROADMAP.md on 2026-09-12 — high severity, a reproduced measurement under
it, **Status** 🔶 **open**, filed … — and the counters read 0/0 the next day against it. So the
rule from 2026-08-21 held for the shape of a status line and not for whether one was written at
all, and the second repair's own reasoning is what condemns it: a rule stated in prose is only as
good as the next person reading it, and the person filing an item is writing that line rather than
reading two documents about it. Fixed the way @registry-completeness says to fix it — as an
equality over a walked set, in schema/tests/test_triage_tools.py: every ## RMn between
# Active items and # Not format scope must be one the counter can see, and a third test runs
§4's own greps and asserts they agree with the walk, so the rule and its instrument cannot drift
apart. Run red against RM235's real line before the token was written, which is the only way to know
a guard guards. A prose rule that has been restated twice is asking for a test; this is the third
restatement and the last one.
- A marker can be stamped with a git commit sha, and it fails twice over. S36's read
<!-- triaged: 0.6.0 · sha cbeeb8f -->, which is a real commit in this repo and not a fingerprint at all.MARKER_REwants exactly twelve hex characters, so seven made the marker invisible: the section readunmarked-reply, indistinguishable from one answered before the ledger existed. The second failure is the one worth remembering, because it compounds silently — with no marker visible,reply_end()falls back to the single-paragraph rule, so paragraphs two onward of a long reply leaked back into the fingerprint. The value the ledger reported for S36 was therefore also wrong, and restamping to it left the sectionreviseda second time.
So "stamp the sha the ledger prints" is wrong for any reply longer than one paragraph, and that was
this file's own prescription until 2026-08-21. The contamination is not special to a malformed
marker — it is the normal state of a section you have just replied to and not yet stamped, which is
every section at the moment you go to stamp it. The value to write is the one the section carried
before the reply existed — the ledger prints it when the item arrives. This entry used to add
"and --backfill is correct for the same reason", which is false and is corrected in the entry
below: the tool computes the contaminated value too. Writing S61's reply reproduced it exactly — unmarked, the section hashed to
143952a318dc, and stamping that read revised (was 143952a318dc) against the true 9e026ea11585.
Checked in a scratch copy rather than argued, which is cheap and is the habit this file keeps asking
for. So: stamp the arrival fingerprint, then re-run the ledger and confirm current. The
confirmation is the whole check either way and it is two seconds — it is also what catches this if you
stamp the wrong value regardless, which is why the failure has never survived a pass that ran it.
--backfillis for replies older than the ledger, and stamps the wrong sha on a reply you just wrote — silently, and permanently. Adopted from the gist on 2026-08-21 and reproduced here against our own copy before being written down. It computes the fingerprint from the file as it stands, and with no marker present yetreply_end()falls back to the single-paragraph rule, so paragraphs two onward of a fresh multi-paragraph reply are hashed as if the consumer had written them. Step 3 warns against hand-copying that value; the tool computes it the same way, so reaching for--backfillto avoid the hazard walks straight into it.
One root cause, two symptoms, and which one you get depends on where the marker ends up. On a
three-paragraph fixture, arrival fingerprint 5c9ee39c5c6f:
- Hand-stamping the printed value, marker at the end of the reply where the Step 3 example puts
it: the marker terminates the reply properly, the next run recomputes over the consumer's prose
alone, and the section reads
revised (was 8fb615edf9ad)against text nobody edited. Loud, and self-announcing. --backfillitself:backfill()walks to the end of the**Statusparagraph — thewhile lines[end + 1].strip() != ""loop — and appends the marker there, not at the end of the reply.reply_end()therefore stops at paragraph one on the next run exactly as it did on this one, the contamination is stable, and the section readscurrent. Nothing ever reports it. The stored fingerprint permanently covers our own reply prose, and the first thing to surface it is a later edit to our paragraphs two onward arriving as a consumer revision.
So the reassuring advice — re-run the ledger and confirm current — is exactly the check that
passes in the worse case. It catches the hand-stamp and is blind to the tool. A stable wrong answer
outranks an unstable one, which is why the twelve-zero placeholder in Step 3 is the recipe rather
than either of these: with a marker already present the reply is excluded whole, and the ledger prints
the true arrival value for you. On the same fixture it printed 5c9ee39c5c6f back, matching what the
section hashed to before the reply was written.
- A section span is bounded by headings, and a document has furniture at both ends that is not a heading. Anything above the first item or below the last one belongs to an item as far as the span logic is concerned. Two instances, arrived from opposite directions, and the fingerprint check is blind to both for the same reason — fingerprints cover the consumer's prose, so anything that is neither reply nor prose travels wherever the span puts it, silently.
Furniture moving in (adopted from the gist, found in another repo): an inbox ending with a line
like *One item is open: S13.* hands that line to the last item, and the move verifies clean because
the footer sat inside the fingerprint on both sides. Repair: delete the footer from the history file,
re-run the ledger for the new value, hand-edit the marker — the one case where revised over
unchanged prose is expected, because the removed line was never the consumer's — then restore the
inbox's own footer, which the archived copy took with it.
Prose moving out (found here, 2026-08-21, and the more expensive half): S62's report contained a
fenced python block with # RebuildHints( written flush left inside it. That is the fence gotcha above, whose known
cost was a blinded ledger — but triage-archive.py shares BOUNDARY_RE, so the span it cut ended
there too. Half of S62 went to the history file, ending mid-block with an unclosed fence; the
commented expected-return, the closing fence and seven further paragraphs including the "Reproduced
against" attestation stayed in the live inbox, orphaned under the preamble, where they sat through a
commit and a release. The archiver reported "every fingerprint intact" and was right to. So the
fence gotcha is not only a ledger problem: it can split a report across two files, and neither the
verification nor the ledger can see it. Repaired in 51e6bba, verbatim and verified by byte
occurrence rather than by eye; S62 read current before and after at the same sha, which is the proof
that the ledger never saw either half.
Both are now guarded in code rather than described here. fence_findings() reports an unclosed
block and a ## Sn heading swallowed by one; the ledger prints them as STRUCTURE lines and
triage-archive.py refuses to move anything while either is true, before it writes. The historical
shapes are pinned in schema/tests/test_triage_tools.py, including a test that walks the old
boundary scan over S62's real shape and shows it stopping inside the fence — a guard nobody has
watched fail is a guess. What is still yours to notice: an archived group's last item leaves its
# heading behind with nothing under it, and each pass leaves the --- that preceded the section
it cut, so those accumulate one per pass — three had piled up before anyone looked.
7. State, and what the first full run found¶
As of 2026-08-11 the backlog is empty. S1–S18 all carry a reply and all sit in CONSUMER_SUGGESTIONS_HISTORY.md with a contents line each. The next consumer item is S19 and the next roadmap item is RM47 — but read both off the tools rather than off this sentence, which is exactly the kind that goes stale:
.claude/triage-state.py # the live inbox — empty means nothing owed
.claude/triage-state.py docs/CONSUMER_SUGGESTIONS_HISTORY.md # every answered item, all `current`
.claude/triage-state.py --next # the next unclaimed Sn, over BOTH files
An emptied inbox breaks id numbering unless the next id is pinned, and this is the one hazard the
split introduced. Once answered items move out, the live file's highest visible id is not the corpus's
highest — with the inbox empty it shows none at all, so the obvious next id is S1, which already
exists and already has a reply. Two defences, and keep both: the live file states the next id in a
heading, and --next computes it from both files so it cannot drift from them. Ids are never reused,
not even for an item answered as a non-issue — the reply is part of the record, and a recycled id would
collide with it. Same rule as RMn, and the same failure mode as reading the RM counter from memory
during a long pass.
The pass produced six code fixes, two roadmap items (RM45, RM46), four documentation fixes and three reasoned non-issues; the suite went 1382 → 1410 tests and every reference example still recompiles byte-identically. CHANGELOG.md's 0.5.4 entry is the per-item record. What is worth carrying into the next run is above, in the algorithm — Step 0 exists because of this run — plus these:
- The backlog was not what it looked like. Of eleven
newitems, two were already fixed, one had its preferred option shipped from another report, and three were non-issues. Half of an unanswered inbox can be answered without writing code, and none of that is discoverable without reading the source first. Budget the pass for establishing rather than for building. - Six items were reported by consumers who had already fixed their half, and each of those fixes is
evidence about the right shape: S15's
ServiceGate, S1'sstrip_registry_owned_keys(which we had upstreamed a release earlier), S16's reliance on sibling files, S12's search-instead-of-recall. Read what they built before deciding what to build; twice here the consumer's own argument against their first option was the reason the item was filed rather than patched (S10's per-article terms, S8's trust caveat). - The reply is the deliverable even when nothing is filed and nothing is fixed. A non-issue reply that cannot show what was probed is worthless, and a hedged one is worse than silence.
- Answering an item is not finishing it. RM43, RM44, RM45 and RM46 are open with S9, S13, S8 and S10
as their motivating cases.
RM_TOC.mdis the index for that half — this file's history is the record of what a consumer was told.