Skip to content

FAQ — questions that already have answers

What this is. A question-shaped index into decisions that are already made. Every entry is a question somebody actually asked — a consumer in CONSUMER_SUGGESTIONS, a dogfooding round, an audit, or one of our own sessions — together with the one-line answer and a pointer to the entry that holds the reasoning.

What this is not: a third index. RM_TOC.md is the complete list of every RMn and CONSUMER_SUGGESTIONS_HISTORY.md carries every answered Sn. Both are keyed by item. This file is keyed by question, which is the only thing a person arriving cold actually has — and the gap is not hypothetical: writing MODULE_LIFECYCLE.md re-derived S7 from scratch, with the same probe, because nothing connected "why did my digest move when nothing changed?" to an item called "fetched_at in the digest breaks find-by-hash".

Rules for this file, and they are the reason it can exist beside the other two.

  • One or two sentences, then a link. Never the reasoning. The linked entry is the one to edit; a second copy of an argument is exactly what made RM33 unfindable.
  • Only settled questions. An open item belongs in the roadmap, not here. If the answer is "it depends" or "not decided", leave it out.
  • A refusal is an answer, and usually the most useful kind — most of what follows is a repair somebody proposed that was checked and rejected for a reason worth knowing.
  • When a question keeps returning after being answered, say so in the entry. That is the signal that the answer is written down in the wrong place.

Identity, digests and signatures

Why did artifact.digest change when I did not change any data? Because the digest is the byte identity — these bytes, from this compiler — not the content identity. content_signature is the content one. A moved digest beside an unmoved signature is a provenance-only change and is the intended reading. → SCHEMAS § Identity & integrity, and S7, which was filed after somebody spent an afternoon looking for the content change that had not happened.

I reworded the reason on an overrides.csv row. Did my content_signature move? No, since 0.7 (RM180): reason, decided_by and decided_at are provenance beside the correction and sit outside the content identity, the way a README caveat does. Changing the value an override writes does move it. artifact.digest and the manifest.inputs entry for overrides.csv move either way, and so does the verification binding — a reworded reason still un-closes a module. → S87, and SCHEMAS § The authored overlay.

Then should the digest exclude the timestamp column, the way a build system excludes mtimes? No — unsound rather than unwanted. verify_manifest re-hashes every artifact.files[] entry from disk before recomputing the root, so a digest over anything but the shipped bytes is one no consumer can check. The mtime analogy misleads because an excluded mtime is not inside the artifact; this timestamp is a column in the parquet. → S7, which also rejects the two other proposed repairs.

Does a rebuild always mint a new digest? No. An untouched spec recompiled under a fixed compiler reproduces it exactly. What moves it across rebuilds is a toolchain change — parquet is not byte-deterministic across polars/arrow versions, which is why P4 scopes the guarantee to a fixed compiler_version. → CONSTITUTION P4, and MODULE_LIFECYCLE § 6.4.

Does re-running the enricher restamp fetched_at and churn my digest? No. Every sidecar merge is never-clobber, so a recorded row wins and its stamp is never rewritten; only deleting the sidecar re-stamps. → S7. (The name is wrong and the rename is planned for 1.0 — ROADMAP § the 1.0 cleanup.)

Which identity should a dedup or find-by-hash surface key on? content_signature, and just-dna-compiler signature <spec> computes it without compiling. Key a "these exact bytes" claim on artifact.digest. → S7.

Would recompiling my stored artifact produce different output than the one I am holding? Ask just_dna_format.release_records.needs_recompile(compiled_under, current) — the interval-keyed record of what each release changed about compiled output. It answers per axis and tri-state, so an interval it does not cover reads unknown rather than nothing changed — and so does a manifest that stamped no compiler_version at all (None or blank), since RM183; a stamp that is present but unreadable still raises, quoting all of it. → SCHEMAS § The release record.

Then does it tell me whether to rebuild? No, deliberately. The same fact costs a registry an immutable PATCH and a local cache a free rebuild, so the decision is yours and only the fact is ours — there is no should_rebuild. The declared correction-versus-addition flag is not that verdict: it says whether a value we published was wrong, which is upstream knowledge only this repo holds. → SCHEMAS § The release record.

My stored manifest.stats.genes disagrees with what I recompute from the authored rows — is that a bug? Not if compilation.dropped_rows is non-empty. compile_module re-derives stats over the survivors when the symbolic-allele drop removed something, so a recomputation from authored rows is the pre-drop side and the two disagree permanently under any compiler. That counter is the check. → SCHEMAS § The release record.

Is content_signature build-independent? No — reference-independent, not build-independent. Two modules with identical CSVs on different assemblies describe loci hundreds of bases apart, so a non-default genome_build feeds the hash. Every GRCh38 module's signature is unchanged by this. → RM36.

Does adding a new optional column break published modules? No, and the "digest window" argument that said otherwise expired in 0.4.1. An unset optional column is omitted from content_signature, so the authored identity is untouched; only a recompile's artifact.digest moves, which P4 already scopes. Removal, promotion to required, and retyping are the major-only moves. → CONSTITUTION P3, amended 2026-08-11.

Can I reorder rows in an authored CSV? Yes. It moves artifact.digest and leaves content_signature untouched. The counter-argument was probed and failed: an author reordering rows in their editor is already legal and already moves the digest, so forbidding a tool the same move proves too much. → ROADMAP, the row-placement decision.


Sidecars, re-runs and regeneration

Why can't the enricher just fill my empty weight cells from a GWAS effect? Because a null weight means the author has not modelled this, not nobody has computed this yet — MODULE_LIFECYCLE § Stage 3 lists weight/direction/effect_size among the cells no tool fills, and every check in the tier reports rather than repairs. There is also a sign trap: weight is documented positive=protective while a GWAS beta is positive on the effect allele, so a silent fill would invert the claim on exactly the rows nobody re-reads. What you want instead is gwas_effects.csv (RM90): the published effects sit beside the authored column with their units and their effect alleles, and your consumer picks one source or the other wholesale. Declare which, and on what scale, in module_spec.yaml's weighting: block (RM92). → SCHEMAS § The GWAS-effect table

Why is a per-row "use the GWAS value where weight is null" rule refused? It reintroduces the very thing the report was about. Two methodologies in one summable column is a number that means nothing, and a module whose weights vary in provenance row by row has no single scale left for weighting: to declare. Splitting the module in two was also considered and refused: the split criterion would be source coverage, not methodology (module B would be "the variants with no published GWAS"), and it would make module membership churn every time a new paper lands — routing an upstream fact straight into authored identity, which is what the derived-fact category exists to avoid.

Why did re-running a pass not pick up the newer data? Every derived sidecar is merge-not-clobber: an existing row is authoritative because a human may have overridden it. Delete the file to re-derive. → ENRICHER.md, and MODULE_LIFECYCLE § 6.3 for what deleting costs.

Can resolution.csv be published as a parquet like the other tables? No, and this is the first repair everyone proposes. It is a build-time lookup, not a published fact table: its provenance columns are outside the fact set by design, reverse_module cannot reconstruct half of them, and a consumer keying on it would be reading the lookup rather than the answer — which is materialized into weights.parquet and the positional tables where it belongs. → SCHEMAS § The resolution table.

Does reverse give me my spec back? No. It is a fixed point, not a backup. Manifest-only fields (authorship, provenance, logo, readme, and since 0.6 weighting), the whole verification attestation and its closure, and resolution.csv's provenance are all lost — deliberately, in each case. → COMPILER § Reverse and MODULE_LIFECYCLE § 6.9.

Reverse drops rsid_alternates — is that a bug to file? No, closed, do not re-flag. Reverse rebuilds the table from weights.parquet, which carries no provenance at all; those columns are outside the fact set precisely so they never reach the artifact, so the data does not exist for reverse to emit. Recovering them means re-running the enricher.

Why does a re-draft report differs instead of fixing the row? Because rewriting your value would destroy the evidence of the disagreement, and only you know which side is right. Drafting appends and never mutates; drift on existing rows is a cross-check's job to report. → ROADMAP § Parked in 0.5, where the one-word line between append and mutate is drawn.

Can the enricher just fix the authored cell it found wrong? Not in 0.x, and the reason is stronger than tidiness: content_signature is defined as pre-resolution and reference-independent, so if a network fetch could edit variants.csv that documented property would simply be false. Also unresolved: what such an edit does to authorship. → ROADMAP § Parked in 0.5, "Enricher co-authoring".


Validation, checks and trust

--strict passed, so the module is correct? No. strict means reproducible, never right. The compiler never fetches, so it has no reference to check your coordinates against; a whole file shifted by one base passes validate, strict compile, fully_resolved: true, and mints VRS ids that verify. → COMPILER § What the compiler can and cannot validate.

One exception since 0.7 (RM143), and it does not move that line. A strict compile refuses when verification.json records a finding on genome_build_agreement — the enricher having established that the rows are on a different assembly than the genome_build the module declares. That is internal consistency, one authored file contradicting another, not the compiler checking a coordinate: it acts on a record already written and adds no reference of its own. Every other recorded finding still only warns, because those are disagreements with an outside archive and the archive is the stale side often enough. No attestation means silence, findings=0 is a clean bill, and a skipped record — what --offline writes — is unknown.

Why does the ClinVar clin_sig cross-check warn even under --strict? Deliberate, and not an inconsistency to fix: failing a compile over a clinical disagreement would make the format arbitrate a clinical dispute, and a curator who read the primary literature and disagrees with a submission is doing their job. Same reasoning for the allele-function check and the article-licence warning. → ROADMAP § Parked in 0.5.

Should it escalate when the disagreement is with an expert panel rather than a lone submitter? Tempting and not taken, for the same reason. The confidence is surfaced (ClinSigFinding.confidence) — surface it, let the consumer route on it, do not decide for them. → ROADMAP § Parked in 0.5.

Why can't validate stamp the closure when everything passes? Because a record stamped by whatever happened to execute says only someone ran a tool, which is the exact defect the closure exists to fix. Closing is a deliberate act; validate stays read-only. → SCHEMAS § The closure.

fully_resolved is true — is the module trustworthy? Not on its own. The flag is all() over variants.csv, so it is vacuously true for a module with no variants.csv. The safe rule is resolution_subjects > 0 and (resolution_mode == "strict" or fully_resolved). A catalog followed the old rule and had to migrate a stored projection. → RM44 / S13.

Why is a blank cell not false? Three-valued is the house algebra: true / false / unknown, and None is never False. Combine with Kleene semantics rather than withholding on any unknown, because unknown AND false really is false. → CONSTITUTION and the tri-state rules throughout SCHEMAS.

A check reported zero findings — does that mean it passed? Only if it ran. A check that could not run reports why (clin_sig_not_checked, gene_loci_not_checked, verification.json's skipped key), because an empty finding list otherwise says both "compared everything" and "never compared". → ENRICHER.md.

The manifest says quotes_found: 0, quotes_unchecked: 0 — were the quotes checked and missed? Not if only the abstract was searched: an abstract miss settles nothing, and quotes_unchecked counts only citations with no text at all. abstract_only_count > 0 beside quotes_found < quotes_authored is the case to treat as unchecked; per citation, quote_source = abstract beside quotes_found = 0 in literature.csv means unsettled. A counter that settles it outright is a minor's, on the 0.8 branch. → RM264 / RM256 / S109.


Schema shape — repairs that were checked and rejected

Can alts accept IUPAC ambiguity codes, or can Y be expanded to C,T? No to both. No ref/alt/alts column has a nucleotide grammar (eleven columns, six models), so adding one would reject N too and break P3 for modules that compile today. And the expansion reading has no instantiation: probed across 4,439,382 ClinVar rows and all sixteen modules, R/Y/S/W/K/M/B/D/H/V appear in REF or ALT zero times. → RM5 and the 0.6 idea-book probe.

Should the compiler fill a blank direction from state? No. direction_from_state is sound as a consumer's read-time fallback and a fabricated fact in a published table — state='significant' names no direction at all. A state-only module correctly ships an empty direction column. → COMPILER § Upgrade derivation.

Should the compiler materialize a gene panel declared in module_spec.yaml? Dead, not deferred. The compiler must not create rows no curator wrote, and expansion at compile leaves reverse choosing between the declaration and the rows — neither a fixed point. Drafting writes those rows as authored bytes instead. → RM4; panel: is deprecated in 0.6, removed at 1.0.

Can there be a --non-commercial compile flag? No — charter-illegal. A flag cannot be recorded in the artifact, and reverse_module rebuilds module_spec.yaml from parquet alone, so compile → reverse → compile would refuse on the third step (P7). The declaration has to be data. → ROADMAP / RM21.

Should _check_misspelled_tables just search any subdirectory? No. Tolerating a location without extending the guard puts a typo'd derived/varaints.csv exactly where the check written to catch it cannot see. An authored name under derived/ is reported as misplaced by exact match; everything else is fuzzy-matched against the full known set. → RM49.

Two annotations.parquet rows with the same key collapse — isn't that a P7 loss? No, and this is the standing example of a mechanically-possible finding with no real instantiation. Sharing a variant_key forces a single locus (a one-to-many rsid is expanded to distinct keys), hence one gene; identical conclusion+negatives means the same effect, hence the same phenotype. The constraint set is empty. The genuine poly-effect case is what the keying already fixes. → RM80 / S29; and see Working practice below.


Naming and renames

Can we rename a column / a file / a parameter? A rename is a removal plus an addition, so it is major-only under P3 whatever the thing is. What is legal in a minor is landing the new name as an accepted alias first, so the major only has to remove — that is RM51 (sources.csv → licensing.csv), and it is the pattern to copy. → RM51, and the 1.0 cleanup tracker for what is queued.

Then why not rename resolve_with_ensembl, which everyone agrees is misnamed? Adding a differently-named alias would be legal and additive; it was declined because the only honest alias is --no-resolution, which buys a better name at the cost of two flags meaning one thing. The rename itself waits for 1.0. → S14, answered with a refusal.

Can a field's meaning be corrected in place? No — add the corrected field beside it. Finding.line was added next to Finding.row rather than redefining row, because a consumer already compensating for the old meaning would have started reporting line 4 for line 3 with no signal. → S18.


Scope boundaries

Will the format ever carry per-sample results, coverage, or a report card? No — a measurement is the consumer's, by the data-agnostic charter. RM7 (the evaluation-output schema) is listed in the roadmap only so it is not mistaken for format scope. Callability is expressed as pointers into a VCF the format never sees. → ROADMAP § Not format scope, and SCHEMAS § the consumer join contract.

Can a module contain an expression, a script, or a predicate? Not code. A module is data (P1). The sanctioned escapes if tables are ever outgrown are a non-Turing-complete boolean predicate over genotypes and declarative pattern grammars such as regular expressions — neither is needed yet, and both are escape hatches rather than defaults. → CONSTITUTION P1.

Can the compiler check my coordinates against the reference? No, and it never will: format and compiler never fetch (P2). A check that needs a reference belongs in the enricher. → COMPILER § the inescapable blind spots.

Should there be an offline gnomAD frequency snapshot, like ClinVar's? No — v4.1's sites VCFs are 58 GB (exomes) and 742 GB (genomes), so there is no slice to ship. Frequency is the one online-only link, and that is not a reproducibility hole because frequencies.csv is the pin once written. → ROADMAP § Parked in 0.5.


Working practice

Is start 0-based? No. Every start in this codebase is the 1-based VCF position. The docstring that said otherwise was the bug, and it shifted 3,038 variants across four modules past every offline gate. Pinned by schema/tests/test_coordinate_convention.py. → CLAUDE.md.

Do I need to file a round-trip or dedup loss I can construct mechanically? Only after you have built a real, sensible example against the actual code paths. A mechanically-possible loss with no real instantiation is noise — see the annotations.parquet entry above, where the constraint set turns out to be empty.

enrich said an authority call under one of my overrides.csv answers has moved. Was my answer wrong? Nothing in that message says so, and nothing in this format can. It is a statement about the record: the reason you wrote was written about a particular disagreement, and the archive now publishes a different value, so the reason describes a disagreement that is no longer the one on file. Often the reason still says exactly what you mean; sometimes the archive has moved toward you and the row is now unnecessary; occasionally it has moved somewhere you would argue with differently. All three are yours to decide. It warns in both modes and escalates in neither — see ENRICHER § Has the disagreement you answered moved?.

Why did that message not come back on the next run? Because it cannot. The baseline is the previous run's clin_sig_authority_calls.csv, and the run that reports the move is the run that replaces it — so the next run compares against the new value and finds nothing. That is a limit of an observation: saying it every run would need your overlay row bound to the value it justifies, which is a column this format does not have and a change to the authored surface. Act on the message when you see it.

Can I get the same check for frequencies.csv or resolution.csv? No, and the reason is worth knowing rather than working around. A currency check needs a value recorded at the time to compare against, and only two tables carry one: clin_sig_authority_calls.csv (derived — the clin_sig, the verbatim token, and the release it came from) and, since RM160, studies.csv's authored confidence/confidence_unit pair (the citing source's own review state, re-asked by evidence_status_currency). Every other sidecar holds the source's current answer, so there is no prior value and a general "the value moved" check would be answerable for two tables and silently absent for the rest. Each finding names its table for exactly that reason (RM160).

A message cites an RMn — is that a bug? No, it means known and deliberate. Leave the data honest and note the limitation rather than inventing a workaround. → RM_TOC.md for what any given number is.

I edited module_spec.yaml and the compile says my attestation is stale — is that a bug? No, it is the design: verification.json's module_hash binds the authored bytes, so any authored edit un-closes the module and drops the block, deliberately and for free. Note the reach — appending an authorship: entry counts, though it moves no identity. What no longer counts is a change of line endings: since 0.6 the binding reads \r\n as \n (RM82), because an editor or core.autocrlf rewriting them is not an edit anybody made. It stops at newlines — a BOM, trailing whitespace and a missing final newline are still edits. → SCHEMAS § the closure and MODULE_LIFECYCLE § 6.2, where both are measured.

Can the binding exclude the display metadata, so fixing a card subtitle does not un-close my module? No, and the reason is the partition rather than the cost. Splitting it along content_signature's line looks right and is not: that line excludes name, version and namespace too, so a binding drawn there lets a closure survive a rename — the reviewer's attestation travels to a module with a different identity. content_signature excludes them so a registry strip does not move content identity, which is correct for a dedup key and exactly wrong for an attestation. The binding answers is this the same document a named person signed off, which is why it is coarse. The fix for an already-published subtitle is registry-owned metadata stored beside the module (the IDENTITY_AUTHORITY_KEYS shape), not a smaller binding. → ROADMAP § RM133.

Why does the binding cover only the authored files and not the sidecars? Because the derived sidecars carry a fetched_at per row, so binding to them would perish the attestation on a re-enrichment that changed nothing anyone claimed. Read source currency off each record's own release, never off the binding. → SCHEMAS § the verification attestation.

MITOMAP marks a variant [VUS*] — why is that withheld instead of drafted as uncertain_significance? Because the withhold has to happen before the normalizer: normalize_clin_sig's own default is other, a definite member, so an unmapped token falling through would have become a confident call. The bracket's five VCEP abbreviations are mapped; VUS* is MITOMAP's own asterisk and the row is counted, never drafted (RM171).

mitomap miss says built=None naming ClinVar — why not False, and why not an empty result? A derived lane's absent parent is could-not-run, not failed and not "no misses": False would file another lane's absence as this one breaking, and an empty miss is the strongest possible claim, from a comparison that never ran. The reason names the parent it lacks (RM171; the tri-state is RM176's).

A MITOMAP-drafted module will not compile until I write every genotype — why does ClinVar's fill and this one not? The stub is not about the contig. On chrMT the compiler fills the sole expressible genotype for a ClinVar row because that record is a claim about an allele; MITOMAP's is a claim about a literature corpus, some of it reported only heteroplasmically, so the cell is a curator's decision and the placeholder protects it (RM171).

CIViC says two variants are pathogenic together — why is that not a HaplotypeRow, and why not just drop the row? The source's own description says heterozygous compound mutation: the two are in trans, and a haplotype is cis, so the brick would assert the opposite of what was observed. Dropping the row was refused because it deletes two of RM170's three subjects; the row stays, keyed on the variant's own profile, with the composite profile beside it. The representation is RM28's, parked (history/ROADMAP_HISTORY_0_6.md#rm174--a-claim-about-two-variants-in-trans-is-written-as-two-single-variant-rows-because-no-brick-holds-the-real-subject)).

ClinPGx's old clinicalAnnotations.zip still downloads — why does clinpgx build refuse it instead of reading both? Because a retired filename that still answers 200 is a frozen 2025 object, and a reader that parses both vintages is a reader that can still publish 2025 data. The builder refuses the old member names and says which archive to fetch (RM175).

check-repeat-bands reports STRchive's pathogenic_max but never writes it — is that a gap? No. An upper bound on pathogenicity is not a claim the band table makes, and the check reports and never repairs; the catalogue's own boundaries are compared, and the one it lacks (FMR1's premutation threshold) is a disagreement the author keeps (RM165).

cache rebuild printed not run for some lanes — did it fail? No. The outcome is three-valued: ACMG needs an Elsevier workbook, PharmVar a personal key, CIViC a pinned release date, Ensembl is built elsewhere, and a derived lane without its parents names them. Each prints the reason from the registry field, and the exit code counts only real failures (RM176).

check_declared_use(PUBMIND_TERMS, …) skips under every declared use — is the PubMind drafter unreachable? No. That function gates a fetch, and PubMind is never fetched: the operator builds the snapshot with pubmind build, and draft_gene_panel_from_pubmind reads it without consulting the gate, warning in the source's own words that every licence cell on the pubmind row is null (@acquisition-gate-is-not-a-read-gate; a test pins that the draft is not skipped). The terms are None on every axis because they are genuinely unsettleable, not unrecorded — the software licence, the paper's CC BY-NC-ND and the coordinate table's silence are three statements and none is about the bytes (PUBMIND_ASSESSMENT). What the unknown answer governs is publishing a module that carries those values, RM27's undesigned axis. A wrapper that runs the gate itself before calling the drafter is what makes it unreachable (S99).

The consumer-suggestions inbox is empty — were my notes lost? No. An answered item moves byte-for-byte to CONSUMER_SUGGESTIONS_HISTORY.md with a row in its index. An empty live file means nothing is owed. And never read the next free Sn off a written-down number; run .claude/triage-state.py --next.