Skip to content

Use cases & dogfooding — feasibility and blocker analysis

For each real or desired use case: is it enabled by the format as it stands, and if not, what is missing? (The verdicts were first written against 0.4 and are revised in place as items ship; the tree is at 0.7.) The point is to separate three things that get conflated —

  • work the format already enables (author it today),
  • work that is a consumer concern the format deliberately does not own (the data-agnostic north star — see CLAUDE.md), so there is nothing to add, and
  • genuine gaps that need an additive format change.

Every gap below is tagged → RMn and collected in Roadmap items surfaced at the end; those are the items to migrate into ROADMAP.md.

The feedback → schema cycle (where this doc sits)

The design docs are stages of one loop; an idea moves left-to-right as it matures:

  1. Feedback — a consumer's field report. → CONSUMER_SUGGESTIONS.md (open) / CONSUMER_SUGGESTIONS_HISTORY.md (answered), the runbook in CONSUMER_TRIAGE_LOOP.md. The pre-Sn rounds 1 and 2 were their own threads and are both retired — their dispositions are recorded in history/CONSUMER_SUGGESTIONS_HISTORY_PRE_0_6.md
  2. Usage → blockers → solvability — run each use case against the current bricks: is it enabled, consumer-side, or a gap; and is the gap closable additively? → this doc
  3. Means → draft schema → decision — for a gap worth closing, the proposed shape + a charter check
  4. the open questions to settle it. → the PROPOSAL_* thread for the release under proposals/ (eight so far, 0.4.1 through 0.7 PT3, every one now a record). (Shipped decisions → CHANGELOG.md.)
  5. Conclusion — "how to do it now, with these bricks" — the distilled worked example once the shape is settled. → REFERENCE_EXAMPLES.md
  6. Terminal, one of two:
  7. Fixed — schema + compiler shipped (the models; COMPILER.md marks it validated/materialized), or
  8. Deferred — a recognised gap parked as a roadmap item (RMn → ROADMAP.md) when the means aren't worth building yet.

So this doc and REFERENCE_EXAMPLES.md are the same use cases at two points in the loop — here they are questions (what blocks?), there they are answers (author it like this). An ENABLED / schema-ready row here graduates to a REFERENCE_EXAMPLES entry; a GAP row graduates to the release's proposal thread (if being closed now) or an RMn roadmap item (if deferred). The loop is why a "blocker" is never a dead end: it is either dissolved (it was consumer-side all along), closed additively, or explicitly parked.

Verdict legend - ENABLED — authorable/usable now, no format change. - CONSUMER-SIDE — a consumer (runner/app) feature; the format correctly owns nothing here (per the north star). No format change — often no blocker at all. - SCHEMA-READY / COMPILER-PENDING — the schema models express it; only the deferred compiler materialization (new parquet + round-trip) is missing. One known gap, not per-use-case. - GAP — needs an additive format addition (the interesting rows).

A recurring result: most "can the format do X?" questions dissolve into CONSUMER-SIDE, because the module is declarative annotation and the doing is a consumer's. The format earns its keep by the properties it froze (declarative-not-code, integrity-as-identity, the unresolved/callability contract), which are what make a consumer's X safe and reproducible — not by hosting X.


1. The consumer's 0.5 suggestions, each run through the lens

1a. Verification harness — run a module against N VCFs, emit report-card diffs (round-2 §3b)

Verdict: CONSUMER-SIDE — no blocker, no format deliverable. This is the headline dogfooding use case and it is already enabled. Walk the requirements:

  • A panel is a set of loci + expected genotypes/bins → already a module (a variants.csv of genotype→conclusion rows, or a binning table of measure→phenotype). No new "panel type".
  • Deterministic extraction of the observed value from a VCF → the consumer's job, and source_field (0.4) now names the exact VCF FORMAT/INFO field to read, removing the last glue.
  • Trustworthy before/after (per-caller, ±liftover) diffs → artifact.digest already gives the module a content identity, so a report-card diff is a byte-level diff of two deterministic runs.
  • No-call ≠ mismatch → the mandatory unresolved outcome + the callability contract already stop a "no-call under caller B" from masquerading as a "mismatch vs caller A".

The one thing that looked like a format artifact — a standardized evaluation-output / report-card schema ({locus, observed, callability, bin_selected, verdict}) — is per-sample results, i.e. a measurement, so the north star keeps it consumer-side. It belongs in just-dna-lite, not just-dna-format. Nothing to schedule in the format. → RM7 records the consumer-side schema so it is not mistaken for a format item.

1b. Augmented-VCF as the landing pad for cracked short-read loci (round-2 §3c)

Verdict: ENABLED at the format boundary; emission is CONSUMER-SIDE. A caller emits its niche genotype (PER3 span, DAT1 motif-path, MAOA half-repeat) as a synthetic VCF record (<STR> with INFO/RU, FORMAT/REPCN, custom evidence fields); a repeat_alleles.csv module consumes it via the same source_field=REPCN path as ExpansionHunter. The format does not invent a representation — it binds to the VCF one. Producing the augmented VCF is the caller's job (consumer-side). The only format touch-point (source_field) shipped in 0.4. Consuming symbolic alleles (<STR n>, <CNV>) at the count/dosage layer is enabled via the binning tables, and since 0.6 a symbolic allele is expressible inside a VariantRow.genotype too — VCF's five types with the length in the token, <CNV:TR:30> included (see §3b; RM5, shipped).

1c. Callability three-state, phasing-aware panels, trio/multi-sample (round-2 §3d)

  • Callability three-state (covered-hom-ref vs no-call): CONSUMER-SIDE, derivable from VCF DP/GQ/FT (or a gVCF ref-block). The format's part: requires_callable (reserved flag) marks rows where absence is informative; promoting it to a typed boolean column and reserving callable_from (the DP,GQ,FT signal) are the format-side follow-ups → RM6. Both shipped — requires_callable as a tri-state VariantRow column in 0.4, callable_from in 0.5 — and 0.7 widened the first to haplotypes.csv and pharm_variants.csv, the other two tables that name a locus, so a star-allele module can state the assumption its sources make in prose → RM70.
  • Phasing-aware panels: ENABLED — the phased flag + the phased genotype form A|G (0.3 item 5b) already let a runner do cis/trans for compound-het and star-allele phasing (*2x2/*4 vs *2/*4x2). No gap.
  • Trio / multi-sample (Mendelian / de-novo assertions): CONSUMER-SIDE — VCF is natively multi-sample; the assertion runner is a consumer. An optional declarative inheritance-expectation field on a panel row would let the module carry the assertion as data rather than consumer lore — small, additive, optional → RM10 (only if a real module needs it).

1d. Authoring-support suggestions (from the just-dna-agents integration, ROADMAP obs 2026-07-10)

  • A canonical machine/LLM-facing authoring reference. Consumers (MCP servers, agents, docs) hard-code prose summaries of the DSL that drift from the real schema. ADOPTED (RM8, shipped in the 0.4 sample): just_dna_format.reference.authoring_reference() returns a JSON-serialisable summary — every model's field list + all vocabularies + reserved names + the palette — generated from the live models, so it cannot drift; json_schemas() gives the full JSON Schema. A consumer's get_spec_format renders this instead of a hand-maintained blob.
  • A recommended icon/color palette. Display validates icon_set/color but shipped no recommended enumerated palette, so each authoring tool invented one. ADOPTED (RM9, shipped in the 0.4 sample): manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS (curated semantic-use → value maps, recommendation-only — not enforced), surfaced through authoring_reference().

2. Reference-data-backed module types

2a. ClinVar gene-panel (flag pathogenic variants in a gene set)

Verdict: ENABLED — and the materialization moved off the compiler (RM4, closed). The surface is just-dna-enricher draft-panel --source clinvar --gene …: the enricher drafts the matching ClinVar rows into variants.csv from a pinned snapshot, and the author's no-op over the drafted subset is the authorial act, so the compiler never resolves a panel and never needs a reference mixin. Worked end to end in reference_examples/hfe_hemochromatosis (REFERENCE_EXAMPLES §9b). The panel: block GenePanelSpec declared is deprecated, removal at 1.0. The earlier verdict here — schema-ready, gap on native compile-time materialization gated on a content-pinned ClinVar mixin — was the 0.4 reading; what closed it was deciding that a drafted, curated row is the right subject, not a predicate the compiler evaluates.

2b. PharmGKB drug-response annotation (item 9)

Verdict: ADOPTED (RM3, shipped in the 0.4 sample). A PharmGKB row maps a variant/diplotype → a drug + a response/phenotype + a PharmGKB evidence level (1A…4, VALID_EVIDENCE_LEVELS) — a different axis from a risk weight. Built as a dedicated PharmVariantRow (pharm_variants.csv) for single-variant drug response (keeps the SNP core clean — one CSV = one concern), plus optional drug/response/evidence_level columns on DiplotypeRow for the diplotype-keyed case. A PharmGKB module has no empty variants.csv. evidence_level is a third significance-flavoured axis, distinct from stat_significance/clin_sig (orthogonal-axes discipline, Principle 5). Materialization deferred with the rest of 0.4.

One gap in it closed in 0.7 (RM132): the table could make a clinical claim per row and cite only per variant. A ClinPGx-drafted module carried 1,482 drug-response rows with nowhere to ground any of them, because studies.csv keys on (variant_key, pmid) while these rows key on drug, genotype and category too — so one study row would ground every claim recorded for that variant. PharmVariantRow now carries its own optional pmid under the rule RM47 settled for the bins: a row cites when its claim is finer-grained than a study row's, and the citation table describes. evidence_level was never the handle — it grades evidence rather than pointing at it — and provenance_quote deliberately did not follow.


2d. GWAS effect sizes as grounding for a curated panel (0.6)

Verdict: ENABLED (RM90 shipped), and the interesting part is what it refused. A curator with a trait panel wants the published effect sizes beside their own annotations — partly as evidence, partly because a consumer reported that hand-set weight values "construct nonsense" across a corpus and that GWAS effects are often better grounded (S36).

The enabled shape: just-dna-enricher gwas fills gwas_effects.csv, one row per published association, and the module ships it beside variants.csv. A consumer joins on variant_key and reads effect_size with effect_unit and effect_allele.

The blocker that was not dissolved, stated so nobody re-proposes it. Filling weight from those effects is barred (MODULE_LIFECYCLE § Stage 3), and a per-row precedence rule was refused as putting two methodologies in one summable column. What closed the gap additively was two tables and a declaration, not an exception to the rule.

And a caution the real data supplies better than the design note could. rs1800562's 186 associations span 62 EFO traits in 12 distinct effect units — three of them spellings of one unit, two more differing only in case — with 138 rows in the Catalog's uninformative unit and 42 of 195 naming no effect allele at all. A "GWAS effects are better than curator weights" pipeline that pools those is worse than the weights it replaces. Read them per trait, and read manifest.gwas_effects.units before pooling anything.

2e. Republishing a licence-encumbered third-party corpus as modules (2026-09-13)

Verdict: ENABLED, and this is the case the licence machinery was built for — with one condition that lands on the marketplace rather than on the format, and one that shapes how the modules are cut.

The case, named by the maintainer. Take a large existing annotation corpus somebody else curated, rework it, and publish it as modules. The concrete instance is SNPedia: 106,603 SNP entries plus its genosets, reachable as a MediaWiki dump (the extraction is already done in zhaofengli/snappy — see ROADMAP_0_8 § RM188). This is a different shape from 2a and 2b: those draft rows from a source into an author's own module, where this republishes somebody's whole corpus under terms that travel with it.

The terms are established, and that is the first thing that had to be true. Read 2026-09-13 off three primary pages, all naming the same licence in the same words:

Page Says
SNPedia:Copyrights "The content in SNPedia is licensed under a Creative Commons Attribution-Noncommercial-Share Alike 3.0 United States License."
SNPedia:General_disclaimer the same sentence, plus "For more details, see our Terms of Use."
Bulk the same sentence, and it invites bulk access — a bots. subdomain, code samples in four languages, a published GFF

Contrast this with the HPO refusal (ENRICHER.md), because the two look alike and are not. HPO's licence URL 404s and OBO records a bare label with no SPDX id, so nothing machine-readable establishes the terms and an unestablished permission is not a permission. SNPedia names a versioned, SPDX-identifiable licence (CC-BY-NC-SA-3.0-US) on three pages. Its SNPedia:Terms_of_Use page is a 404 — the elaboration is missing — but the grant is not the elaboration, and @no-named-licence is about a source that names none. This one names one.

Two acquisition notes. www.snpedia.com is behind Incapsula and answers a 212-byte challenge stub to curl; bots.snpedia.com serves the same pages plainly, which is the sanctioned route and is what the Bulk page tells you to use. Its Forbidden section bans two scraping patterns — every version of every page, and every possible rs# — and says nothing about reuse; the dump route avoids both.

What the three letters cost, in this format's own fields.

Term The cell The consequence
NC commercial_use=False, declared_use=non_commercial The compile gate refuses a module declaring commercial use. This is @gate-is-data-driven doing its job, not an obstacle.
SA share_alike=True The module itself goes out under CC BY-NC-SA 3.0 US, and share-alike is viral over whatever it is mixed with. So a SNPedia corpus is its own module, never rows blended into a general one — which the one CSV = one concern rule already wanted.
BY attribution, license_url Credit as designated, the URI, and an indication that it was modified — "after rework" is a modification and the licence requires saying so.

redistribution=True: CC BY-NC-SA permits it. Recorded, not gated — @redistribution-ungated, RM27.

The one condition that is not the format's to answer. NC in CC 3.0 bars use "primarily intended for or directed toward commercial advantage or private monetary compensation". Whether just-dna-registry distributing an NC module is itself commercial use is a real question, and it decides whether this use case works end to end — a free catalog is a different answer from a paid one, and the answer may differ again for a commercially-operated free one. That belongs to the marketplace, and it should be settled there before the first NC corpus is published, not after.

Two things to check before building, neither of which blocks the design. SNPedia has been owned by MyHeritage since 2019, and their announcement said it would remain free "for academic and non-profit use" — phrasing narrower than the CC grant, which permits any non-commercial use. A CC licence is irrevocable for the versions published under it, so the practical answer is to pin the dump revision you took, exactly as every other source here is pinned. Separately: a genoset's boolean structure is arguably a fact and its summary prose is plainly expression, so "rework" launders neither — and stripping the prose to dodge share-alike would throw away most of what makes the corpus worth having. Neither point is a legal opinion, and neither should be treated as one.

What this case actually proves. Every earlier entry in this section takes a permissive source and produces a sellable module. This is the first that takes an encumbered one and produces an honestly-labelled unsellable module — and the format holds it without a new column, a new flag, or a mode. That is the whole argument for licensing-as-data over a --non-commercial switch (@licensing-as-data, @gate-is-data-driven), with a real corpus standing on it instead of a hypothetical.

3. Composite modules (the real pipeline shapes)

3a. SNP + PRS in one module

Verdict: ENABLED (RM1 + RM2 shipped). A module is a directory of CSVs carrying both variants.csv (VariantRow) and pgs.csv (PgsRow), joined on the shared trait_efo_id (item 5) so a variant panel and its PRS companion sit in one content-addressed unit. The compiler now materializes every present table kind to parquet (round-trip lossless) and treats variants.csv as optional, so composed and single-domain modules both compile. No blocker.

0.5 revisit: the 0.4 shape was modelled at the wrong grain, and real data proved it. RM3 was declared shipped against a hand-authored sample. Run against the actual ClinPGx corpus it does not hold, in two stages:

  1. A clinical annotation is published per genotype — the summary table names the variant and drug, a child table gives one row per call, and the large majority carry exactly three. PharmVariantRow had no genotype, and the compiler deduped on (variant_key, drug), so authoring the real SLCO1B1/simvastatin annotation produced duplicate row for key ('rs4149056', 'simvastatin'). ~97% of the corpus was unauthorable.
  2. Adding genotype was not enough. One variant and one drug carry several distinct annotations — rs4149056 + simvastatin is Metabolism/PK at 1A, Efficacy at 3 and Toxicity at 1A. 1,199 of 17,380 triples collide; 839 separate by phenotype category, 283 by neither category nor level.

Closed additively with genotype, phenotype_category (closed vocabulary) and annotation_id (a source accession as identity, like PgsRow.pgs_id) → RM20. The lesson is the dogfood rule in CLAUDE.md read from the other side: a shape validated against a sample rather than a corpus is not validated. See reference_examples/pgx_slco1b1_simvastatin/.

2c. Star alleles and drug response from the live authorities (0.5)

Verdict: ENABLED, with the licensing made legible (RM21). The enricher can now cross-check a module's allele_function.csv against PharmVar and CPIC and its pharm_variants.csv against a ClinPGx snapshot, and resolution reaches pharm_variants.csv/haplotypes.csv so a PGx module gets coordinates without carrying a variants.csv.

The blocker that turned up was not technical. Every pharmacogenomics upstream is copyleft and none is sellable: ClinPGx, CPIC and PharmVar are each CC BY-SA 4.0 plus a separate contractual bar on sale. api.pharmgkb.org was retired 2026-07-20, and CPIC sits inside the ClinPGx merger with its licence page redirecting to the ClinPGx policy — so swapping sources does not escape the terms. That is consumer-relevant rather than a format gap, but it is unrepresentable in a 0.4 module, so it was closed additively as sources.csv (RM21) rather than left to a README nobody can query.

Composition principle (settled during the PharmGKB decision, now in CLAUDE.md): a module composes from optional table kinds — one CSV = one concern — so the SNP core (variants.csv+studies.csv) stays minimal and no module ever carries an empty variants.csv or a foreign domain's columns just to host one table. This is the human-authorable half of the RM2 work.

3b. SNP + indels

Verdict: ENABLED for small ACGT indels; GAP for structural/symbolic. VariantRow alleles are ^[ACGT]+$ multi-base, so a small insertion/deletion is expressible today (ref=A, alts=AT, genotype A/AT) on the same variants.csv as SNPs — a mixed SNP+indel panel is authorable now. Symbolic/large structural alleles shipped in 0.6 (RM5): VCF's five closed first-level types with the length inside the token (<DEL:4977>, <CNV:TR:30>), worked on a real module in reference_examples/mt_common_deletion. A lengthless token is dropped with a warning and refused under strict. Dosage and count still route through the copy-number / repeat binning tables; the earlier "recognised gap" here is closed, and what stays unexpressible is deliberate (CPIC's IUPAC codes).

3c. SNP + PRS + PGx + CNV in one "personal panel"

Verdict: ENABLED (RM1 + RM2 shipped). The generalization of 3a: a personal/curated module mixing variants.csv, pgs.csv, activity_phenotype.csv/diplotypes.csv, and copynumbers.csv, all joined on trait_efo_id, compiles today — each present kind materializes to parquet with round-trip. This is exactly the "personal module re-checked deterministically on every pipeline change" the verification harness (§1a) wants — and RM1/RM2 are what unlocked it.


4. Network-first validation & enrichment (external-source scrutiny)

4a. just-module-validator — deterministic source-checks + provenance enrichment against public sources

Verdict: CONSUMER-SIDE of the format and compiler (Principle 2 keeps it out of those two tiers) — and the sibling it proposed exists: it is just-dna-enricher (RM13, realized). The two additive format anchors it needed shipped in 0.4 (RM11/RM12), and one requiredness fix waits for 1.0.

The sibling library as proposed, network-first: given a module it checks the authored claims against public sources and enriches them — every one of these is now an enricher check with a VALID_VERIFICATION_CHECKS member (citation_existence, rsid_currency and rsid_coordinate_agreement, citation_identifier, provenance_quote) —

  • validate every pmid resolves in PubMed, and every rsid resolves in dbSNP at the authored chrom:start (flag coord/liftover drift);
  • cross-fill provenance ids — derive a doi from a PMID and vice-versa;
  • confirm a study's claim actually appears in the cited article's fulltext (imagine further source-checks in the same spirit).

By the data-agnostic north star and Principle 2 (no network; inject-only), the doing — every fetch and lookup — is a consumer's, and can never live in just-dna-format/just-dna-compiler (that would pull the network dependency the tiers forbid, Goal 2). So the validator is a new consumer/enricher sibling to just-dna-lite, recorded as → RM13 so it is not mistaken for format scope — exactly as the report-card harness is (§1a / RM7). Crucially, most of what it checks needs nothing from the format: rsid, chrom, start already exist, so validating them against dbSNP is pure consumer work — enabled today. Two things it wants to anchor are genuine additive format gaps, and one is a 1.0 requiredness fix:

  • doi as a provenance id — additive, shipped in 0.4 (RM11). StudyRow previously carried only pmid (required, and it must contain ≥1 real PubMed id). DOI is wider: it covers preprints (bioRxiv/medRxiv), books, theses, and datasets that have no PMID. The optional doi column now lets the validator record and cross-fill it, and lets a module cite a DOI-bearing source. Purely additive → P3/P8 clean (new optional field; existing data still validates); validated against the DOI grammar and kept verbatim. → RM11.

  • A provenance locator — search-phrase/regex pointing at the passage in fulltext — additive, shipped in 0.4 (RM12). So the validator can answer "does the cited article's fulltext actually contain this claim?" in a yes/no manner, a study row now carries optional provenance_quote (keyword phrase) and provenance_regex. The regex sits squarely inside Principle 1's sanctioned escape hatch: a declarative pattern grammar is data, not code — the module ships the pattern, the consumer supplies the fulltext and runs the match, evaluated by a linear-time / ReDoS-safe engine (P1's explicit requirement; the compiler only re.compile-checks it at author time). It is the provenance analogue of source_field (0.4): source_field is a declarative pointer to where the measurement lives in a VCF; the locator is a declarative pointer to where the claim lives in the article. Neither holds the data it points at (north star ✓). Primarily an aid for LLM-authors (which can emit a precise pattern), yet a plain keyword phrase is legible enough to clear the human-authorability gate for a human author too. → RM12.

  • pmid is mandatory today — the DOI-only case cannot be closed additively (a 1.0 fix). pmid: str is required and must parse to a real PubMed id (extract_pmids), so a preprint/book/thesis with only a DOI is unauthorable right now — and demoting a required field to optional is precisely the move Principle 8 forbids within a major. Adding doi (RM11) is necessary but not sufficient: while pmid stays required-and-PMID-shaped, DOI-only provenance is still rejected. The full fix is doi-first at 1.0 — make pmid optional/legacy and require at least one of {doi, pmid} ("not every citation has a PMID, but every citation has a stable id" — the reverse of today's rule). That is a requiredness change → major-only, parked as a 1.0-cleanup candidate, not an RMn. Until 1.0, DOI-only provenance is an explicitly-parked gap.


5. Module-level authorship & provenance (author-kind → scrutiny calibration)

5a. Structured per-version authorship — who created / edited / audited, and whether each is AI or a human expert

Verdict: SHIPPED in 0.4 (RM14) — an additive, digest-neutral authorship record; the old flat fields stay for compat. This is the module-level companion to the network-first validator (§4a): the validator — and a marketplace review queue, and a human auditor — needs to route its scrutiny by who authored the version, because AI and human error-spectra overlap but differ. An AI author fabricates plausible-but-wrong PMIDs / rsids / effect-sizes (exactly the checks RM11–RM13 automate); a human expert makes transcription / off-by-one / stale-reference slips. The format never performs the scrutiny (consumer-side, north star) — it must carry the author-kind so the consumer can select the right profile, the same "annotate so the consumer's X is safe and reproducible" contract as everywhere else.

What exists today, and why it does not cover it:

  • ModuleManifest.authors: list[str] — a flat list: no role (created/edited/audited), no kind (AI/human). The overloaded-axis anti-pattern (P5), at the list level.
  • curator / method — single free-form strings; Defaults.curator even defaults to "ai-module-creator", smuggling author-kind into a string a consumer cannot reliably facet on. This is precisely the axis-overload Principle 5 exists to unwind.
  • Provenance (generator/model/agent_version) + per-variant ProvenanceItem.human_reviewed — captures AI-generation and per-variant human review, but not module-level role attribution (who edited vs. audited this version), and it names only the AI side.

So the axes were half-present and tangled. The shipped shape is a structured, per-version authorship list (Contribution model) unbundling three orthogonal axes (P5): identity (who), role (created | edited | audited | reviewed, a closed vocab), and kind — a multi-valued, open tag set with a recommended seed: a human ladder of assurance human → human_expert → human_certified (medically / board-certified, e.g. a clinical geneticist), or ai plus a scale tag agent/team/swarm. There is deliberately no hybrid tag — it was rejected as non-explicit (hybrid what?); a joint contribution is two entries (a human and an ai), each with its own kind, so the mix is always spelled out. Each entry is optionally timestamped (at). "Per-version" falls out of immutability (P4): a version's manifest records its own authorship, and cross-version history is the union via aggregate_provenance.

Why it is cheap. artifact.digest is a Merkle root over the parquet files only — manifest metadata (logs, provenance, logo, and this) is deliberately out of it. So two versions with identical annotation content but different authorship keep the same content identity (correct: who authored ≠ what the annotation is). Adding it is additive/optional (P3/P8), touches no parquet column, and is digest-neutral even after 0.4 freezes. curator/authors/provenance stay working; folding the flat authors into the structured record is a 1.0-cleanup candidate. authoring_reference() picks up the new vocabularies automatically.

Charter check: data-agnostic ✓ (module metadata, not sample data); declarative ✓; P5 — this is the axis-unbundling; P6 — role/kind are frozenset vocabularies; the human-authorability gate is met by keeping the whole block optional and collapsing it to a single entry for the common one-AI-author case, so a module never reads like an enterprise audit ledger. Like panel, it is manifest metadata and is not reconstructed by the lossy parquet→spec reverse_module (which rebuilds a content skeleton) — the durable per-version record is the manifest itself, which is correct and no P7 issue (P7 governs artifact columns). → RM14 (shipped).


6. One variant, many effects — the variant-effect pair as identity

6a. Genotype-dependent poly-effect: sickle-cell rs334 (HBB Glu6Val)

Verdict: FIXED — was a silent round-trip GAP introduced by the variant_key column, closed by keying annotations.parquet on the variant-effect pair (variant_key, conclusion, negatives). No DSL change: the author still writes ordinary variants.csv rows. The fix is entirely in how the compiler dedups and rejoins annotation.

The scenario. rs334 (HBB, GAG→GTG, β-globin Glu6Val) is the textbook antagonistic-pleiotropy locus: the same variant produces categorically different phenotypes by genotype. The carrier is malaria-resistant; the homozygote has sickle-cell disease. Authored, that is two informative genotype rows at one locus:

rsid,genotype,state,conclusion,gene,phenotype,category
rs334,A/A,ref,No HbS allele — no sickle phenotype,HBB,Normal hemoglobin,hematologic
rs334,A/T,protective,Sickle-cell trait — resistance to severe P. falciparum malaria,HBB,Malaria resistance,infectious-disease
rs334,T/T,risk,Sickle-cell anemia (HbSS) — chronic hemolysis and vaso-occlusion,HBB,Sickle-cell disease,hematologic

The A/T and T/T rows share one variant_key (rs334) but carry different conclusion, phenotype, and category — infectious-disease (a protective trait) versus hematologic (a disease). The effects genuinely do not live in one category: category does not subsume them.

Why one-row-per-variant was wrong (the reasoning). weights.parquet is keyed on (variant_key, genotype), so each genotype row is faithfully distinct there. But annotations.parquet — which carries gene/phenotype/category and exists so a consumer can read a variant's annotation without scanning every genotype row — was deduplicated on variant_key alone. That silently asserts "a variant has one annotation." For a genuine poly-effect variant it is false: the second row (T/T) collapsed onto the first met (A/T), and on reverse_module every rs334 row was rewritten with the surviving row's phenotype/category. The homozygote's Sickle-cell disease / hematologic became Malaria resistance / infectious-disease — a confident, silent inversion of clinical meaning, and a Principle-7 (lossless round-trip) violation. This is not exotic: the same shape recurs wherever developmental / neural loci are pleiotropic and a single category tag cannot hold the effect. The bug was introduced with variant_key — before it, dedup keyed on rsid and had the same latent flaw, just less visible.

The honest identity of an annotation-bearing row is therefore the variant-effect pair, not the variant: variant + effect, where the effect is (conclusion, negatives). (It has to be conclusion, not genotype: annotations.parquet is per-variant-effect, and two genotypes that share an effect should still share one annotation row — dedup on the effect, not on the trigger.)

The mechanics (what actually changed).

  • Dedup key. _build_annotations now dedups on (variant_key, conclusion, negatives) — one row per genuine variant-effect pair (first occurrence wins). The A/T and T/T effects survive as two rows; a truly identical repeat still collapses.
  • Self-joinable table. annotations.parquet now carries conclusion and negatives (alongside variant_key), so the table can be rejoined to weights.parquet on the exact pair. weights already carries variant_key/conclusion/negatives per row, so no new weights column is needed.
  • Reverse probes the same key. reverse_module rebuilds each variant row's (variant_key, conclusion, negatives) triple from its weights row and looks up its own annotation — so T/T gets Sickle-cell disease/hematologic back, not A/T's. An older artifact whose annotations lacks a conclusion column falls back to the legacy variant_key-only probe (backward-compatible read).
  • Digest. artifact.digest moved once because annotations.parquet gained two columns — free at the time, since 0.4 was still unpublished (Principle 4); determinism + round-trip are the held invariants. That window is gone: 0.5.0 published on 2026-08-07, so a further column would be 1.0.

Charter check: data-agnostic ✓ (still pure annotation — no measurement; the sample's genotype is supplied by the consumer at query time); declarative ✓; P5 — this unbundles an overloaded identity (variant ≠ variant-effect); P7 — the whole point is restoring lossless round-trip, proven by test_poly_effect_annotation_survives_roundtrip (both effects survive and the digest is a fixed point). Human-authorability gate ✓: the author writes plain genotype→conclusion rows and never sees the key; the machinery is entirely compiler-side. See COMPILER.md §"Intentionally unimplemented" item 5 (the reverse_module boundary) and the SNV example in REFERENCE_EXAMPLES.md §1.


6. Population context — what the missing numbers actually block (0.5)

Four use cases that the SNP core, as of 0.4, could not serve at all. None is a gap in the annotation model: every one of them needs a reference number about a population, which a module had nowhere to put. They are the reason the frequency and gene-constraint sidecars exist.

6a. "Is this variant actually rare?" — offline carrier-frequency context. A carrier-screening module lists pathogenic HBB alleles. A consumer showing a positive call wants to say how common that allele is, and in which ancestry group — sickle-cell's HBB 11:5227002 T>A sits near 4.8% in African ancestry and near zero in Finnish. Before 0.5 the module could carry the annotation but not the frequency, so either the consumer fetched gnomAD at query time (network at read time, and a different number than the curator saw) or it said nothing. Closed additively: frequencies.csv → one row per (allele, ancestry group). Enabled.

6b. Reproducing an ACMG BA1/BS1 filter against the frequency the curator saw. BA1 ("allele frequency too high for a Mendelian disorder") and BS1 are frequency thresholds, and applying them needs the filtering allele frequency (faf95), not the point estimate. The blocker was never the threshold — a consumer can apply that — it was that a re-run months later hit a different gnomAD release and silently reclassified variants. Closed additively: faf95 is carried on its owning ancestry group's row, and dataset names the release, so the filter is reproducible against the numbers the curator used rather than against whatever the API serves today. This is why dataset is inside the fact set. Enabled.

6c. Out-of-ancestry caveats. A risk annotation derived from a European cohort applied to a South Asian sample may rest on an allele that is common in one group and absent in the other. The module cannot decide what to do about that — that is the consumer's disclosure policy, and the format never makes the call (the data-agnostic north star). What it can now do is carry the per-group numbers so the consumer has something to reason with. Enabled (format supplies the table; consumer supplies the policy).

6d. Gene-level triage on a cardio or cancer panel. A forty-variant panel spanning a dozen genes: which genes are haploinsufficient, and which tolerate loss of function? LOEUF separates them (MYH7 at 0.64 is constrained; a tolerant gene sits above 1), and pLI and missense-Z refine it. Repeating a gene-level fact on every variant row would be the wrong shape — same gene, forty copies, one axis smeared across another. Closed additively: gene_metrics.csv, one row per gene, a separate table (Principle 5). Enabled.

What stayed out. Sex-stratified counts (a second axis — folding nfe_XX into population would be the state-overloading mistake again), and an offline frequency snapshot (58 GB exomes / 742 GB genomes — not a thing that ships; parked in ROADMAP.md).


7. Regulatory effect — which gene a non-coding variant moves, and which way (0.7)

Two use cases, one shipped and one deliberately left open, and the pair is the point: they came from the same measurement round and only one of them has anyone asking for it.

7.1 "This promoter variant lowers TBX1 in most tissues." — ENABLED (RM194 + RM200).

A gene panel wants the variants that matter for its genes, and slicing an artifact by gene position answers a narrower question than it appears to: it catches coding and near-splice variants and silently drops the promoters, enhancers and chromatin-altering variants that act on a gene without sitting in it. AlphaGenome attributes a variant to genes across a 1 Mb window and names the gene itself, which is what @gene-map-is-another-sources-attribution requires — so the attribution is the source's rather than a span the caller drew.

Served by expression_effects.csv: one row per (variant, gene) with a signed magnitude, the majority direction across 371 tissue tracks, the count that agreed, and the distance to the gene. The distance is what makes it usable — distal scores run ~10× lower, so a magnitude threshold without it keeps the proximal rows while looking like it filtered on effect.

Non-commercial: the rows enter under declared_use=non_commercial through alphagenome_atlas, and a module carrying them is gated at compile like any other restricted source.

7.2 "This position sits in open chromatin in these cell types." — GAP, and nobody has asked.

Measurable and not adopted, which is a different verdict from blocked. RM200 measured the Atlas's *_ACTIVE scorers and found they describe the locus rather than the variant — raw assay units barely moved by the ALT, except where DNA geometry is at stake, where ATAC_ACTIVE moves ~16× control inside a Z-DNA former and ~15× in a G-quadruplex. So a locus-accessibility annotation is a coherent thing to record and genuinely new to this format.

What stops it is not the schema. Every number says what the scorers do; none says a consumer wants it, and the track vocabulary mixes cancer cell lines (EFO), anatomical structures (UBERON) and cell types (CL) under one ranking, with three ENCODE no term registered placeholders. Naming a cell type from that ranking would publish a sampling artefact as a mechanism — the same test that sank CHIP_TF and CAGE.

Reopen this with a consumer, never with an argument. The measurement is done and recorded in probes/ALPHAGENOME_ATLAS.md § 6.6; what is missing is somebody who needs the answer.

Roadmap items surfaced

The gaps above, consolidated. Format-side items migrate into ROADMAP.md; the consumer-side one is recorded so it is not mistaken for a format task.

# Item Kind Unblocks Priority
RM1 ✅ shipped — compiler materializes all 0.4 tables → parquet with lossless round-trip (generic _build_table/_write_table_csv over _TABLE_KINDS) format (compiler) 3a, 3c, harness on binned loci done
RM2 ✅ shipped — composed modules: variants.csv optional, a module carries only the kinds it uses (no empty variants.csv); studies.csv required iff variants present format (compiler) SNP+PRS, personal panels done
RM3 ✅ shipped in 0.4 sample — PharmVariantRow (pharm_variants.csv) + drug/response/evidence_level on DiplotypeRow format (schema) 2b done
RM20 ✅ shipped in 0.5 — PharmGKB annotations are per-genotype and per-category: genotype, phenotype_category (closed vocab) and annotation_id on PharmVariantRow; duplicate key (variant_key, drug, genotype, phenotype_category, annotation_id). Corrects RM3, which was validated against a sample rather than the corpus. format (schema + compiler) 2b, the real ClinPGx corpus done
RM21 ✅ shipped in 0.5 — Data-source licensing as data (sources.csv + manifest.sources): per (source, layer) licence, attribution, pinned license_sha256, tri-state share_alike/commercial_use, and the acquirer's declared_use. Compiler refuses annotation-layer content that forbids sale when no declaration is recorded; enricher refuses at acquisition. format (schema + compiler) + enricher 2c, marketplace redistribution done
RM22 ✅ shipped in 0.5 — PGx tables join resolution: enrich() reads pharm_variants.csv and haplotypes.csv, so a module with no variants.csv gets coordinates (it previously enriched to an empty resolution.csv). enricher 2c, 3c done
RM4 ✅ closed — off the compiler: just-dna-enricher draft-panel --source clinvar drafts the panel's rows and the author curates them; the panel: block is deprecated for removal at 1.0 enricher 2a —
RM5 ✅ shipped in 0.6 — symbolic/structural alleles: VCF's five closed first-level types with the length inside the token (<DEL:1500>, <CNV:TR:30>); a lengthless token is dropped, refused under strict. 5-HTTLPR, the motivating case, turned out to be a plain indel ref/alts state directly format (schema) 3b (SV), 1b (symbolic consume), 5-HTTLPR medium
RM6 Promote requires_callable to a typed boolean column; reserve/build callable_from (DP,GQ,FT three-state) format (schema) 1c callability low-medium
RM7 Evaluation-output / report-card schema for the verification harness consumer (just-dna-lite), NOT the format 1a — (not a format task)
RM8 ✅ shipped in 0.4 sample — reference.authoring_reference() + json_schemas(), generated from the live models format (schema) 1d drift done
RM9 ✅ shipped in 0.4 sample — manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS format (schema) 1d palette done
RM10 Optional declarative inheritance-expectation field (trio/de-novo assertion as data) format (schema) 1c trio low (only if needed)
RM11 ✅ shipped in 0.4 — doi provenance column on StudyRow (optional; validated against the DOI grammar, kept verbatim) format (schema) 4a done
RM12 ✅ shipped in 0.4 — Provenance locator: optional provenance_quote (keyword phrase) + provenance_regex (author-time-compiled, matched by a consumer-side linear-time engine — P1 pattern grammar) on StudyRow format (schema) 4a done
RM13 just-module-validator — network-first source-check/enrichment library consumer (new sibling), NOT the format 4a — (not a format task)
RM18 ✅ shipped in 0.5 — Population-frequency + gene-constraint sidecars (frequencies.csv, gene_metrics.csv), produced by the enricher's gnomAD v4.1 passes, compiled to their own optional parquets and fact-hashed. Retires the planned allele_frequency/af_population axes in favour of tables. format (schema + compiler) + enricher 6a–6d done
RM19 ✅ shipped in 0.5 — GA4GH VRS allele identity: stdlib derive_vrs_allele_id, vrs_id/caid cross-reference columns, and variant_key deriving from the VA for a resolved substitution. Satisfies RM15's build-naming condition (GRCh38-only now; multi-build minting remains RM15). format (schema + compiler) + enricher build-naming identity, cross-database joins done
RM14 ✅ shipped in 0.4 — Structured per-version authorship (authorship: [Contribution]): {who, role, kind, at}; role closed {created/edited/audited/reviewed}; kind open, seed = human ladder {human, human_expert, human_certified} / {ai}+scale {agent,team,swarm} (no hybrid — joint = two entries). Manifest metadata → digest-neutral. format (schema) 4a validator, marketplace review done

Takeaway. The two load-bearing items — RM1 + RM2 (compiler materialization + composed modules) — are now shipped: the frozen 0.4 shapes are runnable artifacts, and every composite/personal module compiles with lossless round-trip. What remains open is small and clearly scoped: RM3-adjacent extensions and the two provenance anchors RM11/RM12 (doi + fulltext locator) that let a network-first validator scrutinise a module without the format ever fetching — and those two shipped in 0.4, RM5 in 0.6, RM6 before 0.6, and RM10 folded into RM28, so this paragraph is a record of where the 0.4 round stood rather than a list of open work; the open list is ROADMAP.md. Notably, the format's purpose expansion (the verification harness) still needs no format change — it rides on the properties already frozen, now with the tables materialized under it.