Use cases & dogfooding — feasibility and blocker analysis¶
For each real or desired use case: is it enabled by the format as it stands, and if not, what is missing? (The verdicts were first written against 0.4 and are revised in place as items ship; the tree is at 0.7.) The point is to separate three things that get conflated —
- work the format already enables (author it today),
- work that is a consumer concern the format deliberately does not own (the data-agnostic north
star — see
CLAUDE.md), so there is nothing to add, and - genuine gaps that need an additive format change.
Every gap below is tagged → RMn and collected in Roadmap items surfaced at the end; those are the
items to migrate into ROADMAP.md.
The feedback → schema cycle (where this doc sits)¶
The design docs are stages of one loop; an idea moves left-to-right as it matures:
- Feedback — a consumer's field report. →
CONSUMER_SUGGESTIONS.md(open) /CONSUMER_SUGGESTIONS_HISTORY.md(answered), the runbook inCONSUMER_TRIAGE_LOOP.md. The pre-Snrounds 1 and 2 were their own threads and are both retired — their dispositions are recorded inhistory/CONSUMER_SUGGESTIONS_HISTORY_PRE_0_6.md - Usage → blockers → solvability — run each use case against the current bricks: is it enabled, consumer-side, or a gap; and is the gap closable additively? → this doc
- Means → draft schema → decision — for a gap worth closing, the proposed shape + a charter check
- the open questions to settle it. → the
PROPOSAL_*thread for the release underproposals/(eight so far, 0.4.1 through 0.7 PT3, every one now a record). (Shipped decisions →CHANGELOG.md.) - Conclusion — "how to do it now, with these bricks" — the distilled worked example once the shape
is settled. →
REFERENCE_EXAMPLES.md - Terminal, one of two:
- Fixed — schema + compiler shipped (the models;
COMPILER.mdmarks it validated/materialized), or - Deferred — a recognised gap parked as a roadmap item (
RMn→ROADMAP.md) when the means aren't worth building yet.
So this doc and REFERENCE_EXAMPLES.md are the same use cases at two points in the loop — here
they are questions (what blocks?), there they are answers (author it like this). An ENABLED /
schema-ready row here graduates to a REFERENCE_EXAMPLES entry; a GAP row graduates to the
release's proposal thread (if being closed now) or an RMn roadmap item (if deferred). The loop is why a
"blocker" is never a dead end: it is either dissolved (it was consumer-side all along), closed
additively, or explicitly parked.
Verdict legend - ENABLED — authorable/usable now, no format change. - CONSUMER-SIDE — a consumer (runner/app) feature; the format correctly owns nothing here (per the north star). No format change — often no blocker at all. - SCHEMA-READY / COMPILER-PENDING — the schema models express it; only the deferred compiler materialization (new parquet + round-trip) is missing. One known gap, not per-use-case. - GAP — needs an additive format addition (the interesting rows).
A recurring result: most "can the format do X?" questions dissolve into CONSUMER-SIDE, because the
module is declarative annotation and the doing is a consumer's. The format earns its keep by the
properties it froze (declarative-not-code, integrity-as-identity, the unresolved/callability
contract), which are what make a consumer's X safe and reproducible — not by hosting X.
1. The consumer's 0.5 suggestions, each run through the lens¶
1a. Verification harness — run a module against N VCFs, emit report-card diffs (round-2 §3b)¶
Verdict: CONSUMER-SIDE — no blocker, no format deliverable. This is the headline dogfooding use case and it is already enabled. Walk the requirements:
- A panel is a set of loci + expected genotypes/bins → already a module (a
variants.csvof genotype→conclusion rows, or a binning table of measure→phenotype). No new "panel type". - Deterministic extraction of the observed value from a VCF → the consumer's job, and
source_field(0.4) now names the exact VCFFORMAT/INFOfield to read, removing the last glue. - Trustworthy before/after (per-caller, ±liftover) diffs →
artifact.digestalready gives the module a content identity, so a report-card diff is a byte-level diff of two deterministic runs. - No-call ≠ mismatch → the mandatory
unresolvedoutcome + the callability contract already stop a "no-call under caller B" from masquerading as a "mismatch vs caller A".
The one thing that looked like a format artifact — a standardized evaluation-output / report-card
schema ({locus, observed, callability, bin_selected, verdict}) — is per-sample results, i.e. a
measurement, so the north star keeps it consumer-side. It belongs in just-dna-lite, not
just-dna-format. Nothing to schedule in the format. → RM7 records the consumer-side schema so
it is not mistaken for a format item.
1b. Augmented-VCF as the landing pad for cracked short-read loci (round-2 §3c)¶
Verdict: ENABLED at the format boundary; emission is CONSUMER-SIDE. A caller emits its niche
genotype (PER3 span, DAT1 motif-path, MAOA half-repeat) as a synthetic VCF record (<STR> with
INFO/RU, FORMAT/REPCN, custom evidence fields); a repeat_alleles.csv module consumes it via the
same source_field=REPCN path as ExpansionHunter. The format does not invent a representation — it
binds to the VCF one. Producing the augmented VCF is the caller's job (consumer-side). The only
format touch-point (source_field) shipped in 0.4. Consuming symbolic alleles (<STR n>,
<CNV>) at the count/dosage layer is enabled via the binning tables, and since 0.6 a symbolic
allele is expressible inside a VariantRow.genotype too — VCF's five types with the length in the
token, <CNV:TR:30> included (see §3b; RM5, shipped).
1c. Callability three-state, phasing-aware panels, trio/multi-sample (round-2 §3d)¶
- Callability three-state (covered-hom-ref vs no-call): CONSUMER-SIDE, derivable from VCF
DP/GQ/FT(or a gVCF ref-block). The format's part:requires_callable(reserved flag) marks rows where absence is informative; promoting it to a typed boolean column and reservingcallable_from(the DP,GQ,FT signal) are the format-side follow-ups→ RM6. Both shipped —requires_callableas a tri-stateVariantRowcolumn in 0.4,callable_fromin 0.5 — and 0.7 widened the first tohaplotypes.csvandpharm_variants.csv, the other two tables that name a locus, so a star-allele module can state the assumption its sources make in prose→ RM70. - Phasing-aware panels: ENABLED — the
phasedflag + the phased genotype formA|G(0.3 item 5b) already let a runner do cis/trans for compound-het and star-allele phasing (*2x2/*4vs*2/*4x2). No gap. - Trio / multi-sample (Mendelian / de-novo assertions): CONSUMER-SIDE — VCF is natively
multi-sample; the assertion runner is a consumer. An optional declarative inheritance-expectation
field on a panel row would let the module carry the assertion as data rather than consumer lore —
small, additive, optional
→ RM10(only if a real module needs it).
1d. Authoring-support suggestions (from the just-dna-agents integration, ROADMAP obs 2026-07-10)¶
- A canonical machine/LLM-facing authoring reference. Consumers (MCP servers, agents, docs)
hard-code prose summaries of the DSL that drift from the real schema. ADOPTED (RM8, shipped in
the 0.4 sample):
just_dna_format.reference.authoring_reference()returns a JSON-serialisable summary — every model's field list + all vocabularies + reserved names + the palette — generated from the live models, so it cannot drift;json_schemas()gives the full JSON Schema. A consumer'sget_spec_formatrenders this instead of a hand-maintained blob. - A recommended icon/color palette.
Displayvalidatesicon_set/colorbut shipped no recommended enumerated palette, so each authoring tool invented one. ADOPTED (RM9, shipped in the 0.4 sample):manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS(curatedsemantic-use → valuemaps, recommendation-only — not enforced), surfaced throughauthoring_reference().
2. Reference-data-backed module types¶
2a. ClinVar gene-panel (flag pathogenic variants in a gene set)¶
Verdict: ENABLED — and the materialization moved off the compiler (RM4, closed). The surface is
just-dna-enricher draft-panel --source clinvar --gene …: the enricher drafts the matching ClinVar
rows into variants.csv from a pinned snapshot, and the author's no-op over the drafted subset is the
authorial act, so the compiler never resolves a panel and never needs a reference mixin. Worked end to
end in reference_examples/hfe_hemochromatosis (REFERENCE_EXAMPLES §9b). The panel: block
GenePanelSpec declared is deprecated, removal at 1.0. The earlier verdict here — schema-ready, gap
on native compile-time materialization gated on a content-pinned ClinVar mixin — was the 0.4 reading;
what closed it was deciding that a drafted, curated row is the right subject, not a predicate the
compiler evaluates.
2b. PharmGKB drug-response annotation (item 9)¶
Verdict: ADOPTED (RM3, shipped in the 0.4 sample). A PharmGKB row maps a variant/diplotype → a
drug + a response/phenotype + a PharmGKB evidence level (1A…4, VALID_EVIDENCE_LEVELS)
— a different axis from a risk weight. Built as a dedicated PharmVariantRow (pharm_variants.csv)
for single-variant drug response (keeps the SNP core clean — one CSV = one concern), plus optional
drug/response/evidence_level columns on DiplotypeRow for the diplotype-keyed case. A PharmGKB
module has no empty variants.csv. evidence_level is a third significance-flavoured axis,
distinct from stat_significance/clin_sig (orthogonal-axes discipline, Principle 5). Materialization
deferred with the rest of 0.4.
One gap in it closed in 0.7 (RM132): the table could make a clinical claim per row and cite only per
variant. A ClinPGx-drafted module carried 1,482 drug-response rows with nowhere to ground any of
them, because studies.csv keys on (variant_key, pmid) while these rows key on drug, genotype and
category too — so one study row would ground every claim recorded for that variant. PharmVariantRow
now carries its own optional pmid under the rule RM47 settled for the bins: a row cites when its
claim is finer-grained than a study row's, and the citation table describes. evidence_level was never
the handle — it grades evidence rather than pointing at it — and provenance_quote deliberately did
not follow.
2d. GWAS effect sizes as grounding for a curated panel (0.6)¶
Verdict: ENABLED (RM90 shipped), and the interesting part is what it refused. A curator with a
trait panel wants the published effect sizes beside their own annotations — partly as evidence, partly
because a consumer reported that hand-set weight values "construct nonsense" across a corpus and that
GWAS effects are often better grounded (S36).
The enabled shape: just-dna-enricher gwas fills gwas_effects.csv, one row per published
association, and the module ships it beside variants.csv. A consumer joins on variant_key and reads
effect_size with effect_unit and effect_allele.
The blocker that was not dissolved, stated so nobody re-proposes it. Filling weight from those
effects is barred (MODULE_LIFECYCLE § Stage 3), and a per-row precedence rule was refused as putting two
methodologies in one summable column. What closed the gap additively was two tables and a declaration,
not an exception to the rule.
And a caution the real data supplies better than the design note could. rs1800562's 186
associations span 62 EFO traits in 12 distinct effect units — three of them spellings of one unit, two
more differing only in case — with 138 rows in the Catalog's uninformative unit and 42 of 195 naming
no effect allele at all. A "GWAS effects are better than curator weights" pipeline that pools those is
worse than the weights it replaces. Read them per trait, and read manifest.gwas_effects.units before
pooling anything.
2e. Republishing a licence-encumbered third-party corpus as modules (2026-09-13)¶
Verdict: ENABLED, and this is the case the licence machinery was built for — with one condition that lands on the marketplace rather than on the format, and one that shapes how the modules are cut.
The case, named by the maintainer. Take a large existing annotation corpus somebody else curated,
rework it, and publish it as modules. The concrete instance is SNPedia: 106,603 SNP entries plus
its genosets, reachable as a MediaWiki dump (the extraction is already done in
zhaofengli/snappy — see
ROADMAP_0_8 § RM188).
This is a different shape from 2a and 2b: those draft rows from a source into an author's own
module, where this republishes somebody's whole corpus under terms that travel with it.
The terms are established, and that is the first thing that had to be true. Read 2026-09-13 off three primary pages, all naming the same licence in the same words:
| Page | Says |
|---|---|
SNPedia:Copyrights |
"The content in SNPedia is licensed under a Creative Commons Attribution-Noncommercial-Share Alike 3.0 United States License." |
SNPedia:General_disclaimer |
the same sentence, plus "For more details, see our Terms of Use." |
Bulk |
the same sentence, and it invites bulk access — a bots. subdomain, code samples in four languages, a published GFF |
Contrast this with the HPO refusal (ENRICHER.md), because the two look alike and
are not. HPO's licence URL 404s and OBO records a bare label with no SPDX id, so nothing
machine-readable establishes the terms and an unestablished permission is not a permission. SNPedia
names a versioned, SPDX-identifiable licence (CC-BY-NC-SA-3.0-US) on three pages. Its
SNPedia:Terms_of_Use page is a 404 — the elaboration is missing — but the grant is not the
elaboration, and @no-named-licence is about a source that names none. This one names one.
Two acquisition notes. www.snpedia.com is behind Incapsula and answers a 212-byte challenge stub to
curl; bots.snpedia.com serves the same pages plainly, which is the sanctioned route and is what the
Bulk page tells you to use. Its Forbidden section bans two scraping patterns — every version of
every page, and every possible rs# — and says nothing about reuse; the dump route avoids both.
What the three letters cost, in this format's own fields.
| Term | The cell | The consequence |
|---|---|---|
| NC | commercial_use=False, declared_use=non_commercial |
The compile gate refuses a module declaring commercial use. This is @gate-is-data-driven doing its job, not an obstacle. |
| SA | share_alike=True |
The module itself goes out under CC BY-NC-SA 3.0 US, and share-alike is viral over whatever it is mixed with. So a SNPedia corpus is its own module, never rows blended into a general one — which the one CSV = one concern rule already wanted. |
| BY | attribution, license_url |
Credit as designated, the URI, and an indication that it was modified — "after rework" is a modification and the licence requires saying so. |
redistribution=True: CC BY-NC-SA permits it. Recorded, not gated — @redistribution-ungated, RM27.
The one condition that is not the format's to answer. NC in CC 3.0 bars use "primarily intended
for or directed toward commercial advantage or private monetary compensation". Whether
just-dna-registry distributing an NC module is itself commercial use is a real question, and it
decides whether this use case works end to end — a free catalog is a different answer from a paid one,
and the answer may differ again for a commercially-operated free one. That belongs to the
marketplace, and it should be settled there before the first NC corpus is published, not after.
Two things to check before building, neither of which blocks the design. SNPedia has been owned by MyHeritage since 2019, and their announcement said it would remain free "for academic and non-profit use" — phrasing narrower than the CC grant, which permits any non-commercial use. A CC licence is irrevocable for the versions published under it, so the practical answer is to pin the dump revision you took, exactly as every other source here is pinned. Separately: a genoset's boolean structure is arguably a fact and its summary prose is plainly expression, so "rework" launders neither — and stripping the prose to dodge share-alike would throw away most of what makes the corpus worth having. Neither point is a legal opinion, and neither should be treated as one.
What this case actually proves. Every earlier entry in this section takes a permissive source
and produces a sellable module. This is the first that takes an encumbered one and produces an
honestly-labelled unsellable module — and the format holds it without a new column, a new flag, or a
mode. That is the whole argument for licensing-as-data over a --non-commercial switch
(@licensing-as-data, @gate-is-data-driven), with a real corpus standing on it instead of a
hypothetical.
3. Composite modules (the real pipeline shapes)¶
3a. SNP + PRS in one module¶
Verdict: ENABLED (RM1 + RM2 shipped). A module is a directory of CSVs carrying both
variants.csv (VariantRow) and pgs.csv (PgsRow), joined on the shared trait_efo_id (item 5) so
a variant panel and its PRS companion sit in one content-addressed unit. The compiler now materializes
every present table kind to parquet (round-trip lossless) and treats variants.csv as optional, so
composed and single-domain modules both compile. No blocker.
0.5 revisit: the 0.4 shape was modelled at the wrong grain, and real data proved it. RM3 was declared shipped against a hand-authored sample. Run against the actual ClinPGx corpus it does not hold, in two stages:
- A clinical annotation is published per genotype — the summary table names the variant and drug,
a child table gives one row per call, and the large majority carry exactly three.
PharmVariantRowhad nogenotype, and the compiler deduped on(variant_key, drug), so authoring the real SLCO1B1/simvastatin annotation producedduplicate row for key ('rs4149056', 'simvastatin'). ~97% of the corpus was unauthorable. - Adding
genotypewas not enough. One variant and one drug carry several distinct annotations — rs4149056 + simvastatin is Metabolism/PK at 1A, Efficacy at 3 and Toxicity at 1A. 1,199 of 17,380 triples collide; 839 separate by phenotype category, 283 by neither category nor level.
Closed additively with genotype, phenotype_category (closed vocabulary) and annotation_id (a
source accession as identity, like PgsRow.pgs_id) → RM20. The lesson is the dogfood rule in
CLAUDE.md read from the other side: a shape validated against a sample rather than a corpus is not
validated. See reference_examples/pgx_slco1b1_simvastatin/.
2c. Star alleles and drug response from the live authorities (0.5)¶
Verdict: ENABLED, with the licensing made legible (RM21). The enricher can now cross-check a
module's allele_function.csv against PharmVar and CPIC and its pharm_variants.csv against a
ClinPGx snapshot, and resolution reaches pharm_variants.csv/haplotypes.csv so a PGx module gets
coordinates without carrying a variants.csv.
The blocker that turned up was not technical. Every pharmacogenomics upstream is copyleft and none
is sellable: ClinPGx, CPIC and PharmVar are each CC BY-SA 4.0 plus a separate contractual bar on
sale. api.pharmgkb.org was retired 2026-07-20, and CPIC sits inside the ClinPGx merger with its
licence page redirecting to the ClinPGx policy — so swapping sources does not escape the terms. That
is consumer-relevant rather than a format gap, but it is unrepresentable in a 0.4 module, so it was
closed additively as sources.csv (RM21) rather than left to a README nobody can query.
Composition principle (settled during the PharmGKB decision, now in CLAUDE.md): a module composes
from optional table kinds — one CSV = one concern — so the SNP core (variants.csv+studies.csv)
stays minimal and no module ever carries an empty variants.csv or a foreign domain's columns just to
host one table. This is the human-authorable half of the RM2 work.
3b. SNP + indels¶
Verdict: ENABLED for small ACGT indels; GAP for structural/symbolic. VariantRow alleles are
^[ACGT]+$ multi-base, so a small insertion/deletion is expressible today (ref=A, alts=AT,
genotype A/AT) on the same variants.csv as SNPs — a mixed SNP+indel panel is authorable now.
Symbolic/large structural alleles shipped in 0.6 (RM5): VCF's five closed first-level types with
the length inside the token (<DEL:4977>, <CNV:TR:30>), worked on a real module in
reference_examples/mt_common_deletion. A lengthless token is dropped with a warning and refused under
strict. Dosage and count still route through the copy-number / repeat binning tables; the earlier
"recognised gap" here is closed, and what stays unexpressible is deliberate (CPIC's IUPAC codes).
3c. SNP + PRS + PGx + CNV in one "personal panel"¶
Verdict: ENABLED (RM1 + RM2 shipped). The generalization of 3a: a personal/curated module mixing
variants.csv, pgs.csv, activity_phenotype.csv/diplotypes.csv, and copynumbers.csv, all joined
on trait_efo_id, compiles today — each present kind materializes to parquet with round-trip. This is
exactly the "personal module re-checked deterministically on every pipeline change" the verification
harness (§1a) wants — and RM1/RM2 are what unlocked it.
4. Network-first validation & enrichment (external-source scrutiny)¶
4a. just-module-validator — deterministic source-checks + provenance enrichment against public sources¶
Verdict: CONSUMER-SIDE of the format and compiler (Principle 2 keeps it out of those two tiers) —
and the sibling it proposed exists: it is just-dna-enricher (RM13, realized). The two additive
format anchors it needed shipped in 0.4 (RM11/RM12), and one requiredness fix waits for 1.0.
The sibling library as proposed, network-first: given a module it checks the authored claims
against public sources and enriches them — every one of these is now an enricher check with a
VALID_VERIFICATION_CHECKS member (citation_existence, rsid_currency and
rsid_coordinate_agreement, citation_identifier, provenance_quote) —
- validate every
pmidresolves in PubMed, and everyrsidresolves in dbSNP at the authoredchrom:start(flag coord/liftover drift); - cross-fill provenance ids — derive a
doifrom a PMID and vice-versa; - confirm a study's claim actually appears in the cited article's fulltext (imagine further source-checks in the same spirit).
By the data-agnostic north star and Principle 2 (no network; inject-only), the doing — every
fetch and lookup — is a consumer's, and can never live in just-dna-format/just-dna-compiler
(that would pull the network dependency the tiers forbid, Goal 2). So the validator is a new
consumer/enricher sibling to just-dna-lite, recorded as → RM13 so it is not mistaken for format
scope — exactly as the report-card harness is (§1a / RM7). Crucially, most of what it checks needs
nothing from the format: rsid, chrom, start already exist, so validating them against dbSNP
is pure consumer work — enabled today. Two things it wants to anchor are genuine additive format
gaps, and one is a 1.0 requiredness fix:
-
doias a provenance id — additive, shipped in 0.4 (RM11).StudyRowpreviously carried onlypmid(required, and it must contain ≥1 real PubMed id). DOI is wider: it covers preprints (bioRxiv/medRxiv), books, theses, and datasets that have no PMID. The optionaldoicolumn now lets the validator record and cross-fill it, and lets a module cite a DOI-bearing source. Purely additive → P3/P8 clean (new optional field; existing data still validates); validated against the DOI grammar and kept verbatim.→ RM11. -
A provenance locator — search-phrase/regex pointing at the passage in fulltext — additive, shipped in 0.4 (RM12). So the validator can answer "does the cited article's fulltext actually contain this claim?" in a yes/no manner, a study row now carries optional
provenance_quote(keyword phrase) andprovenance_regex. The regex sits squarely inside Principle 1's sanctioned escape hatch: a declarative pattern grammar is data, not code — the module ships the pattern, the consumer supplies the fulltext and runs the match, evaluated by a linear-time / ReDoS-safe engine (P1's explicit requirement; the compiler onlyre.compile-checks it at author time). It is the provenance analogue ofsource_field(0.4):source_fieldis a declarative pointer to where the measurement lives in a VCF; the locator is a declarative pointer to where the claim lives in the article. Neither holds the data it points at (north star ✓). Primarily an aid for LLM-authors (which can emit a precise pattern), yet a plain keyword phrase is legible enough to clear the human-authorability gate for a human author too.→ RM12. -
pmidis mandatory today — the DOI-only case cannot be closed additively (a 1.0 fix).pmid: stris required and must parse to a real PubMed id (extract_pmids), so a preprint/book/thesis with only a DOI is unauthorable right now — and demoting a required field to optional is precisely the move Principle 8 forbids within a major. Addingdoi(RM11) is necessary but not sufficient: whilepmidstays required-and-PMID-shaped, DOI-only provenance is still rejected. The full fix is doi-first at 1.0 — makepmidoptional/legacy and require at least one of{doi, pmid}("not every citation has a PMID, but every citation has a stable id" — the reverse of today's rule). That is a requiredness change → major-only, parked as a 1.0-cleanup candidate, not anRMn. Until 1.0, DOI-only provenance is an explicitly-parked gap.
5. Module-level authorship & provenance (author-kind → scrutiny calibration)¶
5a. Structured per-version authorship — who created / edited / audited, and whether each is AI or a human expert¶
Verdict: SHIPPED in 0.4 (RM14) — an additive, digest-neutral authorship record; the old flat
fields stay for compat. This is the module-level companion to the
network-first validator (§4a): the validator — and a marketplace review queue, and a human auditor —
needs to route its scrutiny by who authored the version, because AI and human error-spectra
overlap but differ. An AI author fabricates plausible-but-wrong PMIDs / rsids / effect-sizes (exactly
the checks RM11–RM13 automate); a human expert makes transcription / off-by-one / stale-reference
slips. The format never performs the scrutiny (consumer-side, north star) — it must carry the
author-kind so the consumer can select the right profile, the same "annotate so the consumer's X is
safe and reproducible" contract as everywhere else.
What exists today, and why it does not cover it:
ModuleManifest.authors: list[str]— a flat list: no role (created/edited/audited), no kind (AI/human). The overloaded-axis anti-pattern (P5), at the list level.curator/method— single free-form strings;Defaults.curatoreven defaults to"ai-module-creator", smuggling author-kind into a string a consumer cannot reliably facet on. This is precisely the axis-overload Principle 5 exists to unwind.Provenance(generator/model/agent_version) + per-variantProvenanceItem.human_reviewed— captures AI-generation and per-variant human review, but not module-level role attribution (who edited vs. audited this version), and it names only the AI side.
So the axes were half-present and tangled. The shipped shape is a structured, per-version
authorship list (Contribution model) unbundling three orthogonal axes (P5): identity (who),
role (created | edited | audited | reviewed, a closed vocab), and kind — a multi-valued,
open tag set with a recommended seed: a human ladder of assurance human → human_expert →
human_certified (medically / board-certified, e.g. a clinical geneticist), or ai plus a scale tag
agent/team/swarm. There is deliberately no hybrid tag — it was rejected as non-explicit
(hybrid what?); a joint contribution is two entries (a human and an ai), each with its own
kind, so the mix is always spelled out. Each entry is optionally timestamped (at). "Per-version"
falls out of immutability (P4): a version's manifest records its own authorship, and cross-version
history is the union via aggregate_provenance.
Why it is cheap. artifact.digest is a Merkle root over the parquet files only — manifest
metadata (logs, provenance, logo, and this) is deliberately out of it. So two versions with
identical annotation content but different authorship keep the same content identity (correct: who
authored ≠ what the annotation is). Adding it is additive/optional (P3/P8), touches no parquet column,
and is digest-neutral even after 0.4 freezes. curator/authors/provenance stay working;
folding the flat authors into the structured record is a 1.0-cleanup candidate. authoring_reference()
picks up the new vocabularies automatically.
Charter check: data-agnostic ✓ (module metadata, not sample data); declarative ✓; P5 — this is
the axis-unbundling; P6 — role/kind are frozenset vocabularies; the human-authorability gate is
met by keeping the whole block optional and collapsing it to a single entry for the common
one-AI-author case, so a module never reads like an enterprise audit ledger. Like panel, it is
manifest metadata and is not reconstructed by the lossy parquet→spec reverse_module (which rebuilds
a content skeleton) — the durable per-version record is the manifest itself, which is correct and no P7
issue (P7 governs artifact columns). → RM14 (shipped).
6. One variant, many effects — the variant-effect pair as identity¶
6a. Genotype-dependent poly-effect: sickle-cell rs334 (HBB Glu6Val)¶
Verdict: FIXED — was a silent round-trip GAP introduced by the variant_key column, closed by
keying annotations.parquet on the variant-effect pair (variant_key, conclusion, negatives). No
DSL change: the author still writes ordinary variants.csv rows. The fix is entirely in how the
compiler dedups and rejoins annotation.
The scenario. rs334 (HBB, GAG→GTG, β-globin Glu6Val) is the textbook antagonistic-pleiotropy
locus: the same variant produces categorically different phenotypes by genotype. The carrier is
malaria-resistant; the homozygote has sickle-cell disease. Authored, that is two informative genotype
rows at one locus:
rsid,genotype,state,conclusion,gene,phenotype,category
rs334,A/A,ref,No HbS allele — no sickle phenotype,HBB,Normal hemoglobin,hematologic
rs334,A/T,protective,Sickle-cell trait — resistance to severe P. falciparum malaria,HBB,Malaria resistance,infectious-disease
rs334,T/T,risk,Sickle-cell anemia (HbSS) — chronic hemolysis and vaso-occlusion,HBB,Sickle-cell disease,hematologic
The A/T and T/T rows share one variant_key (rs334) but carry different conclusion,
phenotype, and category — infectious-disease (a protective trait) versus hematologic (a
disease). The effects genuinely do not live in one category: category does not subsume them.
Why one-row-per-variant was wrong (the reasoning). weights.parquet is keyed on
(variant_key, genotype), so each genotype row is faithfully distinct there. But
annotations.parquet — which carries gene/phenotype/category and exists so a consumer can read a
variant's annotation without scanning every genotype row — was deduplicated on variant_key alone.
That silently asserts "a variant has one annotation." For a genuine poly-effect variant it is false:
the second row (T/T) collapsed onto the first met (A/T), and on reverse_module every rs334
row was rewritten with the surviving row's phenotype/category. The homozygote's Sickle-cell
disease / hematologic became Malaria resistance / infectious-disease — a confident, silent
inversion of clinical meaning, and a Principle-7 (lossless round-trip) violation. This is not exotic:
the same shape recurs wherever developmental / neural loci are pleiotropic and a single category tag
cannot hold the effect. The bug was introduced with variant_key — before it, dedup keyed on rsid
and had the same latent flaw, just less visible.
The honest identity of an annotation-bearing row is therefore the variant-effect pair, not the
variant: variant + effect, where the effect is (conclusion, negatives). (It has to be conclusion,
not genotype: annotations.parquet is per-variant-effect, and two genotypes that share an effect
should still share one annotation row — dedup on the effect, not on the trigger.)
The mechanics (what actually changed).
- Dedup key.
_build_annotationsnow dedups on(variant_key, conclusion, negatives)— one row per genuine variant-effect pair (first occurrence wins). TheA/TandT/Teffects survive as two rows; a truly identical repeat still collapses. - Self-joinable table.
annotations.parquetnow carriesconclusionandnegatives(alongsidevariant_key), so the table can be rejoined toweights.parqueton the exact pair.weightsalready carriesvariant_key/conclusion/negativesper row, so no newweightscolumn is needed. - Reverse probes the same key.
reverse_modulerebuilds each variant row's(variant_key, conclusion, negatives)triple from itsweightsrow and looks up its own annotation — soT/TgetsSickle-cell disease/hematologicback, notA/T's. An older artifact whoseannotationslacks aconclusioncolumn falls back to the legacyvariant_key-only probe (backward-compatible read). - Digest.
artifact.digestmoved once becauseannotations.parquetgained two columns — free at the time, since 0.4 was still unpublished (Principle 4); determinism + round-trip are the held invariants. That window is gone: 0.5.0 published on 2026-08-07, so a further column would be 1.0.
Charter check: data-agnostic ✓ (still pure annotation — no measurement; the sample's genotype is
supplied by the consumer at query time); declarative ✓; P5 — this unbundles an overloaded identity
(variant ≠ variant-effect); P7 — the whole point is restoring lossless round-trip, proven by
test_poly_effect_annotation_survives_roundtrip (both effects survive and the digest is a fixed
point). Human-authorability gate ✓: the author writes plain genotype→conclusion rows and never sees the
key; the machinery is entirely compiler-side. See COMPILER.md §"Intentionally
unimplemented" item 5 (the reverse_module boundary) and the SNV example in
REFERENCE_EXAMPLES.md §1.
6. Population context — what the missing numbers actually block (0.5)¶
Four use cases that the SNP core, as of 0.4, could not serve at all. None is a gap in the annotation model: every one of them needs a reference number about a population, which a module had nowhere to put. They are the reason the frequency and gene-constraint sidecars exist.
6a. "Is this variant actually rare?" — offline carrier-frequency context. A carrier-screening module
lists pathogenic HBB alleles. A consumer showing a positive call wants to say how common that allele
is, and in which ancestry group — sickle-cell's HBB 11:5227002 T>A sits near 4.8% in African
ancestry and near zero in Finnish. Before 0.5 the module could carry the annotation but not the
frequency, so either the consumer fetched gnomAD at query time (network at read time, and a different
number than the curator saw) or it said nothing. Closed additively: frequencies.csv → one row per
(allele, ancestry group). Enabled.
6b. Reproducing an ACMG BA1/BS1 filter against the frequency the curator saw. BA1 ("allele frequency
too high for a Mendelian disorder") and BS1 are frequency thresholds, and applying them needs the
filtering allele frequency (faf95), not the point estimate. The blocker was never the threshold — a
consumer can apply that — it was that a re-run months later hit a different gnomAD release and silently
reclassified variants. Closed additively: faf95 is carried on its owning ancestry group's row, and
dataset names the release, so the filter is reproducible against the numbers the curator used rather
than against whatever the API serves today. This is why dataset is inside the fact set. Enabled.
6c. Out-of-ancestry caveats. A risk annotation derived from a European cohort applied to a South Asian sample may rest on an allele that is common in one group and absent in the other. The module cannot decide what to do about that — that is the consumer's disclosure policy, and the format never makes the call (the data-agnostic north star). What it can now do is carry the per-group numbers so the consumer has something to reason with. Enabled (format supplies the table; consumer supplies the policy).
6d. Gene-level triage on a cardio or cancer panel. A forty-variant panel spanning a dozen genes:
which genes are haploinsufficient, and which tolerate loss of function? LOEUF separates them (MYH7 at
0.64 is constrained; a tolerant gene sits above 1), and pLI and missense-Z refine it. Repeating a
gene-level fact on every variant row would be the wrong shape — same gene, forty copies, one axis
smeared across another. Closed additively: gene_metrics.csv, one row per gene, a separate table
(Principle 5). Enabled.
What stayed out. Sex-stratified counts (a second axis — folding nfe_XX into population would be
the state-overloading mistake again), and an offline frequency snapshot (58 GB exomes / 742 GB
genomes — not a thing that ships; parked in ROADMAP.md).
7. Regulatory effect — which gene a non-coding variant moves, and which way (0.7)¶
Two use cases, one shipped and one deliberately left open, and the pair is the point: they came from the same measurement round and only one of them has anyone asking for it.
7.1 "This promoter variant lowers TBX1 in most tissues." — ENABLED (RM194 + RM200).
A gene panel wants the variants that matter for its genes, and slicing an artifact by gene position
answers a narrower question than it appears to: it catches coding and near-splice variants and
silently drops the promoters, enhancers and chromatin-altering variants that act on a gene without
sitting in it. AlphaGenome attributes a variant to genes across a 1 Mb window and names the gene
itself, which is what @gene-map-is-another-sources-attribution requires — so the attribution is the
source's rather than a span the caller drew.
Served by expression_effects.csv: one row per (variant, gene) with a signed magnitude, the
majority direction across 371 tissue tracks, the count that agreed, and the distance to the gene.
The distance is what makes it usable — distal scores run ~10× lower, so a magnitude threshold
without it keeps the proximal rows while looking like it filtered on effect.
Non-commercial: the rows enter under declared_use=non_commercial through alphagenome_atlas, and a
module carrying them is gated at compile like any other restricted source.
7.2 "This position sits in open chromatin in these cell types." — GAP, and nobody has asked.
Measurable and not adopted, which is a different verdict from blocked. RM200 measured the Atlas's
*_ACTIVE scorers and found they describe the locus rather than the variant — raw assay units
barely moved by the ALT, except where DNA geometry is at stake, where ATAC_ACTIVE moves ~16× control
inside a Z-DNA former and ~15× in a G-quadruplex. So a locus-accessibility annotation is a coherent
thing to record and genuinely new to this format.
What stops it is not the schema. Every number says what the scorers do; none says a consumer wants
it, and the track vocabulary mixes cancer cell lines (EFO), anatomical structures (UBERON) and
cell types (CL) under one ranking, with three ENCODE no term registered placeholders. Naming a
cell type from that ranking would publish a sampling artefact as a mechanism — the same test that
sank CHIP_TF and CAGE.
Reopen this with a consumer, never with an argument. The measurement is done and recorded in
probes/ALPHAGENOME_ATLAS.md § 6.6; what is missing is somebody who needs the answer.
Roadmap items surfaced¶
The gaps above, consolidated. Format-side items migrate into ROADMAP.md; the
consumer-side one is recorded so it is not mistaken for a format task.
| # | Item | Kind | Unblocks | Priority |
|---|---|---|---|---|
| RM1 | ✅ shipped — compiler materializes all 0.4 tables → parquet with lossless round-trip (generic _build_table/_write_table_csv over _TABLE_KINDS) |
format (compiler) | 3a, 3c, harness on binned loci | done |
| RM2 | ✅ shipped — composed modules: variants.csv optional, a module carries only the kinds it uses (no empty variants.csv); studies.csv required iff variants present |
format (compiler) | SNP+PRS, personal panels | done |
| RM3 | ✅ shipped in 0.4 sample — PharmVariantRow (pharm_variants.csv) + drug/response/evidence_level on DiplotypeRow |
format (schema) | 2b | done |
| RM20 | ✅ shipped in 0.5 — PharmGKB annotations are per-genotype and per-category: genotype, phenotype_category (closed vocab) and annotation_id on PharmVariantRow; duplicate key (variant_key, drug, genotype, phenotype_category, annotation_id). Corrects RM3, which was validated against a sample rather than the corpus. |
format (schema + compiler) | 2b, the real ClinPGx corpus | done |
| RM21 | ✅ shipped in 0.5 — Data-source licensing as data (sources.csv + manifest.sources): per (source, layer) licence, attribution, pinned license_sha256, tri-state share_alike/commercial_use, and the acquirer's declared_use. Compiler refuses annotation-layer content that forbids sale when no declaration is recorded; enricher refuses at acquisition. |
format (schema + compiler) + enricher | 2c, marketplace redistribution | done |
| RM22 | ✅ shipped in 0.5 — PGx tables join resolution: enrich() reads pharm_variants.csv and haplotypes.csv, so a module with no variants.csv gets coordinates (it previously enriched to an empty resolution.csv). |
enricher | 2c, 3c | done |
| RM4 | ✅ closed — off the compiler: just-dna-enricher draft-panel --source clinvar drafts the panel's rows and the author curates them; the panel: block is deprecated for removal at 1.0 |
enricher | 2a | — |
| RM5 | ✅ shipped in 0.6 — symbolic/structural alleles: VCF's five closed first-level types with the length inside the token (<DEL:1500>, <CNV:TR:30>); a lengthless token is dropped, refused under strict. 5-HTTLPR, the motivating case, turned out to be a plain indel ref/alts state directly |
format (schema) | 3b (SV), 1b (symbolic consume), 5-HTTLPR | medium |
| RM6 | Promote requires_callable to a typed boolean column; reserve/build callable_from (DP,GQ,FT three-state) |
format (schema) | 1c callability | low-medium |
| RM7 | Evaluation-output / report-card schema for the verification harness | consumer (just-dna-lite), NOT the format |
1a | — (not a format task) |
| RM8 | ✅ shipped in 0.4 sample — reference.authoring_reference() + json_schemas(), generated from the live models |
format (schema) | 1d drift | done |
| RM9 | ✅ shipped in 0.4 sample — manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS |
format (schema) | 1d palette | done |
| RM10 | Optional declarative inheritance-expectation field (trio/de-novo assertion as data) | format (schema) | 1c trio | low (only if needed) |
| RM11 | ✅ shipped in 0.4 — doi provenance column on StudyRow (optional; validated against the DOI grammar, kept verbatim) |
format (schema) | 4a | done |
| RM12 | ✅ shipped in 0.4 — Provenance locator: optional provenance_quote (keyword phrase) + provenance_regex (author-time-compiled, matched by a consumer-side linear-time engine — P1 pattern grammar) on StudyRow |
format (schema) | 4a | done |
| RM13 | just-module-validator — network-first source-check/enrichment library |
consumer (new sibling), NOT the format | 4a | — (not a format task) |
| RM18 | ✅ shipped in 0.5 — Population-frequency + gene-constraint sidecars (frequencies.csv, gene_metrics.csv), produced by the enricher's gnomAD v4.1 passes, compiled to their own optional parquets and fact-hashed. Retires the planned allele_frequency/af_population axes in favour of tables. |
format (schema + compiler) + enricher | 6a–6d | done |
| RM19 | ✅ shipped in 0.5 — GA4GH VRS allele identity: stdlib derive_vrs_allele_id, vrs_id/caid cross-reference columns, and variant_key deriving from the VA for a resolved substitution. Satisfies RM15's build-naming condition (GRCh38-only now; multi-build minting remains RM15). |
format (schema + compiler) + enricher | build-naming identity, cross-database joins | done |
| RM14 | ✅ shipped in 0.4 — Structured per-version authorship (authorship: [Contribution]): {who, role, kind, at}; role closed {created/edited/audited/reviewed}; kind open, seed = human ladder {human, human_expert, human_certified} / {ai}+scale {agent,team,swarm} (no hybrid — joint = two entries). Manifest metadata → digest-neutral. |
format (schema) | 4a validator, marketplace review | done |
Takeaway. The two load-bearing items — RM1 + RM2 (compiler materialization + composed
modules) — are now shipped: the frozen 0.4 shapes are runnable artifacts, and every
composite/personal module compiles with lossless round-trip. What remains open is small and clearly
scoped: RM3-adjacent extensions and the two provenance anchors RM11/RM12 (doi + fulltext locator)
that let a network-first validator scrutinise a module without the format ever fetching — and those
two shipped in 0.4, RM5 in 0.6, RM6 before 0.6, and RM10 folded into RM28, so this paragraph is a
record of where the 0.4 round stood rather than a list of open work; the open list is ROADMAP.md. Notably,
the format's purpose expansion (the verification harness) still needs no format change — it rides
on the properties already frozen, now with the tables materialized under it.