A machine/LLM-facing authoring reference, generated from the live models (RM8).
Consumers (MCP servers, agents, docs) tend to hard-code a prose summary of the DSL — its columns,
vocabularies, genome build — which then drifts from the real schema (just-dna-agents' MCP
get_spec_format predated the 0.3 columns, for example). authoring_reference() derives that summary
by introspecting the Pydantic models and vocabularies, so it is a single source of truth that cannot
drift: every consumer renders the current field set. For a full JSON Schema, call
json_schemas() (Pydantic's model_json_schema() per model).
Dependency-light: this is a top-level aggregator over spec/binning/pgx/pgs/manifest/vocab;
nothing in the package imports it, so it introduces no cycle.
authoring_reference() -> dict[str, Any]
A drift-proof, JSON-serialisable description of the authored DSL — models (field lists),
vocabularies, reserved names, and the recommended display palette — generated from the live
schema. Consumers render this instead of a hand-maintained prose summary (RM8).
Source code in schema/src/just_dna_format/reference.py
| def authoring_reference() -> dict[str, Any]:
"""A drift-proof, JSON-serialisable description of the authored DSL — models (field lists),
vocabularies, reserved names, and the recommended display palette — generated from the live
schema. Consumers render this instead of a hand-maintained prose summary (RM8)."""
return {
"schema_version": SCHEMA_VERSION,
"genome_build_default": ModuleSpecConfig.model_fields["genome_build"].default,
"models": {name: _describe_model(model) for name, model in _ALL_MODELS.items()},
# Both blocks are generated from the fields' own `vocabulary` markers (`base.py`), not listed
# here. The list they replace drifted twice: it never learned about `recommendation_strength`
# or `phenotype_category` when 0.5 added them, and it filed `actionability` under
# `open_recommended` although `VariantRow` rejects a non-member — a drift in *closedness*,
# which is why the marker carries that flag and why `actionability` now appears below.
"vocabularies": _collect_vocabularies(closed=True),
"open_recommended": _collect_vocabularies(closed=False),
# Per-member prose, for the vocabularies where the member *name* cannot carry the whole rule.
# The element rules (RM54) are the case that forced it: on a `Number=R` VCF field the
# reference is element zero, so "the larger of the two" has two answers, and a vocabulary that
# is silent about which one it means repeats the defect it was added to fix one level down.
# Keyed by the same names as `vocabularies` above, so a consumer that found a member there
# can look up what it means without a second lookup table of its own.
#
# Derived from the fields' own markers, like both blocks above it. It was a comprehension over
# `VCF_POINTER_COMPANIONS` naming `ELEMENT_RULE_MEANINGS` in place — correct, and the only
# channel that composition could reach was this block, so the prose never got to the per-table
# `describe` an author filling one table runs (D1-4).
"vocabulary_notes": _collect_vocabulary_notes(),
# Alternative identity-column sets, any ONE of which satisfies a row's requirement. Field-level
# `required` cannot express "rsid OR chrom+start" — that rule is a model validator — so a tool
# listing required columns told an author a `variants.csv` row needed no identifier at all.
"required_any_of": {
name: [sorted(group) for group in model.REQUIRED_ANY_OF]
for name, model in _ALL_MODELS.items()
if getattr(model, "REQUIRED_ANY_OF", ())
},
"reserved_names": sorted(RESERVED_NAMES_0_4),
# Field-ownership boundary for the `module:` block (S2): keys the format knows about but that
# a *publishing registry* stamps, so an author must omit them. Distinct from `reserved_names`
# (future module columns) — these will never be authored. A consumer strips them before
# validation via `normalize.strip_authority_keys`; `module.version` is NOT here (it is a
# genuine advisory authored field).
"registry_stamped_keys": {
key: IDENTITY_AUTHORITY_REASONS[key] for key in sorted(IDENTITY_AUTHORITY_KEYS)
},
"recommended_palette": {"colors": RECOMMENDED_COLORS, "icons": RECOMMENDED_ICONS},
}
|
json_schemas() -> dict[str, Any]
Full JSON Schema per model (Pydantic model_json_schema()), for consumers that want the
machine-validatable form rather than the compact authoring_reference() summary.
Source code in schema/src/just_dna_format/reference.py
| def json_schemas() -> dict[str, Any]:
"""Full JSON Schema per model (Pydantic `model_json_schema()`), for consumers that want the
machine-validatable form rather than the compact `authoring_reference()` summary."""
return {name: model.model_json_schema() for name, model in _ALL_MODELS.items()}
|