Skip to content

just_dna_format.reference

just_dna_format.reference

A machine/LLM-facing authoring reference, generated from the live models (RM8).

Consumers (MCP servers, agents, docs) tend to hard-code a prose summary of the DSL — its columns, vocabularies, genome build — which then drifts from the real schema (just-dna-agents' MCP get_spec_format predated the 0.3 columns, for example). authoring_reference() derives that summary by introspecting the Pydantic models and vocabularies, so it is a single source of truth that cannot drift: every consumer renders the current field set. For a full JSON Schema, call json_schemas() (Pydantic's model_json_schema() per model).

Dependency-light: this is a top-level aggregator over spec/binning/pgx/pgs/manifest/vocab; nothing in the package imports it, so it introduces no cycle.

authoring_reference

authoring_reference() -> dict[str, Any]

A drift-proof, JSON-serialisable description of the authored DSL — models (field lists), vocabularies, reserved names, and the recommended display palette — generated from the live schema. Consumers render this instead of a hand-maintained prose summary (RM8).

Source code in schema/src/just_dna_format/reference.py
def authoring_reference() -> dict[str, Any]:
    """A drift-proof, JSON-serialisable description of the authored DSL — models (field lists),
    vocabularies, reserved names, and the recommended display palette — generated from the live
    schema. Consumers render this instead of a hand-maintained prose summary (RM8)."""
    return {
        "schema_version": SCHEMA_VERSION,
        "genome_build_default": ModuleSpecConfig.model_fields["genome_build"].default,
        "models": {name: _describe_model(model) for name, model in _ALL_MODELS.items()},
        # Both blocks are generated from the fields' own `vocabulary` markers (`base.py`), not listed
        # here. The list they replace drifted twice: it never learned about `recommendation_strength`
        # or `phenotype_category` when 0.5 added them, and it filed `actionability` under
        # `open_recommended` although `VariantRow` rejects a non-member — a drift in *closedness*,
        # which is why the marker carries that flag and why `actionability` now appears below.
        "vocabularies": _collect_vocabularies(closed=True),
        "open_recommended": _collect_vocabularies(closed=False),
        # Per-member prose, for the vocabularies where the member *name* cannot carry the whole rule.
        # The element rules (RM54) are the case that forced it: on a `Number=R` VCF field the
        # reference is element zero, so "the larger of the two" has two answers, and a vocabulary that
        # is silent about which one it means repeats the defect it was added to fix one level down.
        # Keyed by the same names as `vocabularies` above, so a consumer that found a member there
        # can look up what it means without a second lookup table of its own.
        #
        # Derived from the fields' own markers, like both blocks above it. It was a comprehension over
        # `VCF_POINTER_COMPANIONS` naming `ELEMENT_RULE_MEANINGS` in place — correct, and the only
        # channel that composition could reach was this block, so the prose never got to the per-table
        # `describe` an author filling one table runs (D1-4).
        "vocabulary_notes": _collect_vocabulary_notes(),
        # Alternative identity-column sets, any ONE of which satisfies a row's requirement. Field-level
        # `required` cannot express "rsid OR chrom+start" — that rule is a model validator — so a tool
        # listing required columns told an author a `variants.csv` row needed no identifier at all.
        "required_any_of": {
            name: [sorted(group) for group in model.REQUIRED_ANY_OF]
            for name, model in _ALL_MODELS.items()
            if getattr(model, "REQUIRED_ANY_OF", ())
        },
        "reserved_names": sorted(RESERVED_NAMES_0_4),
        # Field-ownership boundary for the `module:` block (S2): keys the format knows about but that
        # a *publishing registry* stamps, so an author must omit them. Distinct from `reserved_names`
        # (future module columns) — these will never be authored. A consumer strips them before
        # validation via `normalize.strip_authority_keys`; `module.version` is NOT here (it is a
        # genuine advisory authored field).
        "registry_stamped_keys": {
            key: IDENTITY_AUTHORITY_REASONS[key] for key in sorted(IDENTITY_AUTHORITY_KEYS)
        },
        "recommended_palette": {"colors": RECOMMENDED_COLORS, "icons": RECOMMENDED_ICONS},
    }

json_schemas

json_schemas() -> dict[str, Any]

Full JSON Schema per model (Pydantic model_json_schema()), for consumers that want the machine-validatable form rather than the compact authoring_reference() summary.

Source code in schema/src/just_dna_format/reference.py
def json_schemas() -> dict[str, Any]:
    """Full JSON Schema per model (Pydantic `model_json_schema()`), for consumers that want the
    machine-validatable form rather than the compact `authoring_reference()` summary."""
    return {name: model.model_json_schema() for name, model in _ALL_MODELS.items()}