Skip to content

just_dna_format.base

just_dna_format.base

Shared base for every authored-DSL row model (spec/binning/pgx/pgs).

Consolidates the boilerplate that was copy-pasted across the row models into one place:

  • the reserved-namespace guard — extra="forbid" plus the reject_reserved before-validator, so a reserved name fails with a specific diagnosis and any other unknown/misspelled column fails with the generic message (see vocab.reject_reserved); and
  • the field validators for the shared authored vocabulary — rsid, trait_efo_id, direction, clin_sig, stat_significance, evidence_level, finite-effect_size, genotype, the VCF field-pointer grammar (source_field/callable_from/quality_from) and the element rule that qualifies each of those (source_element/callable_element/quality_element).

Each field validator uses check_fields=False, so a subclass runs it only for the fields it actually declares (a model without clin_sig simply never runs the clin_sig check) and a model that adds one of these fields gets the correct validation for free — the per-field rules cannot drift model to model, which is exactly what the previous copy-paste risked. Field-specific rules (star-allele strings, measure bounds, PGS ancestry, the mtDNA legacy-reference guard, identifier completeness) stay on their own models.

genotype moved here in 0.5 when PharmVariantRow gained one: the grammar is the same grammar (a PharmGKB per-genotype clinical annotation describes the same diploid call a VariantRow does), and the DRY rule applies as soon as a validator is shared by two models. It is deliberately not widened for the symbolic alleles PharmGKB also carries (C/del, del/del) — those are RM5, and the enricher skips them rather than coercing them into a nucleotide grammar that cannot express them.

Dependency-light: imports only pydantic + the stdlib vocab leaf, and nothing in the package imports it back, so it introduces no cycle.

AuthoredModel

Bases: BaseModel

Base for authored-DSL rows: reserved-namespace guard + shared-vocabulary field validators.

genome_build property

genome_build: str

The assembly this row was loaded as being on. Read-only; see _genome_build.

with_genome_build

with_genome_build(genome_build: str) -> AuthoredModel

Tell this row which assembly it is on. Returns self, so a loader can map over rows.

Deliberately a method rather than a settable property: injecting the build is something a loader does once, at a known point, and making it look like an ordinary attribute assignment would invite it being done anywhere. just_dna_compiler.compiler._load_csv_rows is the caller.

A model that stamps an identity (_KEY_INCLUDES_ALTS set) re-derives its variant_key here, which is this loader's answer to the problem _restamp_for_build solves for VariantRow one tier up: a row is constructed from a CSV dict where module_spec.yaml is not in scope, so the key it froze at construction took derive_variant_key's GRCh38 default. Doing it at the point the build arrives means a positional table needs no compiler-side correction pass at all. VariantRow deliberately does not join in — it leaves _KEY_INCLUDES_ALTS unset and keeps its existing restamp, which also emits the "keyed by coordinate instead" warning a GRCh37 module must hear.

Source code in schema/src/just_dna_format/base.py
def with_genome_build(self, genome_build: str) -> "AuthoredModel":
    """Tell this row which assembly it is on. Returns `self`, so a loader can map over rows.

    Deliberately a method rather than a settable property: injecting the build is something a
    *loader* does once, at a known point, and making it look like an ordinary attribute assignment
    would invite it being done anywhere. `just_dna_compiler.compiler._load_csv_rows` is the caller.

    A model that stamps an identity (`_KEY_INCLUDES_ALTS` set) re-derives its `variant_key` here,
    which is this loader's answer to the problem `_restamp_for_build` solves for `VariantRow` one
    tier up: a row is constructed from a CSV dict where `module_spec.yaml` is not in scope, so the
    key it froze at construction took `derive_variant_key`'s GRCh38 default. Doing it at the point
    the build arrives means a positional table needs no compiler-side correction pass at all.
    `VariantRow` deliberately does not join in — it leaves `_KEY_INCLUDES_ALTS` unset and keeps its
    existing restamp, which also emits the "keyed by coordinate instead" warning a GRCh37 module
    must hear.
    """
    self._genome_build = genome_build
    if type(self)._KEY_INCLUDES_ALTS is not None:
        stamp_identity(self, keys_on_alts=self._KEY_INCLUDES_ALTS, freeze_authored=False)
    return self

case_insensitive_allele_fields

case_insensitive_allele_fields(
    model: type[BaseModel],
) -> frozenset[str]

The columns of model that content_signature case-folds — those marked CASE_INSENSITIVE_ALLELE.

Walked off the fields for the same reason content_identity_exclusions is: a hand-kept set one file over is the thing that drifts (@fieldnames-from-model).

Source code in schema/src/just_dna_format/base.py
def case_insensitive_allele_fields(model: type[BaseModel]) -> frozenset[str]:
    """The columns of `model` that `content_signature` case-folds — those marked `CASE_INSENSITIVE_ALLELE`.

    Walked off the fields for the same reason `content_identity_exclusions` is: a hand-kept set one
    file over is the thing that drifts (`@fieldnames-from-model`)."""
    return frozenset(
        name
        for name, field in model.model_fields.items()
        if isinstance(field.json_schema_extra, dict)
        and field.json_schema_extra.get("case_insensitive_allele")
    )

content_identity_exclusions

content_identity_exclusions(
    model: type[BaseModel],
) -> frozenset[str]

The columns of model that content_signature omits — those marked OUTSIDE_CONTENT_IDENTITY.

Walked off the fields, never listed beside them, for the same reason authored_field_names is: a hand-kept set one file over is the thing that drifts. Empty for every model but the overlay row today, and a test pins that equality over the whole registry.

Source code in schema/src/just_dna_format/base.py
def content_identity_exclusions(model: type[BaseModel]) -> frozenset[str]:
    """The columns of `model` that `content_signature` omits — those marked `OUTSIDE_CONTENT_IDENTITY`.

    Walked off the fields, never listed beside them, for the same reason `authored_field_names` is:
    a hand-kept set one file over is the thing that drifts. Empty for every model but the overlay
    row today, and a test pins that equality over the whole registry."""
    return frozenset(
        name
        for name, field in model.model_fields.items()
        if isinstance(field.json_schema_extra, dict)
        and field.json_schema_extra.get("outside_content_identity")
    )

authored_field_names

authored_field_names(model: type[BaseModel]) -> list[str]

The columns a human actually authors for model, in field-declaration order.

model_fields minus anything marked COMPILER_MANAGED. Every generator over the authored surface — a blank template, a drafted CSV's header — reads its columns through here, so a field that stops being authored stops being offered in the same commit that marks it.

Source code in schema/src/just_dna_format/base.py
def authored_field_names(model: type[BaseModel]) -> list[str]:
    """The columns a human actually authors for `model`, in field-declaration order.

    `model_fields` minus anything marked `COMPILER_MANAGED`. Every generator over the authored
    surface — a blank template, a drafted CSV's header — reads its columns through here, so a field
    that stops being authored stops being offered in the same commit that marks it."""
    return [
        name
        for name, field in model.model_fields.items()
        if not (isinstance(field.json_schema_extra, dict) and field.json_schema_extra.get("compiler_managed"))
    ]

merge_key

merge_key(row: BaseModel) -> tuple

The key that decides two rows of a machine-produced table are the same row.

The enricher pass that writes each sidecar merges rather than clobbers, so it holds an existing dict and needs this tuple; the pass reads it from here rather than restating it, which is the half that makes the pass and hints.key_fields unable to disagree (S51). Restated, they were two statements of one fact — and the fact was legible nowhere outside the pass's own body, so every consumer re-deriving a sidecar was guessing it.

_KEY_FALLBACK_FIELDS, where a kind declares one, is used when every primary member is null: GeneValidityRow keys on the source's own assertion_id and falls back to the gene's grain when the source published none. The two levels are tagged ("id" / "grain") so a grain tuple can never collide with an id that happens to equal it.

Raises AttributeError for a model declaring no key, which is the honest failure: a caller reaching here for an unkeyed kind has a bug, and a silent () would merge every row into one.

An EMPTY _KEY_FIELDS is the same bug and used to slip through (RM213). MeasureBinRow declares () as a base-class default meaning subclasses set this, and every subclass does — but a bare MeasureBinRow, or a future kind that inherits the default and forgets, reached here and got () back. Two rows differing in every column returned equal keys, which is exactly the collapse the paragraph above says this function raises to prevent. Measured, not argued.

hints.table_key treats the same falsy value as no declared key and returns None; that is the right answer to a different question (does this table publish a key?) and is deliberately not changed. Here the question is what two rows' identity IS, and there is no such thing as an empty answer to it.

Source code in schema/src/just_dna_format/base.py
def merge_key(row: BaseModel) -> tuple:
    """The key that decides two rows of a machine-produced table are the same row.

    The enricher pass that writes each sidecar merges rather than clobbers, so it holds an `existing`
    dict and needs this tuple; **the pass reads it from here rather than restating it**, which is the
    half that makes the pass and `hints.key_fields` unable to disagree (S51). Restated, they were two
    statements of one fact — and the fact was legible nowhere outside the pass's own body, so every
    consumer re-deriving a sidecar was guessing it.

    `_KEY_FALLBACK_FIELDS`, where a kind declares one, is used when **every** primary member is null:
    `GeneValidityRow` keys on the source's own `assertion_id` and falls back to the gene's grain when
    the source published none. The two levels are tagged (`"id"` / `"grain"`) so a grain tuple can
    never collide with an id that happens to equal it.

    Raises `AttributeError` for a model declaring no key, which is the honest failure: a caller
    reaching here for an unkeyed kind has a bug, and a silent `()` would merge every row into one.

    **An EMPTY `_KEY_FIELDS` is the same bug and used to slip through** (RM213). `MeasureBinRow`
    declares `()` as a base-class default meaning *subclasses set this*, and every subclass does — but
    a bare `MeasureBinRow`, or a future kind that inherits the default and forgets, reached here and
    got `()` back. Two rows differing in every column returned equal keys, which is exactly the
    collapse the paragraph above says this function raises to prevent. Measured, not argued.

    `hints.table_key` treats the same falsy value as *no declared key* and returns `None`; that is the
    right answer to a different question (does this table publish a key?) and is deliberately not
    changed. Here the question is what two rows' identity IS, and there is no such thing as an empty
    answer to it.
    """
    declared: tuple[str, ...] = row._KEY_FIELDS  # type: ignore[attr-defined]
    if not declared:
        raise AttributeError(
            f"{type(row).__name__} declares an empty `_KEY_FIELDS`, so it has no merge identity. "
            f"A subclass pins the real key; merging on `()` would collapse every row into one."
        )
    primary = tuple(getattr(row, name) for name in declared)
    fallback: tuple[str, ...] = getattr(row, "_KEY_FALLBACK_FIELDS", ())
    if not fallback or any(v is not None and v != "" for v in primary):
        return ("id", *primary) if fallback else primary
    return ("grain", *(getattr(row, name) for name in fallback))

accepts_none

accepts_none(annotation: Any) -> bool

Does this annotation admit None? (Optional[str] yes; a defaulted bare str/bool no.)

Source code in schema/src/just_dna_format/base.py
def accepts_none(annotation: Any) -> bool:
    """Does this annotation admit `None`? (`Optional[str]` yes; a defaulted bare `str`/`bool` no.)"""
    return annotation is type(None) or type(None) in get_args(annotation)

field_category

field_category(model: type[BaseModel], name: str) -> str

required | defaulted | optional — the three-way split an authoring surface must respect.

The middle category is the one that bites. MeasureBinRow.measure_kind (str, default "repeat_count") and unresolved (bool, default False) are not required, so pydantic's is_required() says False — but _load_csv_rows turns an empty cell into None and keeps the key, so the model receives None instead of its default and fails on type. An author who filled exactly the columns a two-way required flag named got a rejection about a column nobody had mentioned. A defaulted cell has to be written out with its default rather than left blank.

It lives here rather than in the compiler because two surfaces answer this question and they drifted: just_dna_compiler.draft was fixed to the three-way split and reference.authoring_reference — the drift-proof description consumers render instead of a hand-kept spec dump — was still emitting the two-way one. Both now read this. The format tier is the only place both can import from, and this needs nothing but pydantic.

Source code in schema/src/just_dna_format/base.py
def field_category(model: type[BaseModel], name: str) -> str:
    """`required` | `defaulted` | `optional` — the three-way split an authoring surface must respect.

    The middle category is the one that bites. `MeasureBinRow.measure_kind` (`str`, default
    `"repeat_count"`) and `unresolved` (`bool`, default `False`) are *not* required, so pydantic's
    `is_required()` says `False` — but `_load_csv_rows` turns an empty cell into `None` and **keeps
    the key**, so the model receives `None` instead of its default and fails on type. An author who
    filled exactly the columns a two-way `required` flag named got a rejection about a column nobody
    had mentioned. A `defaulted` cell has to be written out with its default rather than left blank.

    It lives **here** rather than in the compiler because two surfaces answer this question and they
    drifted: `just_dna_compiler.draft` was fixed to the three-way split and `reference.authoring_reference`
    — the drift-proof description consumers render *instead of* a hand-kept spec dump — was still
    emitting the two-way one. Both now read this. The format tier is the only place both can import
    from, and this needs nothing but pydantic.
    """
    field = model.model_fields[name]
    if field.is_required():
        return "required"
    return "optional" if accepts_none(field.annotation) else "defaulted"

vocabulary

vocabulary(
    name: str,
    options: frozenset[str],
    *,
    closed: bool = True,
    notes: dict[str, str] | None = None,
) -> dict[str, object]

Mark a field as drawn from options, for tools that offer an author the valid values.

closed=True means a validator rejects anything outside the set; closed=False marks the recommended-but-open sets (RECOMMENDED_EFFECT_MEASURES, RECOMMENDED_AUTHOR_KINDS), where the members are suggestions and a novel value is legal. VALID_ACTIONABILITY used to be listed here and is closed — the drift this very flag exists to prevent, written into the docstring that defines it (RM225). A consumer must be able to tell "pick one of these" from "these are suggestions", so the two are one marker with a flag rather than two markers.

notes is per-member prose, for the vocabularies where the member name cannot carry the whole rule — ELEMENT_RULE_MEANINGS is the case that forced it, since largest has two answers on a Number=R field and only a sentence can say which one it means. It rides on the marker for the same reason options does: a name to look up would need a central registry, and vocab cannot import pgx (the cycle this module's dependency note exists to avoid), so a registry anywhere else is a second hand-kept list. Carrying the prose here is what lets every surface that reads field_vocabularies print it — before this, the meanings reached the whole-schema reference and nothing else, so describe <kind>, the command an author filling one table runs, listed eight members and could not say which of them counted the reference element.

Omitted from the marker when there is none, so the shape of the 20-odd vocabularies that need no prose is unchanged. Members keep their declaration order rather than being sorted: it is the reading order the constant was written in, and it is as deterministic as sorted(options) is.

Source code in schema/src/just_dna_format/base.py
def vocabulary(
    name: str,
    options: frozenset[str],
    *,
    closed: bool = True,
    notes: dict[str, str] | None = None,
) -> dict[str, object]:
    """Mark a field as drawn from `options`, for tools that offer an author the valid values.

    `closed=True` means a validator rejects anything outside the set; `closed=False` marks the
    recommended-but-open sets (`RECOMMENDED_EFFECT_MEASURES`, `RECOMMENDED_AUTHOR_KINDS`), where the
    members are suggestions and a novel value is legal. `VALID_ACTIONABILITY` used to be listed here
    and is **closed** — the drift this very flag exists to prevent, written into the docstring that
    defines it (RM225). A
    consumer must be able to tell "pick one of these" from "these are suggestions", so the two are one
    marker with a flag rather than two markers.

    `notes` is per-member prose, for the vocabularies where the member *name* cannot carry the whole
    rule — `ELEMENT_RULE_MEANINGS` is the case that forced it, since `largest` has two answers on a
    `Number=R` field and only a sentence can say which one it means. It rides on the marker for the
    same reason `options` does: a name to look up would need a central registry, and `vocab` cannot
    import `pgx` (the cycle this module's dependency note exists to avoid), so a registry anywhere
    else is a second hand-kept list. Carrying the prose here is what lets *every* surface that reads
    `field_vocabularies` print it — before this, the meanings reached the whole-schema reference and
    nothing else, so `describe <kind>`, the command an author filling one table runs, listed eight
    members and could not say which of them counted the reference element.

    Omitted from the marker when there is none, so the shape of the 20-odd vocabularies that need no
    prose is unchanged. Members keep their declaration order rather than being sorted: it is the
    reading order the constant was written in, and it is as deterministic as `sorted(options)` is."""
    marker: dict[str, object] = {"name": name, "options": sorted(options), "closed": closed}
    if notes:
        marker["notes"] = dict(notes)
    return {"vocabulary": marker}

since

since(version: str) -> dict[str, object]

Mark the release a field first appeared in — Field(json_schema_extra=since("0.6.5")) (RM146).

The finding this answers. A module authored on 0.6.6 was sent to a deployment running 0.6.1, which runs validate_spec server-side and reported, verbatim: studies.csv line 2 [curator]: Extra inputs are not permitted. StudyRow.curator is ours, added in 0.6.5. A genuine typo produces the byte-identical shape — [curatr] — and the two want opposite actions from an author: upgrade the reader, or fix the cell. The message is pydantic's under extra="forbid", so it cannot be reworded into carrying the distinction: the information was not in the model at all.

On the field, not in a roster. A list keyed like release_records was the alternative and loses on the rule this repo keeps relearning — a hand-kept list beside a model is a second statement of one fact, and it is the copy that goes stale (@fieldnames-from-model, @registry-completeness). Declared here it travels with the field through every rename and move, and test_first_seen.py asserts an equality over the walked registry, so the next column added cannot omit one.

The answer is per (model, field), never per name, which curator is the worked example of: it is on VariantRow from 0.2.0 and gains its StudyRow twin only in 0.6.5. A roster keyed by column name would give one answer for two facts.

Composes with vocabulary() rather than replacing it — both are entries in one json_schema_extra dict, so a field can carry either or both:

Field(json_schema_extra={**vocabulary("state", VALID_STATES), **since("0.2.0")})
Source code in schema/src/just_dna_format/base.py
def since(version: str) -> dict[str, object]:
    """Mark the release a field first appeared in — `Field(json_schema_extra=since("0.6.5"))` (RM146).

    **The finding this answers.** A module authored on 0.6.6 was sent to a deployment running 0.6.1,
    which runs `validate_spec` server-side and reported, verbatim:
    `studies.csv line 2 [curator]: Extra inputs are not permitted`. `StudyRow.curator` is ours, added
    in 0.6.5. A genuine typo produces the byte-identical shape — `[curatr]` — and the two want
    **opposite actions** from an author: upgrade the reader, or fix the cell. The message is pydantic's
    under `extra="forbid"`, so it cannot be reworded into carrying the distinction: the information was
    not in the model at all.

    **On the field, not in a roster.** A list keyed like `release_records` was the alternative and
    loses on the rule this repo keeps relearning — a hand-kept list beside a model is a second
    statement of one fact, and it is the copy that goes stale (`@fieldnames-from-model`,
    `@registry-completeness`). Declared here it travels with the field through every rename and move,
    and `test_first_seen.py` asserts an **equality over the walked registry**, so the next column added
    cannot omit one.

    **The answer is per (model, field), never per name**, which `curator` is the worked example of: it
    is on `VariantRow` from 0.2.0 and gains its `StudyRow` twin only in 0.6.5. A roster keyed by column
    name would give one answer for two facts.

    Composes with `vocabulary()` rather than replacing it — both are entries in one
    `json_schema_extra` dict, so a field can carry either or both:

        Field(json_schema_extra={**vocabulary("state", VALID_STATES), **since("0.2.0")})
    """
    return {"first_seen": version}

field_first_seen

field_first_seen(model: type[BaseModel]) -> dict[str, str]

{field_name: release} for every field of model that declares one (RM146).

The public reader, so a consumer rendering our findings can answer when did this column appear offline rather than parsing model_fields themselves — which is what the reporter would otherwise have had to do, and what their own rulebook forbids.

Source code in schema/src/just_dna_format/base.py
def field_first_seen(model: type[BaseModel]) -> dict[str, str]:
    """`{field_name: release}` for every field of `model` that declares one (RM146).

    The public reader, so a consumer rendering our findings can answer *when did this column appear*
    **offline** rather than parsing `model_fields` themselves — which is what the reporter would
    otherwise have had to do, and what their own rulebook forbids.
    """
    found: dict[str, str] = {}
    for name, field in model.model_fields.items():
        extra = field.json_schema_extra
        if isinstance(extra, dict) and isinstance(extra.get("first_seen"), str):
            found[name] = extra["first_seen"]
    return found

field_vocabularies

field_vocabularies(
    model: type[BaseModel],
) -> dict[str, dict]

{field_name: {name, options, closed[, notes]}} for every vocabulary-bound field of model.

The single route to "what may this cell contain" — authoring_reference() and any authoring tool read it, and neither keeps a list of its own. Covers both binding sites: a field marked at its own declaration, and a field whose vocabulary is enforced by AuthoredModel's shared validators.

Source code in schema/src/just_dna_format/base.py
def field_vocabularies(model: type[BaseModel]) -> dict[str, dict]:
    """`{field_name: {name, options, closed[, notes]}}` for every vocabulary-bound field of `model`.

    The single route to "what may this cell contain" — `authoring_reference()` and any authoring tool
    read it, and neither keeps a list of its own. Covers both binding sites: a field marked at its own
    declaration, and a field whose vocabulary is enforced by `AuthoredModel`'s shared validators."""
    found: dict[str, dict] = {}
    for name, field in model.model_fields.items():
        marker = (
            field.json_schema_extra.get("vocabulary") if isinstance(field.json_schema_extra, dict) else None
        )
        if isinstance(marker, dict):
            found[name] = marker
        elif name in SHARED_VOCABULARIES:
            found[name] = vocabulary(
                name, SHARED_VOCABULARIES[name], notes=SHARED_VOCABULARY_NOTES.get(name)
            )["vocabulary"]
    return found

derive_variant_key

derive_variant_key(
    rsid: str | None,
    chrom: str | None,
    start: int | None,
    ref: str | None,
    alts: str | None = None,
    *,
    build: str = "GRCh38",
) -> str

The natural identity for a variant-ish row: the rsid when present, else the coordinate.

Three cases, in precedence order:

  1. rsid — an rsid row keeps its rsid, unchanged. (A dbSNP id is position/multi-allelic-level, not per-allele, which is why clinical identity keys on this plus genotype, never on rsid.)
  2. A resolved single-base substitution (0.5) — the key is its GA4GH VRS allele id, ga4gh:VA.…, minted by vrs.derive_vrs_allele_id. This is a content-addressed identity: it names the allele by the digest of the exact reference sequence it sits on, so it is build-naming rather than build-ambiguous — the property RM15 was waiting for before coordinate identity could be reconsidered. Byte-identical to the ids gnomAD and ClinGen serve, so it joins against them directly instead of needing a translation table.
  3. Everything else — the coordinate key chrom:start:ref, or chrom:start:ref:alts when an alt is given, so two distinct alleles at one locus (an insertion C>CAAAG beside a deletion C>CA, a benign C>G beside a pathogenic C>A) do not collide. alts is normalized (its comma-separated alleles sorted) so the key is stable regardless of authored order.

Case 2 covers substitutions only — an indel, an MNV, a multi-allelic cell and a contig outside the primary assembly all fall through to case 3, because a VRS allele id is defined over the fully justified allele and justifying an indel needs the reference sequence, which this tier will never fetch (Principle 2). See vrs.derive_vrs_allele_id. That split is deliberate and permanent-shaped: an id is minted only where it can be minted correctly, and the fallback is the same key these rows already had.

Single source of truth shared by VariantRow (which freezes the result into a stored column so resolution can never re-key a row) and the one-to-many expansion re-keying. StudyRow and the position-level matching helpers deliberately call this without alts — a study is position/rsid evidence and matches a variant at chrom:start:ref regardless of which allele it carries — and so are never affected by case 2. See docs/COMPILER.md: the frozen variant_key keeps a position-only row that later resolves to an rsid from flipping its identity, and lets a one-to-many rsid expand to distinct coord-keyed rows (Principle 7).

build is the assembly the coordinate is in; only GRCh38 has a refget table today, and any other build falls through to case 3 rather than minting an id that would claim the wrong sequence.

Source code in schema/src/just_dna_format/base.py
def derive_variant_key(
    rsid: str | None,
    chrom: str | None,
    start: int | None,
    ref: str | None,
    alts: str | None = None,
    *,
    build: str = "GRCh38",
) -> str:
    """The natural identity for a variant-ish row: the rsid when present, else the coordinate.

    Three cases, in precedence order:

    1. **rsid** — an rsid row keeps its rsid, unchanged. (A dbSNP id is position/multi-allelic-level,
       not per-allele, which is why clinical identity keys on this *plus* genotype, never on rsid.)
    2. **A resolved single-base substitution** (0.5) — the key is its **GA4GH VRS allele id**,
       `ga4gh:VA.…`, minted by `vrs.derive_vrs_allele_id`. This is a *content-addressed* identity: it
       names the allele by the digest of the exact reference sequence it sits on, so it is
       build-naming rather than build-ambiguous — the property RM15 was waiting for before coordinate
       identity could be reconsidered. Byte-identical to the ids gnomAD and ClinGen serve, so it joins
       against them directly instead of needing a translation table.
    3. **Everything else** — the coordinate key `chrom:start:ref`, or `chrom:start:ref:alts` when an
       alt is given, so two distinct alleles at one locus (an insertion `C>CAAAG` beside a deletion
       `C>CA`, a benign `C>G` beside a pathogenic `C>A`) do not collide. `alts` is normalized (its
       comma-separated alleles sorted) so the key is stable regardless of authored order.

    Case 2 covers substitutions only — an indel, an MNV, a multi-allelic cell and a contig outside the
    primary assembly all fall through to case 3, because a VRS allele id is defined over the *fully
    justified* allele and justifying an indel needs the reference sequence, which this tier will never
    fetch (Principle 2). See `vrs.derive_vrs_allele_id`. That split is deliberate and permanent-shaped:
    an id is minted only where it can be minted *correctly*, and the fallback is the same key these
    rows already had.

    Single source of truth shared by `VariantRow` (which *freezes* the result into a stored column so
    resolution can never re-key a row) and the one-to-many expansion re-keying. `StudyRow` and the
    position-level *matching* helpers deliberately call this **without** `alts` — a study is
    position/rsid evidence and matches a variant at `chrom:start:ref` regardless of which allele it
    carries — and so are never affected by case 2. See docs/COMPILER.md: the frozen `variant_key` keeps
    a position-only row that later resolves to an rsid from flipping its identity, and lets a
    one-to-many rsid expand to distinct coord-keyed rows (Principle 7).

    `build` is the assembly the coordinate is in; only GRCh38 has a refget table today, and any other
    build falls through to case 3 rather than minting an id that would claim the wrong sequence.
    """
    if rsid is not None:
        return rsid
    if alts and "," not in alts:
        vrs_id = _mint_vrs_key(chrom, start, ref, alts.strip(), build)
        if vrs_id is not None:
            return vrs_id
    base = f"{chrom}:{start}:{ref}"
    alts_norm = ",".join(sorted(a.strip() for a in alts.split(",") if a.strip())) if alts else ""
    return f"{base}:{alts_norm}" if alts_norm else base

stamped_identity_field

stamped_identity_field(
    description: str,
    *,
    default: Any = None,
    first_seen: str,
) -> Any

A Field(...) for a compiler-stamped column that is not part of the authored content.

Three properties, and the middle one is the non-obvious one:

  • COMPILER_MANAGED, so authored_field_names drops it — no author is offered the column, no drafted template writes it, and reverse_module's generic _write_table_csv never re-emits it.
  • exclude=True, so it stays out of model_dump() and therefore out of content_signature. A stamped value is a pure function of the authored cells, so it adds nothing to a content identity — and including it would move the signature of every already-published module carrying one of these tables, which is the one thing a content-dedup key may not do. _build_table reads the field off the model directly (not through model_dump()), so the column still reaches parquet. VariantRow.variant_key/authored_ident are not excluded and are inside content_signature today; that is a grandfathered inconsistency, not a precedent — un-excluding them here, or excluding them there, moves published signatures either way, so the asymmetry is carried until a major. Anything declared through this helper is on the right side of it, including VariantRow's own locus_index/locus_count (RM87), which are stamped on the same model without repeating the defect.
  • a fresh FieldInfo per call, because pydantic binds one to the model that declares it.

first_seen is required rather than defaulted (RM146). A compiler-stamped column is still a column an older reader refuses under extra="forbid", so it owes the same answer as an authored one — and the guard walks model_fields, which does not distinguish them. Defaulting it would let the next stamped column inherit a version nobody measured, which is the whole failure mode.

default is the caller's, not this helper's. It was hard-coded to None while the only users were the three 0.4-family positional models (RM43), where an unstamped identity column genuinely has no value. RM87's locus_count defaults to 1 instead, because a row that was never expanded really does resolve to one locus and locus_count > 1 has to be a predicate a reader can apply while holding a single row — a None/0 default would put a second "undetermined" state back into the column the item exists to make determinate.

Source code in schema/src/just_dna_format/base.py
def stamped_identity_field(description: str, *, default: Any = None, first_seen: str) -> Any:
    """A `Field(...)` for a compiler-stamped column that is not part of the authored content.

    Three properties, and the middle one is the non-obvious one:

    * `COMPILER_MANAGED`, so `authored_field_names` drops it — no author is offered the column, no
      drafted template writes it, and `reverse_module`'s generic `_write_table_csv` never re-emits it.
    * **`exclude=True`, so it stays out of `model_dump()` and therefore out of `content_signature`.**
      A stamped value is a pure function of the authored cells, so it adds nothing to a *content*
      identity — and including it would move the signature of every already-published module carrying
      one of these tables, which is the one thing a content-dedup key may not do. `_build_table` reads
      the field off the model directly (not through `model_dump()`), so the column still reaches
      parquet. `VariantRow.variant_key`/`authored_ident` are **not** excluded and are inside
      `content_signature` today; that is a grandfathered inconsistency, not a precedent — un-excluding
      them here, or excluding them there, moves published signatures either way, so the asymmetry is
      carried until a major. Anything declared *through this helper* is on the right side of it,
      including `VariantRow`'s own `locus_index`/`locus_count` (RM87), which are stamped on the same
      model without repeating the defect.
    * a fresh `FieldInfo` per call, because pydantic binds one to the model that declares it.

    `first_seen` is **required rather than defaulted** (RM146). A compiler-stamped column is still a
    column an older reader refuses under `extra="forbid"`, so it owes the same answer as an authored
    one — and the guard walks `model_fields`, which does not distinguish them. Defaulting it would let
    the next stamped column inherit a version nobody measured, which is the whole failure mode.

    `default` is the **caller's**, not this helper's. It was hard-coded to `None` while the only users
    were the three 0.4-family positional models (RM43), where an unstamped identity column genuinely
    has no value. RM87's `locus_count` defaults to `1` instead, because a row that was never expanded
    really does resolve to one locus and `locus_count > 1` has to be a predicate a reader can apply
    while holding a single row — a `None`/`0` default would put a second "undetermined" state back
    into the column the item exists to make determinate.
    """
    return Field(
        default=default,
        exclude=True,
        json_schema_extra={**COMPILER_MANAGED, **since(first_seen)},
        description=description,
    )

reject_compiler_filled

reject_compiler_filled(
    data: object, fields: dict[str, Any], *, what: str
) -> None

Refuse an identity column this model has the compiler fill (RM43), with a diagnosis.

alts is the only such column today: authored and required on variants.csv, authored and optional on heteroplasmy.csv, and compiler-filled on pharm_variants.csv/haplotypes.csv, where a pharm annotation matches a variant at chrom:start:ref regardless of allele. Writing it there by analogy with variants.csv is a plausible confusion rather than a typo, and before this guard it was silently accepted: extra="forbid" cannot see it, the value entered authored_ident, and _write_table_csv then dropped it — so compile → reverse → compile was not a fixed point. It failed loudly before 0.6 added the column, so a bare accept is a regression however sensible the input.

Scoped to IDENTITY_FIELDS deliberately, which is what leaves variant_key/authored_ident accepted-and-overwritten: those two are stamped rather than filled, and VariantRow has tolerated them since 0.5 on the "authored values ignored, no foot-gun" rule. Nothing is lost by ignoring a stamped value; something is lost by ignoring a filled one.

Source code in schema/src/just_dna_format/base.py
def reject_compiler_filled(data: object, fields: dict[str, Any], *, what: str) -> None:
    """Refuse an **identity** column this model has the compiler fill (RM43), with a diagnosis.

    `alts` is the only such column today: authored and required on `variants.csv`, authored and
    optional on `heteroplasmy.csv`, and **compiler-filled** on `pharm_variants.csv`/`haplotypes.csv`,
    where a pharm annotation matches a variant at `chrom:start:ref` regardless of allele. Writing it
    there by analogy with `variants.csv` is a plausible confusion rather than a typo, and before this
    guard it was silently *accepted*: `extra="forbid"` cannot see it, the value entered
    `authored_ident`, and `_write_table_csv` then dropped it — so `compile → reverse → compile` was
    not a fixed point. It failed loudly before 0.6 added the column, so a bare accept is a regression
    however sensible the input.

    Scoped to `IDENTITY_FIELDS` deliberately, which is what leaves `variant_key`/`authored_ident`
    accepted-and-overwritten: those two are *stamped* rather than filled, and `VariantRow` has
    tolerated them since 0.5 on the "authored values ignored, no foot-gun" rule. Nothing is lost by
    ignoring a stamped value; something is lost by ignoring a filled one.
    """
    if not isinstance(data, dict):
        return
    hits = sorted(
        name
        for name in data
        if name in IDENTITY_FIELDS
        and isinstance((field := fields.get(name)), FieldInfo)
        and isinstance(field.json_schema_extra, dict)
        and field.json_schema_extra.get("compiler_managed")
    )
    if not hits:
        return
    raise ValueError(
        f"{what}: {', '.join(repr(h) for h in hits)} is filled by the compiler on this table, from "
        f"the injected resolution.csv, and is not an authored column here — remove it. It IS authored "
        f"on variants.csv and heteroplasmy.csv, which is the confusion this message exists for; here "
        f"the row's identity is rsid, or chrom + start (+ ref), and the allele list is looked up."
    )

authored_identity

authored_identity(row: AuthoredModel) -> list[str]

Which of IDENTITY_FIELDS this row's model declares and the author actually filled.

Source code in schema/src/just_dna_format/base.py
def authored_identity(row: "AuthoredModel") -> list[str]:
    """Which of `IDENTITY_FIELDS` this row's model declares *and* the author actually filled."""
    fields = type(row).model_fields
    return [name for name in IDENTITY_FIELDS if name in fields and getattr(row, name) is not None]

stamp_identity

stamp_identity(
    row: AuthoredModel,
    *,
    keys_on_alts: bool,
    freeze_authored: bool,
) -> None

Stamp variant_key (and, at construction, authored_ident) onto a positional row.

The key is derived from the authored subset only, never from the row's current cells, which is what makes it survive the compile-time coordinate fill: once authored_ident is frozen, filling chrom/start/ref/alts cannot re-key the row however many times this runs. That is the same guarantee VariantRow gets from _freeze_identity never re-running on model_copy, arrived at from the other direction because these rows are filled in place rather than copied.

freeze_authored is True exactly once, at construction, and authored values are overwritten rather than trusted — a CSV that carries a variant_key/authored_ident column of its own does not get to declare its own identity (VariantRow calls that "no foot-gun"). with_genome_build re-derives the key only, because the assembly changes what a coordinate means and cannot change what the author wrote.

Source code in schema/src/just_dna_format/base.py
def stamp_identity(row: "AuthoredModel", *, keys_on_alts: bool, freeze_authored: bool) -> None:
    """Stamp `variant_key` (and, at construction, `authored_ident`) onto a positional row.

    The key is derived from the **authored subset only**, never from the row's current cells, which is
    what makes it survive the compile-time coordinate fill: once `authored_ident` is frozen, filling
    `chrom`/`start`/`ref`/`alts` cannot re-key the row however many times this runs. That is the same
    guarantee `VariantRow` gets from `_freeze_identity` never re-running on `model_copy`, arrived at
    from the other direction because these rows are filled **in place** rather than copied.

    `freeze_authored` is `True` exactly once, at construction, and authored values are overwritten
    rather than trusted — a CSV that carries a `variant_key`/`authored_ident` column of its own does
    not get to declare its own identity (`VariantRow` calls that "no foot-gun"). `with_genome_build`
    re-derives the *key* only, because the assembly changes what a coordinate means and cannot change
    what the author wrote.
    """
    if freeze_authored:
        row.authored_ident = authored_identity(row)
    authored = set(row.authored_ident or ())

    def cell(name: str) -> Any:
        return getattr(row, name, None) if name in authored else None

    # A row that names no variant at all keys as `None` — the pre-0.5 `heteroplasmy.csv` shape, where
    # the bins are about a gene. The two PGx models forbid it by validator (rsid, or chrom+start), so
    # for them this branch is unreachable and the key is always a string.
    if cell("rsid") is None and cell("start") is None:
        row.variant_key = None
        return
    row.variant_key = derive_variant_key(
        cell("rsid"),
        cell("chrom"),
        cell("start"),
        cell("ref"),
        cell("alts") if keys_on_alts else None,
        build=row.genome_build,
    )

genotype_allele_ok

genotype_allele_ok(allele: str) -> bool

Whether one member of a genotype is spellable: a nucleotide string, a symbolic allele, or *.

The single place _validate_genotype decides what an allele is, so the three arms of the grammar (phased pair, hemizygous single, unphased pair) cannot drift apart — they did not, but the check was written out three times and a widening had to be applied three times with it.

Symbolic alleles joined in 0.6 (RM5): a genotype may name one (<DEL:1500>/A — a heterozygous deletion), because that is exactly the call a consumer reading a structural VCF has in hand. A lengthless one passes here and is refused by the compiler; see vocab.validate_allele for why the split is forced rather than chosen.

* joined beside them in 0.6 (RM59), and the two arms are deliberately not one arm. They are checked by different predicates against different modules of meaning, and the separation is the item rather than an implementation detail: parse_symbolic_allele answers which variant is this, unspelled, and is_unobservable_allele answers whether this sample's allele could be seen at all — the callability axis requires_callable already owns (P5). Folding * into SYMBOLIC_ALLELE_TYPES would give one syntax two meanings and hand * a length, an event and a place it does not have. See alleles' RM59 section for the full argument.

Why the genotype column and not the allele columns: vocab.validate_allele (HaplotypeRow.allele, VariantRow.effect_allele) still refuses *, and must. Those columns name the allele a rule is about, and a rule about an allele nobody observed states nothing. A genotype is the one place the observation itself is written down, which is why it is the one place * belongs.

Source code in schema/src/just_dna_format/base.py
def genotype_allele_ok(allele: str) -> bool:
    """Whether one member of a genotype is spellable: a nucleotide string, a symbolic allele, or `*`.

    The single place `_validate_genotype` decides what an allele *is*, so the three arms of the
    grammar (phased pair, hemizygous single, unphased pair) cannot drift apart — they did not, but
    the check was written out three times and a widening had to be applied three times with it.

    Symbolic alleles joined in 0.6 (RM5): a genotype may name one (`<DEL:1500>/A` — a heterozygous
    deletion), because that is exactly the call a consumer reading a structural VCF has in hand.
    A *lengthless* one passes here and is refused by the compiler; see `vocab.validate_allele` for
    why the split is forced rather than chosen.

    **`*` joined beside them in 0.6 (RM59), and the two arms are deliberately not one arm.** They are
    checked by different predicates against different modules of meaning, and the separation is the
    item rather than an implementation detail: `parse_symbolic_allele` answers *which variant is this,
    unspelled*, and `is_unobservable_allele` answers *whether this sample's allele could be seen at
    all* — the callability axis `requires_callable` already owns (P5). Folding `*` into
    `SYMBOLIC_ALLELE_TYPES` would give one syntax two meanings and hand `*` a length, an event and a
    place it does not have. See `alleles`' RM59 section for the full argument.

    Why the genotype column and not the allele columns: `vocab.validate_allele` (`HaplotypeRow.allele`,
    `VariantRow.effect_allele`) still refuses `*`, and must. Those columns name the allele a *rule* is
    about, and a rule about an allele nobody observed states nothing. A genotype is the one place the
    observation itself is written down, which is why it is the one place `*` belongs.
    """
    return (
        bool(ALLELE_PATTERN.match(allele))
        or parse_symbolic_allele(allele) is not None
        or is_unobservable_allele(allele)
    )