Skip to content

just_dna_format.vocab

just_dna_format.vocab

Shared constrained vocabularies, identifier patterns, and reusable validator helpers.

A dependency-light leaf (stdlib only) so every authored-DSL model — spec (variants/studies), binning (the measure→phenotype primitive), pgx (star-alleles), and pgs — validates against one source of truth for the orthogonal axes and identifier grammars. Per CONSTITUTION Principle 6, constrained vocabularies are frozenset[str] + a validator, never Enum/Literal.

spec re-exports the names it historically owned, so existing imports (from just_dna_format.spec import VALID_DIRECTIONS) keep working unchanged.

validate_phenotype_categories

validate_phenotype_categories(
    value: str | None,
    field_name: str = "phenotype_category",
) -> str | None

Validate a multi-valued phenotype-category cell against VALID_PHENOTYPE_CATEGORIES.

Accepts ClinPGx's own spellings (Metabolism/PK) case-insensitively and normalizes them to the vocabulary member (metabolism_pk), because the authored DSL should not make a human transcribe a slash-and-caps token exactly.

Source code in schema/src/just_dna_format/vocab.py
def validate_phenotype_categories(value: str | None, field_name: str = "phenotype_category") -> str | None:
    """Validate a multi-valued phenotype-category cell against `VALID_PHENOTYPE_CATEGORIES`.

    Accepts ClinPGx's own spellings (`Metabolism/PK`) case-insensitively and normalizes them to the
    vocabulary member (`metabolism_pk`), because the authored DSL should not make a human transcribe
    a slash-and-caps token exactly.
    """
    if value is None:
        return value
    normalized: list[str] = []
    for token in MULTI_SEP.split(value):
        token = token.strip()
        if not token:
            continue
        canonical = token.lower().replace("/", "_").replace(" ", "_").replace("-", "_")
        if canonical not in VALID_PHENOTYPE_CATEGORIES:
            raise ValueError(
                f"{field_name} tokens must be one of {sorted(VALID_PHENOTYPE_CATEGORIES)}, got: {token!r}"
            )
        normalized.append(canonical)
    return ";".join(normalized) if normalized else None

reject_template_placeholders

reject_template_placeholders(
    data: object, *, what: str = "row"
) -> object

A mode="before" guard: refuse any cell still carrying TEMPLATE_PLACEHOLDER.

Runs before field coercion so an unreplaced stub in a typed column (start: int, a closed vocabulary, the genotype grammar) is diagnosed as an unfilled template rather than as "Input should be a valid integer" — the author is told what to do, not what pydantic wanted.

This tightens validation: a module carrying the literal string <<REPLACE>> in a free-text cell would now be invalid. Recorded deliberately rather than slipped in; the token is chosen so no curated prose contains it.

Source code in schema/src/just_dna_format/vocab.py
def reject_template_placeholders(data: object, *, what: str = "row") -> object:
    """A `mode="before"` guard: refuse any cell still carrying `TEMPLATE_PLACEHOLDER`.

    Runs *before* field coercion so an unreplaced stub in a typed column (`start: int`, a closed
    vocabulary, the genotype grammar) is diagnosed as an unfilled template rather than as
    "Input should be a valid integer" — the author is told what to do, not what pydantic wanted.

    This tightens validation: a module carrying the literal string `<<REPLACE>>` in a free-text cell
    would now be invalid. Recorded deliberately rather than slipped in; the token is chosen so no
    curated prose contains it."""
    hits = sorted(_placeholder_paths(data, prefix=""))
    if hits:
        raise ValueError(
            f"unreplaced template placeholder {TEMPLATE_PLACEHOLDER!r} in {what}: "
            f"{', '.join(hits)}. Replace the value, or delete the row if you do not need it."
        )
    return data

reject_misplaced

reject_misplaced(
    data: object, declared: Iterable[str], what: str
) -> object

A mode="before" guard naming a column that is real elsewhere in the DSL but not on this model.

Sits between reject_reserved (a name no model has, held against a future release) and extra="forbid" (an unknown or misspelled name). declared is the model's own field names, so a model that genuinely carries the column — FrequencyRow.source — is never touched, and the check cannot drift out of step with the models the way a second name list would.

Source code in schema/src/just_dna_format/vocab.py
def reject_misplaced(data: object, declared: Iterable[str], what: str) -> object:
    """A `mode="before"` guard naming a column that is real elsewhere in the DSL but not on this model.

    Sits between `reject_reserved` (a name no model has, held against a future release) and
    `extra="forbid"` (an unknown or misspelled name). `declared` is the model's own field names, so a
    model that genuinely carries the column — `FrequencyRow.source` — is never touched, and the check
    cannot drift out of step with the models the way a second name list would."""
    if isinstance(data, dict):
        fields = frozenset(declared)
        hits = sorted(k for k in data if k in MISPLACED_COLUMN_REASONS and k not in fields)
        if hits:
            reasons = "; ".join(f"{h!r} {MISPLACED_COLUMN_REASONS[h]}" for h in hits)
            raise ValueError(f"column(s) not authored on a {what}: {reasons}.")
    return data

reject_reserved

reject_reserved(data: object) -> object

A mode="before" guard for every authored model, layered on top of extra="forbid".

extra="forbid" already rejects any unknown column, but treats a reserved name and a random/typo'd one identically (the generic "extra inputs are not permitted"). This guard runs first and, when the raw input carries a reserved-namespace column (RESERVED_NAMES_0_4), raises a specific error stating what the name is reserved for and that a future release may claim it — so reference_db fails differently from xyzzy. (It said caller until 2026-08-12, which had stopped being true: caller was dropped from the reserved set rather than built, so it takes the generic message like any other stray column, and the example claimed a diagnosis the code no longer produces.) That is the reserved list's build-time (author/compile-time) value: reserved ≠ arbitrary at the point of failure, not merely in a published dictionary. A misspelled or genuinely-unknown column still falls through to extra="forbid"'s generic message (a hint to check the field list). Non-mapping input passes through untouched (pydantic handles it).

Source code in schema/src/just_dna_format/vocab.py
def reject_reserved(data: object) -> object:
    """A `mode="before"` guard for every authored model, layered *on top of* `extra="forbid"`.

    `extra="forbid"` already rejects any unknown column, but treats a reserved name and a random/typo'd
    one identically (the generic "extra inputs are not permitted"). This guard runs first and, when the
    raw input carries a reserved-namespace column (`RESERVED_NAMES_0_4`), raises a *specific* error
    stating what the name is reserved for and that a future release may claim it — so `reference_db`
    fails differently from `xyzzy`. (It said `caller` until 2026-08-12, which had stopped being true:
    `caller` was *dropped* from the reserved set rather than built, so it takes the generic message like
    any other stray column, and the example claimed a diagnosis the code no longer produces.) That is
    the reserved list's build-time (author/compile-time) value:
    reserved ≠ arbitrary at the point of failure, not merely in a published dictionary. A misspelled or
    genuinely-unknown column still falls through to `extra="forbid"`'s generic message (a hint to check
    the field list). Non-mapping input passes through untouched (pydantic handles it)."""
    if isinstance(data, dict):
        hits = sorted(k for k in data if k in RESERVED_NAMES_0_4)
        if hits:
            reasons = "; ".join(
                f"{h!r} {RESERVED_NAME_REASONS.get(h, 'an anticipated future axis')}" for h in hits
            )
            raise ValueError(
                f"reserved column name(s), not authorable fields: {reasons}. Reserved against the "
                f"one-way door (CONSTITUTION P3/P5) — a future release may claim them; do not author "
                f"them into a module. (Reserved now: {sorted(RESERVED_NAMES_0_4)}.)"
            )
    return data

split_field_pointer

split_field_pointer(
    value: str,
) -> list[tuple[str | None, str]]

A pointer cell → its (namespace, key) atoms, in authored order.

| is alternation between fields (try the first, fall back to the next) and never indexing, so a cell yields one atom per alternative. namespace is None for a bare key, which means unqualified — an honest absence, not a default: the whole point of RM53 is that guessing the namespace converts unstated into a stated answer, and it would be wrong for the very first module that used the column.

Source code in schema/src/just_dna_format/vocab.py
def split_field_pointer(value: str) -> list[tuple[str | None, str]]:
    """A pointer cell → its `(namespace, key)` atoms, in authored order.

    `|` is *alternation between fields* (try the first, fall back to the next) and never indexing, so
    a cell yields one atom per alternative. `namespace` is `None` for a bare key, which means
    **unqualified** — an honest absence, not a default: the whole point of RM53 is that guessing the
    namespace converts *unstated* into a *stated* answer, and it would be wrong for the very first
    module that used the column."""
    atoms: list[tuple[str | None, str]] = []
    for part in value.split("|"):
        head, sep, tail = part.partition("/")
        if sep and head in VCF_NAMESPACES:
            atoms.append((head, tail))
        else:
            atoms.append((None, part))
    return atoms

vcf_field_number

vcf_field_number(
    namespace: str | None, key: str
) -> str | None

The Number the VCF spec reserves for this field, or None when it is not knowable here.

Three-valued, and the third value is the common one. A qualified pointer is looked up directly.

A bare key splits on whether it is a known collision. For a key only one namespace reserves, the bare form is not ambiguous at all — no INFO table defines GQ — so the single entry is the answer. For a colliding key both namespaces have to be in the table and agree: AD is R either way, so its cardinality is settled even though its meaning is not, while CN is A under INFO and 1 under FORMAT and has no answer until the pointer says which. AF is the case that makes the distinction load-bearing — the spec reserves INFO/AF and does not reserve FORMAT/AF (which every caller nevertheless emits), so reading the one known entry as the bare key's cardinality would answer a question about a field the spec never described.

Anything the reserved tables do not carry (a caller's own key, REPCN) is unknown, and unknown withholds: asserting a cardinality for a field this tier has never seen described would be a source convention wearing a fact.

Source code in schema/src/just_dna_format/vocab.py
def vcf_field_number(namespace: str | None, key: str) -> str | None:
    """The `Number` the VCF spec reserves for this field, or `None` when it is not knowable here.

    Three-valued, and the third value is the common one. A qualified pointer is looked up directly.

    A **bare** key splits on whether it is a known collision. For a key only one namespace reserves,
    the bare form is not ambiguous at all — no INFO table defines `GQ` — so the single entry is the
    answer. For a colliding key both namespaces have to be in the table *and* agree: `AD` is `R`
    either way, so its cardinality is settled even though its meaning is not, while `CN` is `A` under
    INFO and `1` under FORMAT and has no answer until the pointer says which. `AF` is the case that
    makes the distinction load-bearing — the spec reserves `INFO/AF` and does **not** reserve
    `FORMAT/AF` (which every caller nevertheless emits), so reading the one known entry as the bare
    key's cardinality would answer a question about a field the spec never described.

    Anything the reserved tables do not carry (a caller's own key, `REPCN`) is unknown, and unknown
    withholds: asserting a cardinality for a field this tier has never seen described would be a
    source convention wearing a fact."""
    if namespace is not None:
        return VCF_FIELD_NUMBER.get(f"{namespace}/{key}")
    known = [
        VCF_FIELD_NUMBER[f"{ns}/{key}"] for ns in sorted(VCF_NAMESPACES) if f"{ns}/{key}" in VCF_FIELD_NUMBER
    ]
    if key in VCF_COLLIDING_KEYS and len(known) < len(VCF_NAMESPACES):
        return None
    return known[0] if len(set(known)) == 1 else None

is_multi_valued_number

is_multi_valued_number(number: str | None) -> bool

Whether a Number code describes a value list rather than a single value.

0 is a Flag (present or absent) and 1 is a scalar; everything else — A, R, G, P, . and the fixed counts 2/4 — returns more than one value, so a pointer at it names a list and not a number.

The return is a bare bool on purpose, and the tri-state is preserved by the caller's action rather than by this signature. The sentence here used to read "None (unknown) is not multi-valued: withhold, never negate, and never accuse", which promised three outcomes from a total function that has two — a reader taking it at face value would look for a None this can never return (RM225). What actually holds: the sole caller (compiler._vcf_pointer_warnings) uses this as if is_multi_valued_number(number): to decide whether to raise a warning about an unselected element, so False there means withhold the warning — and withholding is exactly the right move on an unknown cardinality. Answering False for an unknown is the withhold, not a negation of it. The input is already three-valued (vcf_field_number returns None for unknown) and the two states this collapses call for the same action, so the narrowing happens at the point where it is safe.

If a future caller ever needs to tell unknown from scalar, it must read vcf_field_number's answer directly rather than widening this to bool | None — the collapse is the contract here, not an accident.

Source code in schema/src/just_dna_format/vocab.py
def is_multi_valued_number(number: str | None) -> bool:
    """Whether a `Number` code describes a value list rather than a single value.

    `0` is a Flag (present or absent) and `1` is a scalar; everything else — `A`, `R`, `G`, `P`, `.`
    and the fixed counts `2`/`4` — returns more than one value, so a pointer at it names a list and
    not a number.

    **The return is a bare `bool` on purpose, and the tri-state is preserved by the caller's action
    rather than by this signature.** The sentence here used to read "`None` (unknown) is **not**
    multi-valued: withhold, never negate, and never accuse", which promised three outcomes from a
    total function that has two — a reader taking it at face value would look for a `None` this can
    never return (RM225). What actually holds: the sole caller
    (`compiler._vcf_pointer_warnings`) uses this as `if is_multi_valued_number(number):` to decide
    whether to *raise* a warning about an unselected element, so `False` there means **withhold the
    warning** — and withholding is exactly the right move on an unknown cardinality. Answering
    `False` for an unknown is the withhold, not a negation of it. The input is already three-valued
    (`vcf_field_number` returns `None` for unknown) and the two states this collapses call for the
    same action, so the narrowing happens at the point where it is safe.

    If a future caller ever needs to tell *unknown* from *scalar*, it must read `vcf_field_number`'s
    answer directly rather than widening this to `bool | None` — the collapse is the contract here,
    not an accident."""
    return number is not None and number not in {"0", "1"}

validate_field_token

validate_field_token(
    value: str | None, field_name: str
) -> str | None

Validate a VCF field-name pointer: a key, optionally namespace-qualified, optionally |-alternated (FORMAT/REPCN, INFO/DP|FORMAT/DP, CN|DS).

The grammar is what keeps such a column a pointer and not an expression — no operators, no whitespace, no code — which is what lets Principle 1 (declarative, data-not-code) hold while a module still says where in a VCF its quantity, its callability signal or its confidence floor lives.

Two widenings landed in 0.6, both strictly additive (P3 — every cell that validated before still validates and still means the same thing):

  • The namespace qualifier (RM53). A VCF field is identified by namespace and name, and the two reserved-key tables collide on DP, AD, ADF, ADR, MQ, AF and — new in 4.4 — CN. A bare key stays legal and keeps meaning unqualified; the compiler warns when an unqualified one is a known collision, which is what stops the bare spelling from looking like a decision.
  • The spec's own key charset (RM61). A dot is legal inside a key and 1000G is a legal key, both reserved by name; the previous grammar refused them while claiming to describe VCF field names.

The namespace is matched case-sensitively, unlike the -/_ slip match_vocab absorbs: INFO/ and FORMAT/ are how every VCF header, every spec table and bcftools spell it, so there is no established lowercase spelling for an author to slip into.

Shared by binning.MeasureBinRow.source_field (the measured quantity), spec.VariantRow.callable_from (the callability signal) and spec.VariantRow.quality_from (the field the min_quality floor is stated against), so it lives on AuthoredModel rather than being copied per model.

Source code in schema/src/just_dna_format/vocab.py
def validate_field_token(value: str | None, field_name: str) -> str | None:
    """Validate a **VCF field-name pointer**: a key, optionally namespace-qualified, optionally
    `|`-alternated (`FORMAT/REPCN`, `INFO/DP|FORMAT/DP`, `CN|DS`).

    The grammar is what keeps such a column a *pointer* and not an expression — no operators, no
    whitespace, no code — which is what lets Principle 1 (declarative, data-not-code) hold while a
    module still says where in a VCF its quantity, its callability signal or its confidence floor
    lives.

    Two widenings landed in 0.6, both strictly additive (P3 — every cell that validated before still
    validates and still means the same thing):

    * **The namespace qualifier** (RM53). A VCF field is identified by namespace *and* name, and the
      two reserved-key tables collide on `DP`, `AD`, `ADF`, `ADR`, `MQ`, `AF` and — new in 4.4 — `CN`.
      A bare key stays legal and keeps meaning *unqualified*; the compiler warns when an unqualified
      one is a known collision, which is what stops the bare spelling from looking like a decision.
    * **The spec's own key charset** (RM61). A dot is legal inside a key and `1000G` is a legal key,
      both reserved by name; the previous grammar refused them while claiming to describe VCF field
      names.

    The namespace is matched case-sensitively, unlike the `-`/`_` slip `match_vocab` absorbs: `INFO/`
    and `FORMAT/` are how every VCF header, every spec table and `bcftools` spell it, so there is no
    established lowercase spelling for an author to slip into.

    Shared by `binning.MeasureBinRow.source_field` (the measured quantity),
    `spec.VariantRow.callable_from` (the callability signal) and `spec.VariantRow.quality_from` (the
    field the `min_quality` floor is stated against), so it lives on `AuthoredModel` rather than being
    copied per model."""
    if value is not None and not SOURCE_FIELD_PATTERN.match(value):
        raise ValueError(
            f"{field_name} must be a VCF field-name pointer — a key, optionally qualified by its "
            f"namespace and optionally |-alternated (e.g. REPCN, FORMAT/REPCN, CN|DS, "
            f"INFO/DP|FORMAT/DP) — a pointer, not an expression, got: {value!r}. A namespace prefix "
            f"is one of {sorted(VCF_NAMESPACES)} followed by '/'."
        )
    return value

match_vocab

match_vocab(
    value: str, vocab: frozenset[str]
) -> str | None

The vocabulary member value names, treating - and _ as the same separator.

A hyphen where an underscore goes is the most common slip in a hand-written cell, and this format's whole premise is that the DSL is authorable by a human. --use non-commercial was already accepted by the enricher CLI, which normalized the separator on its way in, while the identical string in a licensing.csv cell was refused — so the surface an author learns the vocabulary from taught a spelling the file rejected.

Both directions are tried rather than one, and the exact value first: no vocabulary in this schema carries a hyphenated member today, but trying the value as written before swapping means a future one cannot be broken by this function. A swap can never be ambiguous either — it would take two members differing only in their separators, which would be two spellings of one thing.

Returns the canonical member (so the stored cell is always the declared spelling), or None when the value names nothing. Widening what a field accepts, never narrowing: every value that validated before still validates, which is what keeps this Principle 3-legal.

Source code in schema/src/just_dna_format/vocab.py
def match_vocab(value: str, vocab: frozenset[str]) -> str | None:
    """The vocabulary member `value` names, treating `-` and `_` as the same separator.

    **A hyphen where an underscore goes is the most common slip in a hand-written cell**, and this
    format's whole premise is that the DSL is authorable by a human. `--use non-commercial` was
    already accepted by the enricher CLI, which normalized the separator on its way in, while the
    identical string in a `licensing.csv` cell was refused — so the surface an author learns the
    vocabulary from taught a spelling the file rejected.

    Both directions are tried rather than one, and the exact value first: no vocabulary in this
    schema carries a hyphenated member today, but trying the value as written before swapping means a
    future one cannot be broken by this function. A swap can never be ambiguous either — it would take
    two members differing *only* in their separators, which would be two spellings of one thing.

    Returns the canonical member (so the stored cell is always the declared spelling), or `None` when
    the value names nothing. Widening what a field accepts, never narrowing: every value that
    validated before still validates, which is what keeps this Principle 3-legal.
    """
    if value in vocab:
        return value
    for candidate in (value.replace("-", "_"), value.replace("_", "-")):
        if candidate in vocab:
            return candidate
    return None

check_vocab

check_vocab(
    value: str | None,
    vocab: frozenset[str],
    field_name: str,
) -> str | None

Validate an optional categorical against a closed frozenset vocabulary (Principle 6).

Passes None through (absent = unknown), and canonicalizes a -/_ separator slip to the declared member (see match_vocab). The message format matches the pre-refactor per-field validators exactly (<field> must be one of [...], got: <value>).

Source code in schema/src/just_dna_format/vocab.py
def check_vocab(value: str | None, vocab: frozenset[str], field_name: str) -> str | None:
    """Validate an optional categorical against a closed `frozenset` vocabulary (Principle 6).

    Passes `None` through (absent = unknown), and canonicalizes a `-`/`_` separator slip to the
    declared member (see `match_vocab`). The message format matches the pre-refactor per-field
    validators exactly (`<field> must be one of [...], got: <value>`)."""
    if value is None:
        return None
    matched = match_vocab(value, vocab)
    if matched is None:
        raise ValueError(f"{field_name} must be one of {sorted(vocab)}, got: {value!r}")
    return matched

validate_trait_ids

validate_trait_ids(
    value: str | None, field_name: str = "trait_efo_id"
) -> str | None

Validate a multi-valued CURIE cell: each [,;|]-split token must be an ontology CURIE.

Source code in schema/src/just_dna_format/vocab.py
def validate_trait_ids(value: str | None, field_name: str = "trait_efo_id") -> str | None:
    """Validate a multi-valued CURIE cell: each `[,;|]`-split token must be an ontology CURIE."""
    if value is None:
        return value
    for tok in MULTI_SEP.split(value):
        tok = tok.strip()
        if tok and not TRAIT_ID_PATTERN.match(tok):
            raise ValueError(
                f"{field_name} tokens must be ontology CURIEs like EFO_0004340 / MONDO:0005265, got: {tok!r}"
            )
    return value

validate_allele

validate_allele(
    value: str | None, field_name: str = "allele"
) -> str | None

Validate an optional allele: a nucleotide string (^[ACGT]+$, case-insensitive), or a symbolic/structural allele carrying its length (<DEL:1500>, <CNV:TR:30>) since 0.6 (RM5).

Two users, not one — HaplotypeRow.allele and VariantRow.effect_allele. (alleles.py and CLAUDE.md both said "exactly one" until RM5; the count is what an author of a grammar change reads to size the blast radius.)

A lengthless symbolic allele passes here and is refused later, by the compiler. That split is forced, not chosen: rejecting it at load makes the row fail to parse, which is fatal in both modes, and the decided behaviour is a warning-and-drop under best_effort. So the schema says what the DSL can spell and the compiler says what makes a usable rulebook.

Source code in schema/src/just_dna_format/vocab.py
def validate_allele(value: str | None, field_name: str = "allele") -> str | None:
    """Validate an optional allele: a nucleotide string (`^[ACGT]+$`, case-insensitive), or a
    symbolic/structural allele carrying its length (`<DEL:1500>`, `<CNV:TR:30>`) since 0.6 (RM5).

    **Two users, not one** — `HaplotypeRow.allele` and `VariantRow.effect_allele`. (`alleles.py` and
    CLAUDE.md both said "exactly one" until RM5; the count is what an author of a grammar change reads
    to size the blast radius.)

    A *lengthless* symbolic allele passes here and is refused later, by the compiler. That split is
    forced, not chosen: rejecting it at load makes the row fail to parse, which is fatal in **both**
    modes, and the decided behaviour is a warning-and-drop under `best_effort`. So the schema says
    what the DSL can spell and the compiler says what makes a usable rulebook.
    """
    if value is None or ALLELE_PATTERN.match(value):
        return value
    if parse_symbolic_allele(value) is not None:
        return value
    # The message names the length convention without claiming *this* validator enforces it — it does
    # not, deliberately (see the docstring), and the compiler is what refuses a lengthless one. An
    # author reading this is being rejected on the *type*, so telling them the full spelling here is
    # what stops the second rejection one command later.
    raise ValueError(
        f"{field_name} must be nucleotides (e.g. A, G, AC) or a symbolic/structural allele whose "
        f"first-level type is one of {sorted(SYMBOLIC_ALLELE_TYPES)} — the length belongs inside the "
        f"token (<DEL:1500>, <CNV:TR:30>), and a compile refuses one that states none. "
        f"Got: {value!r}"
    )

validate_rsid

validate_rsid(value: str | None) -> str | None

Validate an optional dbSNP identifier (rs<digits>).

Source code in schema/src/just_dna_format/vocab.py
def validate_rsid(value: str | None) -> str | None:
    """Validate an optional dbSNP identifier (`rs<digits>`)."""
    if value is not None and not RSID_PATTERN.match(value):
        raise ValueError(f"rsid must match rs<digits>, got: {value!r}")
    return value

population_sort_key

population_sort_key(population: str) -> tuple[int, str]

Deterministic sort key for an ancestry group: seeded groups in POPULATION_ORDER, then the rest alphabetically. Total and stable for any label, which is what the open vocabulary needs.

Source code in schema/src/just_dna_format/vocab.py
def population_sort_key(population: str) -> tuple[int, str]:
    """Deterministic sort key for an ancestry group: seeded groups in `POPULATION_ORDER`, then the
    rest alphabetically. Total and stable for any label, which is what the open vocabulary needs."""
    try:
        return (POPULATION_ORDER.index(population), "")
    except ValueError:
        return (len(POPULATION_ORDER), population)

normalize_population

normalize_population(value: str) -> str

Fold an ancestry-group label to its canonical form: stripped, lowercased, and a bare/empty label mapped to global (which is how gnomAD reports the whole-dataset row).

Source code in schema/src/just_dna_format/vocab.py
def normalize_population(value: str) -> str:
    """Fold an ancestry-group label to its canonical form: stripped, lowercased, and a bare/empty
    label mapped to `global` (which is how gnomAD reports the whole-dataset row)."""
    token = (value or "").strip().lower()
    return token or "global"

validate_population

validate_population(value: str) -> str

Validate an ancestry group against the OPEN seeded vocabulary.

Enforces only that the label is a non-empty, well-formed token — membership in RECOMMENDED_ANCESTRY_GROUPS is a recommendation, not a gate (see the comment beside it). A label that is merely unfamiliar is kept, because the next source will bring its own naming; one that is malformed (empty, padded, or carrying a separator that would break a CSV cell or a group-by) is rejected, because that is a data error rather than a new source's naming.

Source code in schema/src/just_dna_format/vocab.py
def validate_population(value: str) -> str:
    """Validate an ancestry group against the OPEN seeded vocabulary.

    Enforces only that the label is a non-empty, well-formed token — membership in
    `RECOMMENDED_ANCESTRY_GROUPS` is a recommendation, not a gate (see the comment beside it). A
    label that is merely unfamiliar is kept, because the next source will bring its own naming; one
    that is *malformed* (empty, padded, or carrying a separator that would break a CSV cell or a
    group-by) is rejected, because that is a data error rather than a new source's naming.
    """
    if not value or value != value.strip():
        raise ValueError(f"population must be a non-empty label without padding, got: {value!r}")
    if not POPULATION_PATTERN.match(value):
        raise ValueError(
            f"population must be a lowercase token (letters, digits, underscore) — recommended: "
            f"{sorted(RECOMMENDED_ANCESTRY_GROUPS)}, got: {value!r}"
        )
    return value

validate_finite

validate_finite(
    value: float | None, field_name: str
) -> float | None

Reject a non-finite float (NaN/inf). A NaN breaks round-trip equality (NaN != NaN makes needs_upgrade/idempotency checks oscillate) and serialises to the non-reloadable cell "nan"; an authored measure is always a finite number. Passes None through.

Source code in schema/src/just_dna_format/vocab.py
def validate_finite(value: float | None, field_name: str) -> float | None:
    """Reject a non-finite float (`NaN`/`inf`). A `NaN` breaks round-trip equality (`NaN != NaN`
    makes `needs_upgrade`/idempotency checks oscillate) and serialises to the non-reloadable cell
    `"nan"`; an authored measure is always a finite number. Passes `None` through."""
    if value is not None and not math.isfinite(value):
        raise ValueError(f"{field_name} must be a finite number, got: {value!r}")
    return value