just_dna_format.vocab¶
just_dna_format.vocab ¶
Shared constrained vocabularies, identifier patterns, and reusable validator helpers.
A dependency-light leaf (stdlib only) so every authored-DSL model — spec (variants/studies),
binning (the measure→phenotype primitive), pgx (star-alleles), and pgs — validates against
one source of truth for the orthogonal axes and identifier grammars. Per CONSTITUTION Principle 6,
constrained vocabularies are frozenset[str] + a validator, never Enum/Literal.
spec re-exports the names it historically owned, so existing imports
(from just_dna_format.spec import VALID_DIRECTIONS) keep working unchanged.
validate_phenotype_categories ¶
validate_phenotype_categories(
value: str | None,
field_name: str = "phenotype_category",
) -> str | None
Validate a multi-valued phenotype-category cell against VALID_PHENOTYPE_CATEGORIES.
Accepts ClinPGx's own spellings (Metabolism/PK) case-insensitively and normalizes them to the
vocabulary member (metabolism_pk), because the authored DSL should not make a human transcribe
a slash-and-caps token exactly.
Source code in schema/src/just_dna_format/vocab.py
reject_template_placeholders ¶
A mode="before" guard: refuse any cell still carrying TEMPLATE_PLACEHOLDER.
Runs before field coercion so an unreplaced stub in a typed column (start: int, a closed
vocabulary, the genotype grammar) is diagnosed as an unfilled template rather than as
"Input should be a valid integer" — the author is told what to do, not what pydantic wanted.
This tightens validation: a module carrying the literal string <<REPLACE>> in a free-text cell
would now be invalid. Recorded deliberately rather than slipped in; the token is chosen so no
curated prose contains it.
Source code in schema/src/just_dna_format/vocab.py
reject_misplaced ¶
A mode="before" guard naming a column that is real elsewhere in the DSL but not on this model.
Sits between reject_reserved (a name no model has, held against a future release) and
extra="forbid" (an unknown or misspelled name). declared is the model's own field names, so a
model that genuinely carries the column — FrequencyRow.source — is never touched, and the check
cannot drift out of step with the models the way a second name list would.
Source code in schema/src/just_dna_format/vocab.py
reject_reserved ¶
A mode="before" guard for every authored model, layered on top of extra="forbid".
extra="forbid" already rejects any unknown column, but treats a reserved name and a random/typo'd
one identically (the generic "extra inputs are not permitted"). This guard runs first and, when the
raw input carries a reserved-namespace column (RESERVED_NAMES_0_4), raises a specific error
stating what the name is reserved for and that a future release may claim it — so reference_db
fails differently from xyzzy. (It said caller until 2026-08-12, which had stopped being true:
caller was dropped from the reserved set rather than built, so it takes the generic message like
any other stray column, and the example claimed a diagnosis the code no longer produces.) That is
the reserved list's build-time (author/compile-time) value:
reserved ≠ arbitrary at the point of failure, not merely in a published dictionary. A misspelled or
genuinely-unknown column still falls through to extra="forbid"'s generic message (a hint to check
the field list). Non-mapping input passes through untouched (pydantic handles it).
Source code in schema/src/just_dna_format/vocab.py
split_field_pointer ¶
A pointer cell → its (namespace, key) atoms, in authored order.
| is alternation between fields (try the first, fall back to the next) and never indexing, so
a cell yields one atom per alternative. namespace is None for a bare key, which means
unqualified — an honest absence, not a default: the whole point of RM53 is that guessing the
namespace converts unstated into a stated answer, and it would be wrong for the very first
module that used the column.
Source code in schema/src/just_dna_format/vocab.py
vcf_field_number ¶
The Number the VCF spec reserves for this field, or None when it is not knowable here.
Three-valued, and the third value is the common one. A qualified pointer is looked up directly.
A bare key splits on whether it is a known collision. For a key only one namespace reserves,
the bare form is not ambiguous at all — no INFO table defines GQ — so the single entry is the
answer. For a colliding key both namespaces have to be in the table and agree: AD is R
either way, so its cardinality is settled even though its meaning is not, while CN is A under
INFO and 1 under FORMAT and has no answer until the pointer says which. AF is the case that
makes the distinction load-bearing — the spec reserves INFO/AF and does not reserve
FORMAT/AF (which every caller nevertheless emits), so reading the one known entry as the bare
key's cardinality would answer a question about a field the spec never described.
Anything the reserved tables do not carry (a caller's own key, REPCN) is unknown, and unknown
withholds: asserting a cardinality for a field this tier has never seen described would be a
source convention wearing a fact.
Source code in schema/src/just_dna_format/vocab.py
is_multi_valued_number ¶
Whether a Number code describes a value list rather than a single value.
0 is a Flag (present or absent) and 1 is a scalar; everything else — A, R, G, P, .
and the fixed counts 2/4 — returns more than one value, so a pointer at it names a list and
not a number.
The return is a bare bool on purpose, and the tri-state is preserved by the caller's action
rather than by this signature. The sentence here used to read "None (unknown) is not
multi-valued: withhold, never negate, and never accuse", which promised three outcomes from a
total function that has two — a reader taking it at face value would look for a None this can
never return (RM225). What actually holds: the sole caller
(compiler._vcf_pointer_warnings) uses this as if is_multi_valued_number(number): to decide
whether to raise a warning about an unselected element, so False there means withhold the
warning — and withholding is exactly the right move on an unknown cardinality. Answering
False for an unknown is the withhold, not a negation of it. The input is already three-valued
(vcf_field_number returns None for unknown) and the two states this collapses call for the
same action, so the narrowing happens at the point where it is safe.
If a future caller ever needs to tell unknown from scalar, it must read vcf_field_number's
answer directly rather than widening this to bool | None — the collapse is the contract here,
not an accident.
Source code in schema/src/just_dna_format/vocab.py
validate_field_token ¶
Validate a VCF field-name pointer: a key, optionally namespace-qualified, optionally
|-alternated (FORMAT/REPCN, INFO/DP|FORMAT/DP, CN|DS).
The grammar is what keeps such a column a pointer and not an expression — no operators, no whitespace, no code — which is what lets Principle 1 (declarative, data-not-code) hold while a module still says where in a VCF its quantity, its callability signal or its confidence floor lives.
Two widenings landed in 0.6, both strictly additive (P3 — every cell that validated before still validates and still means the same thing):
- The namespace qualifier (RM53). A VCF field is identified by namespace and name, and the
two reserved-key tables collide on
DP,AD,ADF,ADR,MQ,AFand — new in 4.4 —CN. A bare key stays legal and keeps meaning unqualified; the compiler warns when an unqualified one is a known collision, which is what stops the bare spelling from looking like a decision. - The spec's own key charset (RM61). A dot is legal inside a key and
1000Gis a legal key, both reserved by name; the previous grammar refused them while claiming to describe VCF field names.
The namespace is matched case-sensitively, unlike the -/_ slip match_vocab absorbs: INFO/
and FORMAT/ are how every VCF header, every spec table and bcftools spell it, so there is no
established lowercase spelling for an author to slip into.
Shared by binning.MeasureBinRow.source_field (the measured quantity),
spec.VariantRow.callable_from (the callability signal) and spec.VariantRow.quality_from (the
field the min_quality floor is stated against), so it lives on AuthoredModel rather than being
copied per model.
Source code in schema/src/just_dna_format/vocab.py
match_vocab ¶
The vocabulary member value names, treating - and _ as the same separator.
A hyphen where an underscore goes is the most common slip in a hand-written cell, and this
format's whole premise is that the DSL is authorable by a human. --use non-commercial was
already accepted by the enricher CLI, which normalized the separator on its way in, while the
identical string in a licensing.csv cell was refused — so the surface an author learns the
vocabulary from taught a spelling the file rejected.
Both directions are tried rather than one, and the exact value first: no vocabulary in this schema carries a hyphenated member today, but trying the value as written before swapping means a future one cannot be broken by this function. A swap can never be ambiguous either — it would take two members differing only in their separators, which would be two spellings of one thing.
Returns the canonical member (so the stored cell is always the declared spelling), or None when
the value names nothing. Widening what a field accepts, never narrowing: every value that
validated before still validates, which is what keeps this Principle 3-legal.
Source code in schema/src/just_dna_format/vocab.py
check_vocab ¶
Validate an optional categorical against a closed frozenset vocabulary (Principle 6).
Passes None through (absent = unknown), and canonicalizes a -/_ separator slip to the
declared member (see match_vocab). The message format matches the pre-refactor per-field
validators exactly (<field> must be one of [...], got: <value>).
Source code in schema/src/just_dna_format/vocab.py
validate_trait_ids ¶
Validate a multi-valued CURIE cell: each [,;|]-split token must be an ontology CURIE.
Source code in schema/src/just_dna_format/vocab.py
validate_allele ¶
Validate an optional allele: a nucleotide string (^[ACGT]+$, case-insensitive), or a
symbolic/structural allele carrying its length (<DEL:1500>, <CNV:TR:30>) since 0.6 (RM5).
Two users, not one — HaplotypeRow.allele and VariantRow.effect_allele. (alleles.py and
CLAUDE.md both said "exactly one" until RM5; the count is what an author of a grammar change reads
to size the blast radius.)
A lengthless symbolic allele passes here and is refused later, by the compiler. That split is
forced, not chosen: rejecting it at load makes the row fail to parse, which is fatal in both
modes, and the decided behaviour is a warning-and-drop under best_effort. So the schema says
what the DSL can spell and the compiler says what makes a usable rulebook.
Source code in schema/src/just_dna_format/vocab.py
validate_rsid ¶
Validate an optional dbSNP identifier (rs<digits>).
Source code in schema/src/just_dna_format/vocab.py
population_sort_key ¶
Deterministic sort key for an ancestry group: seeded groups in POPULATION_ORDER, then the
rest alphabetically. Total and stable for any label, which is what the open vocabulary needs.
Source code in schema/src/just_dna_format/vocab.py
normalize_population ¶
Fold an ancestry-group label to its canonical form: stripped, lowercased, and a bare/empty
label mapped to global (which is how gnomAD reports the whole-dataset row).
Source code in schema/src/just_dna_format/vocab.py
validate_population ¶
Validate an ancestry group against the OPEN seeded vocabulary.
Enforces only that the label is a non-empty, well-formed token — membership in
RECOMMENDED_ANCESTRY_GROUPS is a recommendation, not a gate (see the comment beside it). A
label that is merely unfamiliar is kept, because the next source will bring its own naming; one
that is malformed (empty, padded, or carrying a separator that would break a CSV cell or a
group-by) is rejected, because that is a data error rather than a new source's naming.
Source code in schema/src/just_dna_format/vocab.py
validate_finite ¶
Reject a non-finite float (NaN/inf). A NaN breaks round-trip equality (NaN != NaN
makes needs_upgrade/idempotency checks oscillate) and serialises to the non-reloadable cell
"nan"; an authored measure is always a finite number. Passes None through.