just_dna_format.base¶
just_dna_format.base ¶
Shared base for every authored-DSL row model (spec/binning/pgx/pgs).
Consolidates the boilerplate that was copy-pasted across the row models into one place:
- the reserved-namespace guard —
extra="forbid"plus thereject_reservedbefore-validator, so a reserved name fails with a specific diagnosis and any other unknown/misspelled column fails with the generic message (seevocab.reject_reserved); and - the field validators for the shared authored vocabulary —
rsid,trait_efo_id,direction,clin_sig,stat_significance,evidence_level, finite-effect_size,genotype, the VCF field-pointer grammar (source_field/callable_from/quality_from) and the element rule that qualifies each of those (source_element/callable_element/quality_element).
Each field validator uses check_fields=False, so a subclass runs it only for the fields it actually
declares (a model without clin_sig simply never runs the clin_sig check) and a model that adds
one of these fields gets the correct validation for free — the per-field rules cannot drift model to
model, which is exactly what the previous copy-paste risked. Field-specific rules (star-allele
strings, measure bounds, PGS ancestry, the mtDNA legacy-reference guard, identifier completeness)
stay on their own models.
genotype moved here in 0.5 when PharmVariantRow gained one: the grammar is the same grammar
(a PharmGKB per-genotype clinical annotation describes the same diploid call a VariantRow does), and
the DRY rule applies as soon as a validator is shared by two models. It is deliberately not
widened for the symbolic alleles PharmGKB also carries (C/del, del/del) — those are RM5, and the
enricher skips them rather than coercing them into a nucleotide grammar that cannot express them.
Dependency-light: imports only pydantic + the stdlib vocab leaf, and nothing in the package
imports it back, so it introduces no cycle.
AuthoredModel ¶
Bases: BaseModel
Base for authored-DSL rows: reserved-namespace guard + shared-vocabulary field validators.
genome_build
property
¶
The assembly this row was loaded as being on. Read-only; see _genome_build.
with_genome_build ¶
Tell this row which assembly it is on. Returns self, so a loader can map over rows.
Deliberately a method rather than a settable property: injecting the build is something a
loader does once, at a known point, and making it look like an ordinary attribute assignment
would invite it being done anywhere. just_dna_compiler.compiler._load_csv_rows is the caller.
A model that stamps an identity (_KEY_INCLUDES_ALTS set) re-derives its variant_key here,
which is this loader's answer to the problem _restamp_for_build solves for VariantRow one
tier up: a row is constructed from a CSV dict where module_spec.yaml is not in scope, so the
key it froze at construction took derive_variant_key's GRCh38 default. Doing it at the point
the build arrives means a positional table needs no compiler-side correction pass at all.
VariantRow deliberately does not join in — it leaves _KEY_INCLUDES_ALTS unset and keeps its
existing restamp, which also emits the "keyed by coordinate instead" warning a GRCh37 module
must hear.
Source code in schema/src/just_dna_format/base.py
case_insensitive_allele_fields ¶
The columns of model that content_signature case-folds — those marked CASE_INSENSITIVE_ALLELE.
Walked off the fields for the same reason content_identity_exclusions is: a hand-kept set one
file over is the thing that drifts (@fieldnames-from-model).
Source code in schema/src/just_dna_format/base.py
content_identity_exclusions ¶
The columns of model that content_signature omits — those marked OUTSIDE_CONTENT_IDENTITY.
Walked off the fields, never listed beside them, for the same reason authored_field_names is:
a hand-kept set one file over is the thing that drifts. Empty for every model but the overlay
row today, and a test pins that equality over the whole registry.
Source code in schema/src/just_dna_format/base.py
authored_field_names ¶
The columns a human actually authors for model, in field-declaration order.
model_fields minus anything marked COMPILER_MANAGED. Every generator over the authored
surface — a blank template, a drafted CSV's header — reads its columns through here, so a field
that stops being authored stops being offered in the same commit that marks it.
Source code in schema/src/just_dna_format/base.py
merge_key ¶
The key that decides two rows of a machine-produced table are the same row.
The enricher pass that writes each sidecar merges rather than clobbers, so it holds an existing
dict and needs this tuple; the pass reads it from here rather than restating it, which is the
half that makes the pass and hints.key_fields unable to disagree (S51). Restated, they were two
statements of one fact — and the fact was legible nowhere outside the pass's own body, so every
consumer re-deriving a sidecar was guessing it.
_KEY_FALLBACK_FIELDS, where a kind declares one, is used when every primary member is null:
GeneValidityRow keys on the source's own assertion_id and falls back to the gene's grain when
the source published none. The two levels are tagged ("id" / "grain") so a grain tuple can
never collide with an id that happens to equal it.
Raises AttributeError for a model declaring no key, which is the honest failure: a caller
reaching here for an unkeyed kind has a bug, and a silent () would merge every row into one.
An EMPTY _KEY_FIELDS is the same bug and used to slip through (RM213). MeasureBinRow
declares () as a base-class default meaning subclasses set this, and every subclass does — but
a bare MeasureBinRow, or a future kind that inherits the default and forgets, reached here and
got () back. Two rows differing in every column returned equal keys, which is exactly the
collapse the paragraph above says this function raises to prevent. Measured, not argued.
hints.table_key treats the same falsy value as no declared key and returns None; that is the
right answer to a different question (does this table publish a key?) and is deliberately not
changed. Here the question is what two rows' identity IS, and there is no such thing as an empty
answer to it.
Source code in schema/src/just_dna_format/base.py
accepts_none ¶
Does this annotation admit None? (Optional[str] yes; a defaulted bare str/bool no.)
field_category ¶
required | defaulted | optional — the three-way split an authoring surface must respect.
The middle category is the one that bites. MeasureBinRow.measure_kind (str, default
"repeat_count") and unresolved (bool, default False) are not required, so pydantic's
is_required() says False — but _load_csv_rows turns an empty cell into None and keeps
the key, so the model receives None instead of its default and fails on type. An author who
filled exactly the columns a two-way required flag named got a rejection about a column nobody
had mentioned. A defaulted cell has to be written out with its default rather than left blank.
It lives here rather than in the compiler because two surfaces answer this question and they
drifted: just_dna_compiler.draft was fixed to the three-way split and reference.authoring_reference
— the drift-proof description consumers render instead of a hand-kept spec dump — was still
emitting the two-way one. Both now read this. The format tier is the only place both can import
from, and this needs nothing but pydantic.
Source code in schema/src/just_dna_format/base.py
vocabulary ¶
vocabulary(
name: str,
options: frozenset[str],
*,
closed: bool = True,
notes: dict[str, str] | None = None,
) -> dict[str, object]
Mark a field as drawn from options, for tools that offer an author the valid values.
closed=True means a validator rejects anything outside the set; closed=False marks the
recommended-but-open sets (RECOMMENDED_EFFECT_MEASURES, RECOMMENDED_AUTHOR_KINDS), where the
members are suggestions and a novel value is legal. VALID_ACTIONABILITY used to be listed here
and is closed — the drift this very flag exists to prevent, written into the docstring that
defines it (RM225). A
consumer must be able to tell "pick one of these" from "these are suggestions", so the two are one
marker with a flag rather than two markers.
notes is per-member prose, for the vocabularies where the member name cannot carry the whole
rule — ELEMENT_RULE_MEANINGS is the case that forced it, since largest has two answers on a
Number=R field and only a sentence can say which one it means. It rides on the marker for the
same reason options does: a name to look up would need a central registry, and vocab cannot
import pgx (the cycle this module's dependency note exists to avoid), so a registry anywhere
else is a second hand-kept list. Carrying the prose here is what lets every surface that reads
field_vocabularies print it — before this, the meanings reached the whole-schema reference and
nothing else, so describe <kind>, the command an author filling one table runs, listed eight
members and could not say which of them counted the reference element.
Omitted from the marker when there is none, so the shape of the 20-odd vocabularies that need no
prose is unchanged. Members keep their declaration order rather than being sorted: it is the
reading order the constant was written in, and it is as deterministic as sorted(options) is.
Source code in schema/src/just_dna_format/base.py
since ¶
Mark the release a field first appeared in — Field(json_schema_extra=since("0.6.5")) (RM146).
The finding this answers. A module authored on 0.6.6 was sent to a deployment running 0.6.1,
which runs validate_spec server-side and reported, verbatim:
studies.csv line 2 [curator]: Extra inputs are not permitted. StudyRow.curator is ours, added
in 0.6.5. A genuine typo produces the byte-identical shape — [curatr] — and the two want
opposite actions from an author: upgrade the reader, or fix the cell. The message is pydantic's
under extra="forbid", so it cannot be reworded into carrying the distinction: the information was
not in the model at all.
On the field, not in a roster. A list keyed like release_records was the alternative and
loses on the rule this repo keeps relearning — a hand-kept list beside a model is a second
statement of one fact, and it is the copy that goes stale (@fieldnames-from-model,
@registry-completeness). Declared here it travels with the field through every rename and move,
and test_first_seen.py asserts an equality over the walked registry, so the next column added
cannot omit one.
The answer is per (model, field), never per name, which curator is the worked example of: it
is on VariantRow from 0.2.0 and gains its StudyRow twin only in 0.6.5. A roster keyed by column
name would give one answer for two facts.
Composes with vocabulary() rather than replacing it — both are entries in one
json_schema_extra dict, so a field can carry either or both:
Field(json_schema_extra={**vocabulary("state", VALID_STATES), **since("0.2.0")})
Source code in schema/src/just_dna_format/base.py
field_first_seen ¶
{field_name: release} for every field of model that declares one (RM146).
The public reader, so a consumer rendering our findings can answer when did this column appear
offline rather than parsing model_fields themselves — which is what the reporter would
otherwise have had to do, and what their own rulebook forbids.
Source code in schema/src/just_dna_format/base.py
field_vocabularies ¶
{field_name: {name, options, closed[, notes]}} for every vocabulary-bound field of model.
The single route to "what may this cell contain" — authoring_reference() and any authoring tool
read it, and neither keeps a list of its own. Covers both binding sites: a field marked at its own
declaration, and a field whose vocabulary is enforced by AuthoredModel's shared validators.
Source code in schema/src/just_dna_format/base.py
derive_variant_key ¶
derive_variant_key(
rsid: str | None,
chrom: str | None,
start: int | None,
ref: str | None,
alts: str | None = None,
*,
build: str = "GRCh38",
) -> str
The natural identity for a variant-ish row: the rsid when present, else the coordinate.
Three cases, in precedence order:
- rsid — an rsid row keeps its rsid, unchanged. (A dbSNP id is position/multi-allelic-level, not per-allele, which is why clinical identity keys on this plus genotype, never on rsid.)
- A resolved single-base substitution (0.5) — the key is its GA4GH VRS allele id,
ga4gh:VA.…, minted byvrs.derive_vrs_allele_id. This is a content-addressed identity: it names the allele by the digest of the exact reference sequence it sits on, so it is build-naming rather than build-ambiguous — the property RM15 was waiting for before coordinate identity could be reconsidered. Byte-identical to the ids gnomAD and ClinGen serve, so it joins against them directly instead of needing a translation table. - Everything else — the coordinate key
chrom:start:ref, orchrom:start:ref:altswhen an alt is given, so two distinct alleles at one locus (an insertionC>CAAAGbeside a deletionC>CA, a benignC>Gbeside a pathogenicC>A) do not collide.altsis normalized (its comma-separated alleles sorted) so the key is stable regardless of authored order.
Case 2 covers substitutions only — an indel, an MNV, a multi-allelic cell and a contig outside the
primary assembly all fall through to case 3, because a VRS allele id is defined over the fully
justified allele and justifying an indel needs the reference sequence, which this tier will never
fetch (Principle 2). See vrs.derive_vrs_allele_id. That split is deliberate and permanent-shaped:
an id is minted only where it can be minted correctly, and the fallback is the same key these
rows already had.
Single source of truth shared by VariantRow (which freezes the result into a stored column so
resolution can never re-key a row) and the one-to-many expansion re-keying. StudyRow and the
position-level matching helpers deliberately call this without alts — a study is
position/rsid evidence and matches a variant at chrom:start:ref regardless of which allele it
carries — and so are never affected by case 2. See docs/COMPILER.md: the frozen variant_key keeps
a position-only row that later resolves to an rsid from flipping its identity, and lets a
one-to-many rsid expand to distinct coord-keyed rows (Principle 7).
build is the assembly the coordinate is in; only GRCh38 has a refget table today, and any other
build falls through to case 3 rather than minting an id that would claim the wrong sequence.
Source code in schema/src/just_dna_format/base.py
stamped_identity_field ¶
A Field(...) for a compiler-stamped column that is not part of the authored content.
Three properties, and the middle one is the non-obvious one:
COMPILER_MANAGED, soauthored_field_namesdrops it — no author is offered the column, no drafted template writes it, andreverse_module's generic_write_table_csvnever re-emits it.exclude=True, so it stays out ofmodel_dump()and therefore out ofcontent_signature. A stamped value is a pure function of the authored cells, so it adds nothing to a content identity — and including it would move the signature of every already-published module carrying one of these tables, which is the one thing a content-dedup key may not do._build_tablereads the field off the model directly (not throughmodel_dump()), so the column still reaches parquet.VariantRow.variant_key/authored_identare not excluded and are insidecontent_signaturetoday; that is a grandfathered inconsistency, not a precedent — un-excluding them here, or excluding them there, moves published signatures either way, so the asymmetry is carried until a major. Anything declared through this helper is on the right side of it, includingVariantRow's ownlocus_index/locus_count(RM87), which are stamped on the same model without repeating the defect.- a fresh
FieldInfoper call, because pydantic binds one to the model that declares it.
first_seen is required rather than defaulted (RM146). A compiler-stamped column is still a
column an older reader refuses under extra="forbid", so it owes the same answer as an authored
one — and the guard walks model_fields, which does not distinguish them. Defaulting it would let
the next stamped column inherit a version nobody measured, which is the whole failure mode.
default is the caller's, not this helper's. It was hard-coded to None while the only users
were the three 0.4-family positional models (RM43), where an unstamped identity column genuinely
has no value. RM87's locus_count defaults to 1 instead, because a row that was never expanded
really does resolve to one locus and locus_count > 1 has to be a predicate a reader can apply
while holding a single row — a None/0 default would put a second "undetermined" state back
into the column the item exists to make determinate.
Source code in schema/src/just_dna_format/base.py
reject_compiler_filled ¶
Refuse an identity column this model has the compiler fill (RM43), with a diagnosis.
alts is the only such column today: authored and required on variants.csv, authored and
optional on heteroplasmy.csv, and compiler-filled on pharm_variants.csv/haplotypes.csv,
where a pharm annotation matches a variant at chrom:start:ref regardless of allele. Writing it
there by analogy with variants.csv is a plausible confusion rather than a typo, and before this
guard it was silently accepted: extra="forbid" cannot see it, the value entered
authored_ident, and _write_table_csv then dropped it — so compile → reverse → compile was
not a fixed point. It failed loudly before 0.6 added the column, so a bare accept is a regression
however sensible the input.
Scoped to IDENTITY_FIELDS deliberately, which is what leaves variant_key/authored_ident
accepted-and-overwritten: those two are stamped rather than filled, and VariantRow has
tolerated them since 0.5 on the "authored values ignored, no foot-gun" rule. Nothing is lost by
ignoring a stamped value; something is lost by ignoring a filled one.
Source code in schema/src/just_dna_format/base.py
authored_identity ¶
Which of IDENTITY_FIELDS this row's model declares and the author actually filled.
Source code in schema/src/just_dna_format/base.py
stamp_identity ¶
Stamp variant_key (and, at construction, authored_ident) onto a positional row.
The key is derived from the authored subset only, never from the row's current cells, which is
what makes it survive the compile-time coordinate fill: once authored_ident is frozen, filling
chrom/start/ref/alts cannot re-key the row however many times this runs. That is the same
guarantee VariantRow gets from _freeze_identity never re-running on model_copy, arrived at
from the other direction because these rows are filled in place rather than copied.
freeze_authored is True exactly once, at construction, and authored values are overwritten
rather than trusted — a CSV that carries a variant_key/authored_ident column of its own does
not get to declare its own identity (VariantRow calls that "no foot-gun"). with_genome_build
re-derives the key only, because the assembly changes what a coordinate means and cannot change
what the author wrote.
Source code in schema/src/just_dna_format/base.py
genotype_allele_ok ¶
Whether one member of a genotype is spellable: a nucleotide string, a symbolic allele, or *.
The single place _validate_genotype decides what an allele is, so the three arms of the
grammar (phased pair, hemizygous single, unphased pair) cannot drift apart — they did not, but
the check was written out three times and a widening had to be applied three times with it.
Symbolic alleles joined in 0.6 (RM5): a genotype may name one (<DEL:1500>/A — a heterozygous
deletion), because that is exactly the call a consumer reading a structural VCF has in hand.
A lengthless one passes here and is refused by the compiler; see vocab.validate_allele for
why the split is forced rather than chosen.
* joined beside them in 0.6 (RM59), and the two arms are deliberately not one arm. They are
checked by different predicates against different modules of meaning, and the separation is the
item rather than an implementation detail: parse_symbolic_allele answers which variant is this,
unspelled, and is_unobservable_allele answers whether this sample's allele could be seen at
all — the callability axis requires_callable already owns (P5). Folding * into
SYMBOLIC_ALLELE_TYPES would give one syntax two meanings and hand * a length, an event and a
place it does not have. See alleles' RM59 section for the full argument.
Why the genotype column and not the allele columns: vocab.validate_allele (HaplotypeRow.allele,
VariantRow.effect_allele) still refuses *, and must. Those columns name the allele a rule is
about, and a rule about an allele nobody observed states nothing. A genotype is the one place the
observation itself is written down, which is why it is the one place * belongs.