just_dna_format.spec¶
just_dna_format.spec ¶
The authored module spec DSL (module_spec.yaml + variants.csv + studies.csv).
This is the input half of the module format; manifest.py is the output half. Both live in
this dependency-light package so the compiler is a pure transform between two validated schema
sets, and any consumer can validate a spec or a manifest without pulling the compiler's polars/
duckdb weight.
Identity/display rules reuse the shared helpers in identity and manifest, so the DSL and the
manifest enforce exactly the same constraints.
ModuleInfo ¶
Bases: Display
The module: block of module_spec.yaml: a machine name plus the shared Display
metadata (title/description/report_title/icon/color).
Extends the manifest's Display rather than re-declaring those fields, so the display schema
and its validation (e.g. the hex-colour rule) live in exactly one place. name lives here on
the authoring side; the manifest routes it into Identity instead.
extra="forbid" so an authored-block typo (colour:, nam:) is a hard error, not a silently
dropped key — the same author-time guard the row models carry, applied to the module: block.
(Set here, not on Display, so the manifest side that also uses Display is untouched.)
version_coerced_from
property
¶
The authored version before SemVer coercion, or None if it was already SemVer.
Not a field: it describes what happened during validation, not module content, so it stays out
of model_dump() and out of every CSV.
It does reach the manifest, since RM103 — Identity.version_coerced_from, copied by the
compiler rather than serialized from here. That is the distinction this docstring used to
blur: staying out of the authored surface is a property of the field, while publishing the
fabrication is what makes it auditable in the artifact instead of only in a build log
somebody had to have kept.
Defaults ¶
Bases: BaseModel
Default values applied to variant rows when not explicitly set.
extra="forbid" so a typo'd default key (currator:) is rejected rather than silently ignored
(which would leave the real default in force with no diagnosis).
ModuleSpecConfig ¶
Bases: BaseModel
Top-level model for module_spec.yaml.
extra="forbid" so a misspelled top-level key is a hard error, not a silent no-op. This is
safety-relevant, not merely tidy: a typo'd genome_bild: would otherwise leave genome_build
at its GRCh38 default, silently resolving a GRCh37-intended module against the wrong assembly
— exactly the corruption the resolver's build guard exists to prevent, bypassed by one typo.
VariantRow ¶
Bases: AuthoredModel
One row of variants.csv. At least one identifier (rsid or chrom+start) is required.
Inherits from AuthoredModel: extra="forbid" + the reserved-namespace guard, and the shared
field validators for rsid/trait_efo_id/direction/clin_sig/stat_significance/effect_size.
effective_direction
property
¶
direction if set, else derived from the legacy state (+ weight sign).
effective_stat_significance
property
¶
stat_significance if set, else derived from the legacy state.
effective_clin_sig
property
¶
clin_sig if set, else derived from the legacy ClinVar booleans (lossy).
effective_pathogenic
property
¶
The authoritative pathogenic boolean, or the one implied by clin_sig when unset.
effective_benign
property
¶
The authoritative benign boolean, or the one implied by clin_sig when unset.
needs_upgrade
property
¶
True when a re-publish would materialize a 0.3 column that is currently derived-but-empty
(or would re-align the legacy state). Feeds the marketplace revalidate/needs_upgrade
contract-drift flow (which flags drifted-but-fixable modules for a new PATCH).
upgraded ¶
A copy with the 0.3 axes back-populated from state/booleans and state trimmed to the
legacy set {protective, risk, neutral}. state stays present (never dropped inside a major)
but becomes a derived mirror of direction. Idempotent: r.upgraded().upgraded() ==
r.upgraded().
Source code in schema/src/just_dna_format/spec.py
StudyRow ¶
Bases: AuthoredModel
One row of studies.csv: an (rsid, pmid) evidence link. Grounding evidence is mandatory.
Inherits AuthoredModel (reserved-namespace guard + shared rsid/trait_efo_id/
stat_significance/effect_size validators).
variant_key
property
¶
Stable key matching VariantRow.variant_key, or None when the row names no variant.
StudyRow is never resolved/expanded, so its key stays a derived property (no freezing needed).
None is the 0.6 shape (RM47): a citation row may now describe the module or a gene-keyed
binning table rather than a variant, and derive_variant_key(None, None, None, None) would
otherwise hand back the string "None:None:None" — a key that looks like an identity and
names nothing. Same reasoning as HeteroplasmyRow.variant_key, and the reason the orphan half
of the compiler's study cross-check skips these rows: a row that names no variant cannot
reference one that is missing.
neg_log10_p
property
¶
−log10(p) — derived on write, never stored. None when p_value_num is unset.
The scale a consumer actually filters and plots on (7.3 is genome-wide significance),
materialized into studies.parquet so nobody recomputes it per query. It is deliberately not
authored in this form: that would make the human compute a logarithm to write a row down,
and the DSL exists for the human — the parquet absorbs the convenience, the CSV keeps the
number the paper printed.
extract_pmcids ¶
Pull PubMed Central ids out of a free-form reference string, canonicalized to PMC….
In order, de-duplicated, and tolerant of the spacings a hand-written cell produces (PMC3110566,
PMC 3110566, pmcid: 3110566). This exists so a refusal can name the identifier the author
actually wrote — a generic "must contain at least one PubMed ID" is a dead end where a specific
"that is a PMCID" is a fix — and so the enricher can cross-check an authored PMC id against the
one PubMed reports for the same record.
Source code in schema/src/just_dna_format/spec.py
extract_pmids ¶
Pull digit-only PMIDs out of a free-form reference string, in order, de-duplicated.
Handles bare digits, the bracketed/prefixed [PMID: N] / PMID N forms, and ;-joined
lists. Returns an empty list when the string carries no PMID token (e.g. a dbSNP URL).
A digit run whose immediate context spells PMC is not a PMID and is skipped (RM50). It is a
different registry's number for (usually) a different article, so reading it as a PubMed id is a
confident citation of the wrong paper. A cell carrying both — 21551363; PMC3110566 — still
yields the real PMID; only a cell whose sole numeric content is a PMC id comes back empty, which
is what lets validate_pmid_cell name it.
Source code in schema/src/just_dna_format/spec.py
validate_pmid_cell ¶
The shared grammar for a free-form citation pointer, kept verbatim on success.
Every citation pointer in the schema routes through here, so the rule cannot drift between
them as citation sites are added — and they are added: StudyRow.pmid (required, grounding a
variant), MeasureBinRow.pmid (optional, RM47's bin pointer grounding a boundary) and
PharmVariantRow.pmid (optional, RM132, grounding one drug/genotype claim). required is the
only thing that differs between them; the diagnosis is not.
The PMCID branch is the whole point of separating this out (RM50). Before it, a cell reading
PMC 3110566 was accepted as PMID 3110566 and a cell reading PMC3110566 was refused with a
message that never used the word PMCID — the same mistake, diagnosed two different unhelpful ways.
Now both are refused and the message names the id that was seen. It never repairs: converting a
PMC id to a PMID needs the registry, which this tier does not have and would not use if it did
(just-dna-enricher hint citation --pmcid is the reporting route).