Skip to content

just_dna_format.spec

just_dna_format.spec

The authored module spec DSL (module_spec.yaml + variants.csv + studies.csv).

This is the input half of the module format; manifest.py is the output half. Both live in this dependency-light package so the compiler is a pure transform between two validated schema sets, and any consumer can validate a spec or a manifest without pulling the compiler's polars/ duckdb weight.

Identity/display rules reuse the shared helpers in identity and manifest, so the DSL and the manifest enforce exactly the same constraints.

ModuleInfo

Bases: Display

The module: block of module_spec.yaml: a machine name plus the shared Display metadata (title/description/report_title/icon/color).

Extends the manifest's Display rather than re-declaring those fields, so the display schema and its validation (e.g. the hex-colour rule) live in exactly one place. name lives here on the authoring side; the manifest routes it into Identity instead.

extra="forbid" so an authored-block typo (colour:, nam:) is a hard error, not a silently dropped key — the same author-time guard the row models carry, applied to the module: block. (Set here, not on Display, so the manifest side that also uses Display is untouched.)

version_coerced_from property

version_coerced_from: str | None

The authored version before SemVer coercion, or None if it was already SemVer.

Not a field: it describes what happened during validation, not module content, so it stays out of model_dump() and out of every CSV.

It does reach the manifest, since RM103 — Identity.version_coerced_from, copied by the compiler rather than serialized from here. That is the distinction this docstring used to blur: staying out of the authored surface is a property of the field, while publishing the fabrication is what makes it auditable in the artifact instead of only in a build log somebody had to have kept.

Defaults

Bases: BaseModel

Default values applied to variant rows when not explicitly set.

extra="forbid" so a typo'd default key (currator:) is rejected rather than silently ignored (which would leave the real default in force with no diagnosis).

ModuleSpecConfig

Bases: BaseModel

Top-level model for module_spec.yaml.

extra="forbid" so a misspelled top-level key is a hard error, not a silent no-op. This is safety-relevant, not merely tidy: a typo'd genome_bild: would otherwise leave genome_build at its GRCh38 default, silently resolving a GRCh37-intended module against the wrong assembly — exactly the corruption the resolver's build guard exists to prevent, bypassed by one typo.

VariantRow

Bases: AuthoredModel

One row of variants.csv. At least one identifier (rsid or chrom+start) is required.

Inherits from AuthoredModel: extra="forbid" + the reserved-namespace guard, and the shared field validators for rsid/trait_efo_id/direction/clin_sig/stat_significance/effect_size.

effective_direction property

effective_direction: str

direction if set, else derived from the legacy state (+ weight sign).

effective_stat_significance property

effective_stat_significance: str

stat_significance if set, else derived from the legacy state.

effective_clin_sig property

effective_clin_sig: str | None

clin_sig if set, else derived from the legacy ClinVar booleans (lossy).

effective_pathogenic property

effective_pathogenic: bool | None

The authoritative pathogenic boolean, or the one implied by clin_sig when unset.

effective_benign property

effective_benign: bool | None

The authoritative benign boolean, or the one implied by clin_sig when unset.

needs_upgrade property

needs_upgrade: bool

True when a re-publish would materialize a 0.3 column that is currently derived-but-empty (or would re-align the legacy state). Feeds the marketplace revalidate/needs_upgrade contract-drift flow (which flags drifted-but-fixable modules for a new PATCH).

upgraded

upgraded() -> VariantRow

A copy with the 0.3 axes back-populated from state/booleans and state trimmed to the legacy set {protective, risk, neutral}. state stays present (never dropped inside a major) but becomes a derived mirror of direction. Idempotent: r.upgraded().upgraded() == r.upgraded().

Source code in schema/src/just_dna_format/spec.py
def upgraded(self) -> "VariantRow":
    """A copy with the 0.3 axes back-populated from `state`/booleans and `state` trimmed to the
    legacy set {protective, risk, neutral}. `state` stays present (never dropped inside a major)
    but becomes a derived mirror of `direction`. Idempotent: ``r.upgraded().upgraded() ==
    r.upgraded()``."""
    direction = self.effective_direction
    return self.model_copy(
        update={
            "direction": direction,
            "stat_significance": self.effective_stat_significance,
            "clin_sig": self.effective_clin_sig,
            "pathogenic": self.effective_pathogenic,
            "benign": self.effective_benign,
            "state": trimmed_state(direction),
        }
    )

StudyRow

Bases: AuthoredModel

One row of studies.csv: an (rsid, pmid) evidence link. Grounding evidence is mandatory.

Inherits AuthoredModel (reserved-namespace guard + shared rsid/trait_efo_id/ stat_significance/effect_size validators).

variant_key property

variant_key: str | None

Stable key matching VariantRow.variant_key, or None when the row names no variant.

StudyRow is never resolved/expanded, so its key stays a derived property (no freezing needed).

None is the 0.6 shape (RM47): a citation row may now describe the module or a gene-keyed binning table rather than a variant, and derive_variant_key(None, None, None, None) would otherwise hand back the string "None:None:None" — a key that looks like an identity and names nothing. Same reasoning as HeteroplasmyRow.variant_key, and the reason the orphan half of the compiler's study cross-check skips these rows: a row that names no variant cannot reference one that is missing.

neg_log10_p property

neg_log10_p: float | None

−log10(p) — derived on write, never stored. None when p_value_num is unset.

The scale a consumer actually filters and plots on (7.3 is genome-wide significance), materialized into studies.parquet so nobody recomputes it per query. It is deliberately not authored in this form: that would make the human compute a logarithm to write a row down, and the DSL exists for the human — the parquet absorbs the convenience, the CSV keeps the number the paper printed.

extract_pmcids

extract_pmcids(raw: str) -> list[str]

Pull PubMed Central ids out of a free-form reference string, canonicalized to PMC….

In order, de-duplicated, and tolerant of the spacings a hand-written cell produces (PMC3110566, PMC 3110566, pmcid: 3110566). This exists so a refusal can name the identifier the author actually wrote — a generic "must contain at least one PubMed ID" is a dead end where a specific "that is a PMCID" is a fix — and so the enricher can cross-check an authored PMC id against the one PubMed reports for the same record.

Source code in schema/src/just_dna_format/spec.py
def extract_pmcids(raw: str) -> list[str]:
    """Pull PubMed **Central** ids out of a free-form reference string, canonicalized to `PMC…`.

    In order, de-duplicated, and tolerant of the spacings a hand-written cell produces (`PMC3110566`,
    `PMC 3110566`, `pmcid: 3110566`). This exists so a refusal can name the identifier the author
    actually wrote — a generic "must contain at least one PubMed ID" is a dead end where a specific
    "that is a PMCID" is a fix — and so the enricher can cross-check an authored PMC id against the
    one PubMed reports for the same record.
    """
    seen: dict[str, None] = {}
    for match in PMCID_PATTERN.finditer(raw):
        seen.setdefault(f"PMC{match.group(1)}", None)
    return list(seen)

extract_pmids

extract_pmids(raw: str) -> list[str]

Pull digit-only PMIDs out of a free-form reference string, in order, de-duplicated.

Handles bare digits, the bracketed/prefixed [PMID: N] / PMID N forms, and ;-joined lists. Returns an empty list when the string carries no PMID token (e.g. a dbSNP URL).

A digit run whose immediate context spells PMC is not a PMID and is skipped (RM50). It is a different registry's number for (usually) a different article, so reading it as a PubMed id is a confident citation of the wrong paper. A cell carrying both — 21551363; PMC3110566 — still yields the real PMID; only a cell whose sole numeric content is a PMC id comes back empty, which is what lets validate_pmid_cell name it.

Source code in schema/src/just_dna_format/spec.py
def extract_pmids(raw: str) -> list[str]:
    """Pull digit-only PMIDs out of a free-form reference string, in order, de-duplicated.

    Handles bare digits, the bracketed/prefixed `[PMID: N]` / `PMID N` forms, and `;`-joined
    lists. Returns an empty list when the string carries no PMID token (e.g. a dbSNP URL).

    **A digit run whose immediate context spells `PMC` is not a PMID and is skipped** (RM50). It is a
    different registry's number for (usually) a different article, so reading it as a PubMed id is a
    confident citation of the wrong paper. A cell carrying both — `21551363; PMC3110566` — still
    yields the real PMID; only a cell whose sole numeric content is a PMC id comes back empty, which
    is what lets `validate_pmid_cell` name it.
    """
    excluded = [match.span(1) for match in PMCID_PATTERN.finditer(raw)]
    seen: dict[str, None] = {}
    for match in PMID_PATTERN.finditer(raw):
        start = match.start(1)
        if any(lo <= start < hi for lo, hi in excluded):
            continue
        seen.setdefault(match.group(1), None)
    return list(seen)

validate_pmid_cell

validate_pmid_cell(
    value: str | None, field: str, *, required: bool
) -> str | None

The shared grammar for a free-form citation pointer, kept verbatim on success.

Every citation pointer in the schema routes through here, so the rule cannot drift between them as citation sites are added — and they are added: StudyRow.pmid (required, grounding a variant), MeasureBinRow.pmid (optional, RM47's bin pointer grounding a boundary) and PharmVariantRow.pmid (optional, RM132, grounding one drug/genotype claim). required is the only thing that differs between them; the diagnosis is not.

The PMCID branch is the whole point of separating this out (RM50). Before it, a cell reading PMC 3110566 was accepted as PMID 3110566 and a cell reading PMC3110566 was refused with a message that never used the word PMCID — the same mistake, diagnosed two different unhelpful ways. Now both are refused and the message names the id that was seen. It never repairs: converting a PMC id to a PMID needs the registry, which this tier does not have and would not use if it did (just-dna-enricher hint citation --pmcid is the reporting route).

Source code in schema/src/just_dna_format/spec.py
def validate_pmid_cell(value: str | None, field: str, *, required: bool) -> str | None:
    """The shared grammar for a free-form citation pointer, kept verbatim on success.

    **Every citation pointer in the schema routes through here**, so the rule cannot drift between
    them as citation sites are added — and they are added: `StudyRow.pmid` (required, grounding a
    variant), `MeasureBinRow.pmid` (optional, RM47's bin pointer grounding a *boundary*) and
    `PharmVariantRow.pmid` (optional, RM132, grounding one drug/genotype claim). `required` is the
    only thing that differs between them; the diagnosis is not.

    The PMCID branch is the whole point of separating this out (RM50). Before it, a cell reading
    `PMC 3110566` was accepted as PMID 3110566 and a cell reading `PMC3110566` was refused with a
    message that never used the word PMCID — the same mistake, diagnosed two different unhelpful ways.
    Now both are refused and the message names the id that was seen. **It never repairs**: converting a
    PMC id to a PMID needs the registry, which this tier does not have and would not use if it did
    (`just-dna-enricher hint citation --pmcid` is the reporting route).
    """
    if value is None:
        if required:
            raise ValueError(f"{field} must not be empty")
        return None
    value = str(value).strip()
    if not value:
        if required:
            raise ValueError(f"{field} must not be empty")
        return None
    if extract_pmids(value):
        return value  # kept verbatim; use extract_pmids(value) to recover digit-only ids
    pmcids = extract_pmcids(value)
    if pmcids:
        raise ValueError(
            f"{field} names PubMed Central id(s) {pmcids} and no PubMed ID. PMC and PubMed number "
            f"articles independently, so a PMC id's digits are a different article's PMID — reading "
            f"them as one would cite the wrong paper. Look the PMID up (e.g. "
            f"`just-dna-enricher hint citation --pmcid {pmcids[0]}`) and write that, got: {value!r}"
        )
    raise ValueError(
        f"{field} must contain at least one PubMed ID (bare digits, or a bracketed/prefixed "
        f"form like '[PMID: 9545397]'), got: {value!r}"
    )