Skip to content

just_dna_format.manifest

just_dna_format.manifest

The manifest.json contract — the single source of truth for a compiled annotation module.

Mirrors SPEC §4. Fields known at compile time (display, stats, compilation, inputs, artifact) are filled by the compiler; marketplace-level fields (namespace, version, owner, published_at, canonical_id) are Optional and filled by the marketplace on publish. license reads like one of them and is not: since 0.5 an author declares it in module_spec.yaml, the compiler copies it through and cross-checks it against the licensing table, and no registry stamps over it. version is the field that really is advisory-then-stamped, and describing the two as one pattern was a claim nothing performed (RM111).

This module is intentionally dependency-light (Pydantic + stdlib only) so both just-dna-pipelines (which emits the manifest) and just-dna-marketplace (which consumes and extends it) can share one definition without pulling heavy transitive dependencies.

Identity

Bases: BaseModel

Module identity. namespace/version/canonical_id are filled by the marketplace.

Identity rules are validated here using the shared just_dna_format.identity helpers, so the contract enforces exactly what just-dna-pipelines enforces on module_spec.yaml.

Display

Bases: BaseModel

Shared display metadata for a module. The authoring DSL's spec.ModuleInfo extends this (adding name), so the fields and their validation are defined here once.

Stats

Bases: BaseModel

Card/detail stats derived from the spec at compile time — from the spec, not from one table.

clinvar_count/pathogenic_count/benign_count summarize the per-row ClinVar quality flags that weights.parquet already carries, so consumers can facet on them without reading the artifact (SPEC ROADMAP item 5). They are additive and default to 0 for older manifests. Each counts the legacy boolean, so a count spans a tier pair (pathogenic + likely_pathogenic), and counts authored rows, one per genotype, never distinct variants.

gene_count/genes are a union over every authored table kind with a gene column, which is what the first sentence has always said and what the compiler did not do until RM121: they were read off variants.csv alone, so a star-allele or copy-number module published [] while every one of its rows named a gene, and a registry gene index fed from this field could not find it. Derived fact tables are excluded on purpose — gene_metrics.genes carries those.

Compilation

Bases: BaseModel

Provenance of the compile that produced this artifact (SPEC §5 trust fields).

Frequency

Bases: BaseModel

Summary of a module's injected allele-frequency sidecar (0.5), out of artifact.digest.

A separate block rather than extra fields on Compilation/Resolution: Resolution is about rsID↔coordinate resolution and nothing else, and a frequency table has its own producer, its own release, and its own fact-hash. Absent on a module that carries no frequencies.csv.

GeneMetrics

Bases: BaseModel

Summary of a module's injected gene-constraint sidecar (0.5), out of artifact.digest.

GeneValidity

Bases: BaseModel

Summary of a module's injected gene–disease validity sidecar (0.6, RM24).

Out of artifact.digest, like every sibling block. classifications is emitted sorted rather than in strength order, for the reason Frequency.populations is emitted in canonical order and everything else sorted: a set-like facet has no order of its own, and a sorted list is the only one that cannot drift. A consumer that wants the ladder reads vocab.ORDERED_GENE_VALIDITY, which is published precisely so this block does not have to encode it.

ClinicalAssertions

Bases: BaseModel

Summary of a module's injected clinical-assertion sidecar (0.6, RM25).

The counters exist for the same reason Literature's do: the point of the table is that a one-star single submission and a practice guideline are not the same claim, so a summary that reported only a row count would throw away exactly what was gained. max_review_stars and min_review_stars are the two ends a catalog can filter on without reading the parquet — published as the two counts rather than an average, which would be a number describing no record.

ExpressionEffects

Bases: BaseModel

Summary of a module's injected expression-effect sidecar (0.7, RM194/RM200).

The facets answer one question without reading the parquet: can these effects be thresholded at all? For this table that is a real question with a frequently unwelcome answer, the same way GwasEffects.units is.

without_distance is the one that matters most. AlphaGenome attributes a variant to genes across a 1 Mb window, and scores at the far edge run about an order of magnitude lower than scores at the gene — so a consumer filtering on magnitude alone keeps only the proximal rows and believes it has filtered on effect. distance_to_gene is what makes a distance-aware threshold possible, it is filled from the MANE lane, and a deployment without that lane provisioned produces a table where it is null throughout. Published as a count so that case is visible here rather than discovered after a join.

without_direction is the other. A row whose tracks split evenly carries a magnitude and no direction, which is real evidence about a locus and cannot be used as a directional claim. Counted beside its complement rather than filtered out, because a consumer that silently dropped those rows and one that silently kept them would both be wrong in ways nothing could see.

genes is published in full rather than as a count. The rows are locus-wide by design — most name variants the module does not author — so the gene list is what says what this table is about, and it is bounded by what was queried rather than by what was scored.

GwasEffects

Bases: BaseModel

Summary of a module's injected GWAS-effect sidecar (0.6, RM90).

The facets are chosen so a consumer can decide whether these effects are usable at all without reading the parquet, because for this table that is a real question with a frequently unwelcome answer.

units is the one that matters most. The whole reason the table exists is that a magnitude without its unit is the defect S36 reported one layer down, so the manifest publishes the set of units present: a module whose GWAS effects span umol/l, cm and the Catalog's uninformative unit is one whose effects cannot be pooled, and that is visible here rather than only after a join.

without_effect_allele is the other. The Catalog writes rs4149056-? when a study did not establish which allele carries the effect, and such a row is real evidence about a locus but cannot be used as a weight. Published as a count beside its complement rather than filtered out, because a consumer that silently dropped those rows and one that silently kept them would both be wrong in ways nothing could see.

Literature

Bases: BaseModel

Summary of a module's injected citation sidecar (0.5), out of artifact.digest.

No datasets field, unlike its two siblings: PubMed and Europe PMC publish no release identifier, so there would be nothing true to put in it. The coverage counters take its place — and they are what a reader actually needs, because the fulltext check is partial by nature and a summary that hid that would read as "all citations verified" when most of them were never retrievable.

ClinSigConcordance

Bases: BaseModel

Summary of a module's clinical-significance concordance record (0.7, RM130).

The two counters that matter are opposed_count and unchecked_count, and they answer different questions from row_count. A record of forty subjects where every disagreement is a pathogenic-against-benign pair is a very different module from one where forty rows differ by a confidence step, and the split is the same one the check itself draws.

unchecked_count is published beside them because a run that could not reach an authority is not a run that found agreement. Absent it, a shrinking record would read as an improving module when it may only mean a snapshot went missing.

No consensus field, and the omission is the decision. Nothing here says which authority is right when two disagree: resolving a split needs a weighting model this format does not have, and precomputing one would publish a judgement as though it were a fact.

Sources

Bases: BaseModel

Summary of a module's data-source licensing sidecar (0.5), out of artifact.digest.

The per-layer facets are lists, and collapsing them to booleans would be a defect. A module that used CPIC only to resolve a coordinate and one that embeds ClinPGx annotation prose would render identically under a single share_alike: bool, falsely marking the first as viral. The lists say which layer carries the obligation, so a reader can tell those apart.

commercial_use is the one derived scalar, because it is the one question with a single answer for the module as a whole: most-restrictive-wins. One restricted source at the annotation layer makes the whole module non-sellable, and mixing in a permissive source cannot launder it.

FileEntry

Bases: BaseModel

One hashed file — used for both inputs[] and artifact.files[] (SPEC §5).

Artifact

Bases: BaseModel

The compiled output set plus its Merkle-root digest (the content identity).

GenePanelSpec

Bases: BaseModel

Declares a module derived from a gene set + significance predicate over a reference, rather than an enumerated variant table (SPEC ROADMAP item 7).

Deprecated in 0.6, removed at 1.0 (RM4). Compile-time materialization was dropped rather than built: the compiler must not create rows no curator wrote, and expanding a declaration at compile would make a module's content depend on an external file while leaving reverse to choose between re-emitting the declaration (rows lost) and the rows (declaration lost) — neither a fixed point. The want is served by enricher draft-scaffolding, where the rows are authored bytes before the compiler sees them and the author's no-op over the drafted subset is still an authorial act. The block's one remaining machine reader — the enricher's ClinVar clin_sig cross-check, deciding whether a drafted module is being compared against its own source — now reads the licence row's dataset column instead, which the drafting pass writes itself.

This is the authored interface only: the compiler records it verbatim but does not materialize it (an app-level adapter enumerates the matching variants into variants.csv today). Optional and backwards-compatible — absent on ordinary variant modules.

extra="forbid" so a typo in the authored panel: block is caught, not silently dropped.

ProvenanceItem

Bases: BaseModel

One per-variant provenance record (SPEC ROADMAP item 1). Lives in the full provenance.json document, not in the manifest — the manifest carries only the Provenance summary pointer.

An item stays per variant, and outranks is what carries per-field judgement (S52). A row may outrank an archive on clin_sig while its direction is ordinary and unjustified, so a single rationale string cannot say which field a justification is about — and a check keyed on an item exists for this variant would downgrade every field's mismatch on the row at once. Of the three shapes the reporter put up, a field on the item was refused for a reason outside their list: Provenance.item_count is a published manifest number meaning variants with a record, and making items per-(variant, field) would silently change what it counts for every consumer already reading it — an S14-shaped break, where the addition is legal and the redefinition is not. Accepting the bluntness was refused on the reporter's own argument.

Presence is the machine-readable bit; the prose is for a human. A key in outranks says a justification exists for that column, and that is all a check may act on — never a parse of the text, whose whole justification is that the judgement is not formalizable. Nothing in this format reads outranks today: the capture half is the authoring layer's, and whether a check downgrades a mismatch to INFO is an open question this field does not settle. What it settles is that the record has somewhere to live that is not a string per variant.

ProvenanceDoc

Bases: BaseModel

The full provenance.json authored beside the spec: a header plus per-variant items. The compiler reads and hashes it, then records the lean Provenance summary in the manifest so catalog cards can flag 'AI-authored · rationale available' without inlining the full text.

Provenance

Bases: BaseModel

Lean summary pointer to a version's provenance.json (SPEC ROADMAP item 1). The full items live in the hashed file (kept out of artifact.digest, like logs); this rides in the manifest.

Signature

Bases: BaseModel

Optional detached signature over artifact.digest (SPEC §5 'future'). Defends against a compromised storage backend: a client that pins the marketplace's public key can prove the digest was signed by the trusted party.

Declared here rather than beside ModuleManifest, where it sat until 0.6, because Closure below signs a different hash with the identical shape. Nothing about it is artifact-specific: the signed message is whichever digest string the caller hands signing.sign_digest.

Closure

Bases: BaseModel

The author's statement that this authored set is finished (RM73).

Authoring is a process and until 0.6 it had no end. Every check that needed to know where a value came from had to guess, and each guessed differently, because a flat CSV row records nothing about how it came to be. The provenance half of RM73 answered did this cell move by hashing what a drafting provider wrote; this answers the other question — is the author done — and the two are genuinely different, which is why one is a column on a licence row and this is a signable act.

What it binds is the document's module_hash, and nothing here re-derives it. The closure rides inside VerificationDoc, whose binding the compiler already recomputes and drops on mismatch, so an author who edits a row after closing loses the closure for free — the same perishability the check records have, arrived at by carrying no second copy of the hash. That is the whole mechanism: no new file, no new binding, and no entry in verification.pow_digest, whose payload is deliberately unchanged so that closing a document re-mines nothing and every attestation written before this still verifies.

closed_by is untrusted and signature is not. A name in a JSON file is a claim anyone can type; the Ed25519 signature over module_hash is what makes the act attributable, using the same signing.sign_digest / integrity.verify_signature pair the artifact signature uses. Signing is optional — an unsigned closure is still change-evident, which is the guarantee this format offers (tamper-evidence, never tamper-proofing) — and a present signature that does not verify drops the whole block, because a false claim is worse than silence.

VerificationRecord

Bases: BaseModel

One check, and what putting it produced — the row grain of verification.json (RM45).

A module whose clinical-significance calls were cross-checked against ClinVar and one where that check never ran used to ship identical manifests: not through an oversight in some path, but because no field existed that could differ. This is the record that can.

Two counts, never a boolean, and never one union-typed slot. subjects is what the check was evaluated over and findings is what it turned up, so ran against nothing (subjects=0, skipped=None) and did not run (skipped set) can never occupy the same value. vrs_alleles/vrs_alleles_identified is the precedent and the argument is the same one: an unstated denominator is the defect, because coverage of an unknown fraction is not something anything can key on.

skipped is a closed vocabulary with the sentence beside it, not instead of it. Backfill triage branches on why a pass did not run, so prose in that slot would relocate the substring matching RM44 documents rather than end it. detail is where the good sentence goes — clinical.tautology_reason already writes one, and it stays exactly as it is.

VerificationDoc

Bases: BaseModel

The full verification.json beside the spec: an attestation over a list of check records.

Why this is a document and not a fifth fact CSV. The object has two levels — one attestation covering many records — and a CSV can express that only by carrying a non-data service row (the shape RM36 rejected on genome_build, for exactly the reason it applies here: a data table would hold a row that is not data) or by repeating the attestation on every row, where two rows can then disagree about a per-run fact. provenance.json is the standing precedent for the shape that fits: a JSON document beside the spec, read and hashed by the compiler, summarized into a manifest block.

There is a second reason, and the charter names it. The 0.6 amendment observes that a derived table which is both machine-written and human-overridable can be edited into a state that is not merely stale but is a false claim, and that this "wants a mechanism rather than a convention". Every CSV sidecar is overridable on purpose — a curator correcting a row the enricher wrote is the designed path. An attestation is the one derived thing where that must not silently pass, so it is deliberately not in the family whose overridability is a feature.

The mechanism, and its exact modesty. module_hash binds the record to the authored bytes it was computed over, and nonce is a proof-of-work over that binding plus the records' own signature. Both exist to stop an accidental forgery — an attestation left behind by an edit, or copied between modules — and nothing here is built as though the library were hack-resistant, because it is not and does not claim to be. A reader who wants a guarantee wants manifest.signature, which is a real one.

Verification

Bases: BaseModel

What a module can say about whether anything it asserts was ever CHECKED (RM45).

Absent on a module nothing verified, and absent on one whose attestation no longer matches its bytes. Both read correctly as says nothing, which is the only safe default: a block that survived an edit would say a check passed over rows it never saw.

On Frequency's precedent — a separate block rather than fields on Compilation, because this has its own producer, its own releases and its own fact-hash. It departs from that sibling in one way, deliberately: Frequency carries derived facets (sources, datasets, populations) because its rows stay in the sidecar and never reach the manifest, while these records are few and are embedded whole, so a union list here would restate what is already two lines below it. Keep the parts, compute the convenience.

Nothing in this block is trusted. Every field repeats it, because the first consumer to read a pass off an untrusted manifest will otherwise believe it, and a forged pass is worse than silence.

Contribution

Bases: BaseModel

One authorship contribution to this version of a module (RM14; docs/USE_CASES.md §5a).

Three orthogonal axes (Principle 5), unbundling the flat authors/free-form curator: who (identity), role (what they did — closed vocab), and kind (a multi-valued tag set describing the contributor: a human ladder of assurance human → human_expert → human_certified, or ai with a scale tag agent/team/swarm — open, so new tags may be coined). A joint contribution is two entries (a human and an ai), each with its own kind, so the mix is always spelled out and there is no lossy hybrid tag.

Module metadata: carried in the manifest, out of artifact.digest (like provenance/logs), so two versions with identical annotation content but different authorship share a content identity. A consumer (the network validator, a review queue, a human auditor) routes its scrutiny by kind — the format carries the kind, the consumer picks the profile (the data-agnostic north star). extra="forbid" keeps the record's namespace closed.

Weighting

Bases: BaseModel

What a module's authored weight column means, in the author's own words (0.6, RM92).

VariantRow.weight is a bare float | None described only as "Score (positive=protective)". It has no unit column — unlike effect_size, which has effect_measure beside it — so nothing in the artifact says what scale it runs on, how it was arrived at, or whether two modules' weights are on one scale. A consumer reported the consequence: the weights "construct nonsense" across a corpus, because every module means something different by the column and the artifact cannot say so (S36).

Free text on purpose, all three fields. A closed vocabulary would have to enumerate scales nobody has surveyed, and the one-way-door rule (P5) says do not fix a shape on a guess. More pointedly, the typed field this block deliberately does not carry is a precedence rule — something saying "use the GWAS effect where weight is null". That would put two methodologies in one summable column, which is the defect the report is about; the module states what its weights are, and a consumer chooses a table wholesale rather than blending row by row.

Advisory metadata, exactly like license: out of artifact.digest and out of content_signature (the identity/display half of module_spec.yaml is excluded by design), and not reconstructed by the lossy reverse_module, same class as panel/authorship/license.

ModuleManifest

Bases: BaseModel

Full module manifest (SPEC §4). Written next to the parquets as manifest.json.

read_manifest

read_manifest(path: Path) -> ModuleManifest

Load and validate a manifest.json from disk.

Source code in schema/src/just_dna_format/manifest.py
def read_manifest(path: Path) -> ModuleManifest:
    """Load and validate a `manifest.json` from disk."""
    return ModuleManifest.model_validate_json(Path(path).read_text(encoding="utf-8"))

write_manifest

write_manifest(
    manifest: ModuleManifest, path: Path
) -> Path

Write a manifest to disk as indented JSON. Returns the path written.

Source code in schema/src/just_dna_format/manifest.py
def write_manifest(manifest: ModuleManifest, path: Path) -> Path:
    """Write a manifest to disk as indented JSON. Returns the path written."""
    path = Path(path)
    path.write_text(manifest.model_dump_json(indent=2, exclude_none=False) + "\n", encoding="utf-8")
    return path