just_dna_format.integrity¶
just_dna_format.integrity ¶
Integrity primitives (SPEC §5).
All hashes are SHA-256, lowercase hex, prefixed sha256:. These functions are the shared
implementation the compiler uses to emit integrity fields and a downloader uses to verify
them — keeping both sides byte-for-byte agreement by construction.
Time is never read here: callers pass any timestamps into the manifest. This keeps the module pure and deterministic.
IntegrityError ¶
Bases: Exception
Raised when a file hash, artifact digest, or trust check fails verification.
verify_signature ¶
verify_signature(
digest: str,
signature: Signature,
*,
trusted_public_key: str | None = None,
) -> None
Verify a Signature over the artifact.digest string. Raises IntegrityError on failure.
When trusted_public_key (base64 raw) is given, the signature MUST have been made by that key
— this is the real defense (a self-embedded key proves nothing against a backend that can
rewrite both digest and key). When omitted, only self-consistency is checked.
Source code in schema/src/just_dna_format/integrity.py
sha256_bytes ¶
sha256_file ¶
Streaming SHA-256 of a file's raw bytes, prefixed sha256:.
Source code in schema/src/just_dna_format/integrity.py
file_entry ¶
Build a FileEntry (name, sha256, size) for directory/name.
Source code in schema/src/just_dna_format/integrity.py
file_entries ¶
Build FileEntry rows for each existing name under directory (skips missing).
Source code in schema/src/just_dna_format/integrity.py
newline_normalized_file_entry ¶
A FileEntry over directory/name with \r\n read as \n — for the binding only (RM82).
Rewriting an authored CSV with different line endings changes no value, no digest and no
signature, and used to drop the whole verification attestation and the closure with it: an author
whose editor normalizes newlines, or whose Git does it through core.autocrlf, un-closed a module
without touching a cell. This entry builder is what verification.module_binding is computed over,
through just_dna_compiler.compiler.authored_input_entries, so that rewrite is no longer an edit.
Both halves are normalized, and the second one is the whole trap. artifact_digest hashes
{"name", "sha256", "size"} per file, so a builder that normalized the bytes it hashed while
reporting stat().st_size would still move the binding — by one byte per line, on exactly the
files this exists to protect. So size here is the length of the normalized stream, not the
length on disk. That is sound because these entries are only ever fed to module_binding: they are
hashed, never published as a listing, and nothing reads them as a claim about a file's size. The
entries that are published — manifest.inputs[] and artifact.files[] — keep coming from
file_entry/file_entries with the on-disk bytes and the on-disk size.
A separate function rather than a normalize=True flag on file_entries. A flag must mean the
same thing in every function that takes one, and a boolean that silently changes what a hash is
over is the opposite of that: a caller passing it by habit would re-baseline manifest.inputs[]
with no error and no warning. Two functions cannot be confused at a call site.
The stopping point is newlines, and it is chosen rather than inherited. A BOM, trailing
whitespace and a missing final newline are the obvious next steps, and each makes the binding more
content-ish without making it content — this deliberately implements none of them. Newlines are the
one difference a tool introduces on a file the author did not edit; the others are things a human
typed. A lone \r is left exactly as it is, for the same reason: it is not what an editor or Git
writes when it normalizes. If a real case arrives for one of the others it is additive and gets
argued then, on its own evidence.
Source code in schema/src/just_dna_format/integrity.py
newline_normalized_file_entries ¶
newline_normalized_file_entry for each existing name under directory (skips missing).
The sibling of file_entries, with the same skip-missing contract — a module carries only the
table kinds it uses, so an absent name is the ordinary case rather than a failure.
Source code in schema/src/just_dna_format/integrity.py
artifact_digest ¶
Merkle-style root over the file set (SPEC §5): build the JSON array
[{"name","sha256","size"}, ...] sorted by name, serialized with sorted keys and no
whitespace, then hash. Verifying this one digest verifies the whole set, independent of the order
the files were listed in.
This is the version's immutable byte identity — these bytes, from this compiler (Principle
4) — and not its content identity, which is content_signature. The distinction is the whole
reason there are two hashes: a recompile against a different reference moves the digest while the
authored content is untouched, so reading a moved digest as moved content sends a reader hunting a
change that did not happen. This docstring said "content identity" until 2026-08-12; the same
wording was corrected in the docs when a consumer made exactly that misreading (S7), and the code
copy outlived the fix.
Source code in schema/src/just_dna_format/integrity.py
build_artifact ¶
Hash each output file and compute the artifact digest over the set.
Source code in schema/src/just_dna_format/integrity.py
content_signature ¶
content_signature(
tables: Mapping[str, Sequence[BaseModel]],
genome_build: str = DEFAULT_GENOME_BUILD,
) -> str
Stable content identity over the RAW authored data rows — the canonical, owned algorithm.
Distinct from artifact_digest: that hashes the compiled parquet bytes, which are
GRCh38-coordinate-relative and therefore depend on the Ensembl reference the resolver was given
(and move if the module is recompiled elsewhere). content_signature instead hashes the authored
data rows as parsed — so it is:
- Reference-independent — computed from the rows before resolution (an rsid-only row is
hashed as authored, not as resolved coordinates), so recompiling against a different/complete
reference does not change it. This bullet used to say "build-independent", which was true of the
reference used to resolve and false of the declared assembly, and the two are not the same
thing: for a coordinate-authored module
genome_buildis not a resolution artifact, it is the frame the authored numbers are in. HFE C282Y is 6:26,093,141 on GRCh37 and 6:26,092,913 on GRCh38, so two modules with byte-identical CSVs and different declared builds describe loci 228 bp apart — and hashed equal, which for a content-dedup key is the wrong answer. The realistic way to hit it is not contrived: "lift over" a GRCh37 panel by editing the yaml and not the coordinates, and a registry keyed on this would call the result the same content. - Build-aware, by omitting the default —
genome_buildnow feeds the hash, but only when it is notDEFAULT_GENOME_BUILD. That is the same normalization the bullet below already applies to an unset optional column, not an exception to it, and it is what keeps the fix targeted: every GRCh38 module — which is every module published to date — keeps its existing signature byte for byte, and only the modules that were being misidentified change. - Name/metadata-independent — the identity and display half of
module_spec.yaml(name, version, namespace, title, colour) is excluded, so a metadata edit or a registry strip does not change it.genome_buildis the one key from that file that does feed the hash, because it is not metadata about the module: it is part of what the rows mean. - Normalized — each row is
model_dump(mode="json", exclude_none=True), so CSV reformatting (whitespace, quoting, column reorder, cell canonicalization like1.00→1.0) and additive schema growth (a new optional column left unset) do not change it. Allele case is part of that canonicalization since RM215: a column whose grammar is case-insensitive is upper-cased here, becauseALLELE_PATTERNcarriesre.IGNORECASEand the cell is stored verbatim, soA/G,a/G,A/ganda/gare one heterozygote that used to hash four ways. The fold is driven by theCASE_INSENSITIVE_ALLELEmarker and reaches exactly the four columns whose validator is that grammar — notref/alts, which are not grammar-checked at all (a non-nucleotide there is a spelling defect a later pass diagnoses), so there is no case-insensitivity for them to inherit and folding them would collapse values that differ. Like thegenome_buildbullet, this is targeted: every module whose alleles are upper-case — which is every module published to date — keeps its signature byte for byte, and only the modules that were being misidentified move. - Value cells, not the provenance beside them — a column marked
OUTSIDE_CONTENT_IDENTITY(base.content_identity_exclusions) is dropped from the dump here and nowhere else: the overlay'sreason/decided_by/decided_at(S87) say why a correction was made, and two modules differing only there assert the same thing. It is the same exclusionfact_signatureapplies to a derived table'sfetched_at, applied to the one authored table that carries provenance beside its claims. - Deterministically sorted, order-independent — the normalized rows of each file are sorted by
their canonical JSON, and files are sorted by name, so re-ordering rows yields the same
signature. (This is deliberately unlike
artifact.digest, which preserves authored row order: the two are different identities — a byte-reproducibility digest vs. a content-dedup key.)
tables maps each data-CSV filename (variants.csv, studies.csv, and the 0.4 table kinds) to
its parsed, validated rows. The result is the enabling identity for content-level dedup that
survives import/recompile and metadata-strip. It is the reference algorithm a marketplace's
find_versions_by_content should adopt (see docs/proposals/PROPOSAL_0_4_1.md).
Source code in schema/src/just_dna_format/integrity.py
192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 | |
fact_signature ¶
Stable, producer-independent identity over a derived-fact table's fact columns.
The shared body behind resolution_signature, frequency_signature, and
gene_metrics_signature — three tables with one hashing discipline, so the rule cannot drift
between them as more sidecars land. Each row is reduced to its fact_fields with None dropped,
canonicalized, and the sorted set hashed:
- Fact-only — the provenance columns each table excludes (
source/status/fetched_at) are simply not infact_fields, so a human-filled and a machine-filled table with identical facts hash equal. That producer-independence is the whole reason these tables are hashed here instead of enteringmanifest.inputsas raw-byteFileEntrys. - Normalized —
model_dump(mode="json")withNonedropped, so CSV reformatting and an unset optional column do not change it. - Order-independent — rows are sorted by their canonical JSON, so a producer that emits them in a different order still hashes equal.
Source code in schema/src/just_dna_format/integrity.py
frequency_signature ¶
Fact-hash of frequencies.csv (frequency.FREQUENCY_FACT_FIELDS). See fact_signature.
Recorded in manifest.frequency.signature and kept out of artifact.digest: the compiled
frequencies.parquet is already in the digest, and this is the producer-independent identity of
the same content — the thing that stays equal when the enricher and a human write the same numbers
with different column order and different timestamps.
Source code in schema/src/just_dna_format/integrity.py
gene_metrics_signature ¶
Fact-hash of gene_metrics.csv (gene_metrics.GENE_METRICS_FACT_FIELDS).
literature_signature ¶
Fact-hash of literature.csv (literature.LITERATURE_FACT_FIELDS).
The narrowest fact set of the three, and deliberately so: only which article this is is a fact about the module. Open-access status and quote coverage are the outside world's state on the day the pass ran, so they stay outside — otherwise an embargo lifting would move a module's signature with no authored edit anywhere.
Source code in schema/src/just_dna_format/integrity.py
gene_validity_signature ¶
Fact-hash of gene_validity.csv (gene_validity.GENE_VALIDITY_FACT_FIELDS), 0.6 / RM24.
Wider than its siblings, because the row's identity is wide: a gene–disease assertion is scoped by mode of inheritance and, on an aggregate like GenCC, by who made it.
Two non-provenance columns are excluded, on one rule — a column that locates or describes an
assertion is not the assertion. report_url locates the curation and moves when a site
reorganizes; disease_label is the ontology's wording for a term disease_id already names, and
it churns on its own (one real export carries MONDO:0017146 under two labels at once). The stable
identities — disease_id and assertion_id — are inside.
Source code in schema/src/just_dna_format/integrity.py
clinical_assertion_signature ¶
Fact-hash of clinical_assertions.csv (assertions.CLINICAL_ASSERTION_FACT_FIELDS), RM25.
review_status and review_stars are both inside it although one determines the other. The
mapping between them is a ClinVar convention that Principle 2 keeps out of this tier, so from
here they are two independent inputs — and a table whose stars disagreed with its own prose
should hash differently from one where they agree.
rsid is outside, unlike FREQUENCY_FACT_FIELDS, and the difference is where the value comes
from: gnomAD reports an rsID in its own payload, while the ClinVar lookup is allele-exact and
returns none, so this column is filled from the module's own resolution.csv. Inside the hash it
would make two modules holding the same archive records hash differently on whether their resolver
attached an rsID — the producer-dependence a fact hash exists to exclude.
Source code in schema/src/just_dna_format/integrity.py
gwas_effect_signature ¶
Fact-hash of gwas_effects.csv (gwas.GWAS_FACT_FIELDS), 0.6 / RM90.
effect_unit is inside, and it is the column that most needs to be: the same magnitude in
umol/l and in the Catalog's uninformative unit are different facts, and a table that hashed
equal across them would have thrown away the distinction this table exists to keep.
trait is outside, on gene_validity_signature's rule — a column that describes an
assertion is not the assertion. The Catalog re-words a reported trait between releases for an
unchanged trait_efo_id, so the label churns while the fact does not.
rsid is inside, which inverts clinical_assertion_signature and does so deliberately. There
the rsID is filled from the module's own resolution.csv because the archive returns none; here
the Catalog is queried by rsID and echoes it back inside riskAlleleName, so it is part of what
the source said rather than of what the module knew.
Source code in schema/src/just_dna_format/integrity.py
expression_effect_signature ¶
Fact-hash of expression_effects.csv (expression.EXPRESSION_FACT_FIELDS), 0.7 / RM194+RM200.
gene and gene_id are both inside, which looks like one column twice and is not. The HGNC
symbol is renamed between releases while the Ensembl accession is not, so hashing only the symbol
would move this signature on a rename that changed no claim, and hashing only the accession would
leave the column every authored gene joins against outside the fact.
tracks_total is inside beside tracks_agreeing, on the rule that a cell key carries the
value when two rows may state two claims: 40 of 371 and 40 of 512 are different facts, and a
model shipped with a different track panel must not hash equal to the one that produced these.
distance_to_gene is inside, because it is a property of the (variant, gene) pair rather
than of the variant — the same variant scored against a different gene sits a different distance
away, and the whole reason RM194 records it is that distal scores run an order of magnitude
lower.
The three provenance columns are outside, as everywhere else, so a hand-curated table and a pass-filled one carrying the same claims hash equal.
Source code in schema/src/just_dna_format/integrity.py
clin_sig_concordance_signature ¶
Fact-hash of clin_sig_concordance.csv (concordance.CLIN_SIG_CONCORDANCE_FACT_FIELDS), RM130.
opposed is inside, and it is the column the fact set would be wrong without: the two verdict
columns say the authorities disagree and where the module sits, and neither says whether the
disagreement crosses the pathogenic/benign line. A module whose contested subjects are opposed
asserts something different from one whose subjects merely differ, so the two must not hash equal.
checked_at is outside, on the rule every sibling applies to fetched_at: a re-run that put
the same questions and got the same answers has changed nothing anybody asserted.
Source code in schema/src/just_dna_format/integrity.py
clin_sig_authority_call_signature ¶
Fact-hash of clin_sig_authority_calls.csv
(concordance.CLIN_SIG_AUTHORITY_CALL_FACT_FIELDS), RM130.
confidence and confidence_unit are both inside, together, for gwas_effect_signature's own
reason one table over: an unconverted magnitude and the instrument it was measured on are one
fact in two cells, and hashing the number without the unit would make two authorities' different
scales collide.
status is inside too. no_record and unchecked are different statements about the same
subject — one archive was asked and has nothing, the other could not be asked — and a record that
swapped them says something else about how much of the check actually ran.
Source code in schema/src/just_dna_format/integrity.py
source_signature ¶
Fact-hash of sources.csv (sources.SOURCE_FACT_FIELDS).
Note the one inversion against its siblings: source is inside this fact set, because here it
is the subject of the row rather than the provenance of one. Only fetched_at is excluded, so
re-reading the same terms at a different moment hashes equal, while a changed licence, a changed
permission flag or a changed declaration all move the signature — which is the point.
Source code in schema/src/just_dna_format/integrity.py
resolution_signature ¶
Stable, producer-independent identity over the resolution table's facts (0.5).
resolution.csv is a multi-producer artifact — the enricher, a human, and reverse_module all
write it, with byte-different but fact-identical output (column/row order, provenance columns,
timestamps). So it is hashed HERE, over the fact columns only, and is deliberately NOT added to
manifest.inputs: a raw-bytes FileEntry hash would be unstable across those producers and would
make a reverse → recompile cycle "change the hash" for no real reason. Mirrors
content_signature, restricted to resolution.RESOLUTION_FACT_FIELDS:
- Fact-only — the provenance columns (
source/status/fetched_at) are excluded, so a human-filled and an Ensembl-filled table carrying the same facts hash equal. - Normalized — each row is
model_dump(mode="json")restricted to the fact fields withNonedropped, so CSV reformatting and an unset optional column do not change it. - Deterministically sorted, order-independent — rows are sorted by their canonical JSON.
Together with content_signature and compiler_version it fully determines artifact.digest, so
a holder of the two small CSVs reproduces the artifact byte-for-byte, fully offline.
Source code in schema/src/just_dna_format/integrity.py
verify_manifest ¶
verify_manifest(
module_dir: Path,
manifest: ModuleManifest,
*,
require_marketplace: bool = True,
check_inputs: bool = False,
check_logs: bool = False,
check_provenance: bool = False,
check_logo: bool = False,
check_readme: bool = False,
check_derived: bool = False,
public_key: str | None = None,
) -> None
Verify a downloaded module against its manifest (SPEC §5 verify-then-install).
require_marketplace is a policy switch and its default is the registry one (S34). Step 3
demands compiled_by == "marketplace-server", so with the default a consumer that wires exactly
one call site rejects every locally-compiled module — including one this project's own
compiler produced, which leaves compiled_by null by design. That is correct behaviour and a
surprising default to meet through a parameter named for a requirement rather than for a policy.
A consumer that installs from both places needs two call sites, not one:
require_marketplace=True— a registry install. The bytes came from a party you are trusting to have compiled them, so "who compiled this" is part of what you are checking.require_marketplace=False— a local compile, a sideloaded directory, a file a user handed you. The hashes and the digest are still checked in full; only the provenance claim is dropped, because there is no registry to have made it.
Nothing weaker is on offer and nothing here is a trust root: compiled_by is an unsigned string in
a file the same party wrote. The real guarantee is public_key below — a detached Ed25519
signature over artifact.digest by a key the client pins.
Steps
- Every
artifact.files[]present on disk hashes to its declared value. - The recomputed
artifact.digestmatches the manifest. compile_successis true andcompiled_by == "marketplace-server"(whenrequire_marketplace).- Optionally (
check_inputs) everyinputs[]file on disk matches its declared hash. - Optionally (
check_logs) everylogs[]file present on disk matches its declared hash; absent logs are skipped, since logs are optional and need not be downloaded. - Optionally (
check_provenance) theprovenancedocument, if declared and present on disk, matches its declared hash; an absent provenance file is skipped (it is optional). 6b. Optionally (check_logo) thelogo, if declared and present on disk, matches its declared hash; an absent logo is skipped (it is optional and out ofartifact.digest). 5b. Optionally (check_derived) everyderived[]sidecar CSV present on disk matches its declared byte hash; an absent one is skipped, exactly as for logs — these live beside the spec rather than in the module dir, so a consumer who fetched only the artifact has none of them, which is not a failure. This checks bytes in transit; it is not the tables' identity (that is the fact signatures, which survive a rewrite this check would flag). 6c. Optionally (check_readme) thereadme, on the same terms as the logo. This is the check that makes a served readme verifiable: a registry serving a file whose hash nothing records is serving something nobody can check, which is why the field exists rather than the bytes merely sitting on disk. - Signature (SPEC §5): if
public_key(base64 raw) is given, the manifest MUST carry a signature overartifact.digestmade by that key. If a signature is present but no key is pinned, it is verified for self-consistency only.
Raises IntegrityError on the first failure; returns None on success.
Source code in schema/src/just_dna_format/integrity.py
478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 | |