just_dna_compiler.compiler¶
just_dna_compiler.compiler ¶
Module spec compiler: validates a spec directory and compiles it to a composed multi-parquet
artifact plus a manifest.json. A module composes from optional table kinds (RM2): the three-parquet
SNP core (weights, annotations, studies) when it carries variants, plus one parquet per 0.4 table
kind it includes (diplotypes, pharm_variants, pgs, the binning kinds, …). ARTIFACT_PARQUETS is the
roster and len(ARTIFACT_PARQUETS) the count — stated that way rather than spelled, because this
sentence carried a spelled-out figure for two releases while the tuple grew past it (RM218).
Public API
validate_spec(spec_dir) -> ValidationResult compile_module(spec_dir, output_dir, ...) -> CompilationResult (emits manifest.json) reverse_module(parquet_dir, output_dir, ...) -> Path
The DSL/manifest schema comes from just-dna-format; this package is the transform between them.
SpecError ¶
Bases: ValueError
A module_spec.yaml that could not be loaded — see load_spec.
A ValueError subclass so a caller that already brackets a load with except (OSError,
ValueError), the way read_verification's callers do, keeps working without knowing this type
exists.
table_bindings ¶
csv -> the parquet(s) it becomes, assembled from the registries and never named by hand.
Public because two surfaces outside this module need it and both would otherwise keep a copy: the
docs site's generated per-table reference, and test_artifact_parquet_bindings.py, which asserts the
union of the values is exactly ARTIFACT_PARQUETS. That equality is what makes the map total — a
new table kind whose parquet is written by a literal at its own write site fails the test instead of
quietly missing a reference page.
Four registries, each owning its own slice: SNP_CORE_PARQUETS (the core, where variants.csv
fans out to two), _TABLE_KINDS (the optional authored kinds), _FACT_TABLES (the derived facts)
and OVERRIDES_PARQUET (the overlay, whose CSV is authored but is not a table kind).
Source code in compiler/src/just_dna_compiler/compiler.py
authored_input_entries ¶
The authored files a module is made of, hashed for the verification binding.
Public because two tiers must agree on it byte for byte (the RM41 lesson): the compiler
recomputes the binding from this set when it decides whether to publish manifest.verification,
and the enricher hashes the identical set into the attestation's module_hash so a later compile
can tell whether the spec has been edited since the checks ran. A private symbol here would leave
the enricher choosing between reaching into a private name and re-implementing the list, and a
re-implementation is a place for the two to drift — at which point every attestation this
workspace writes reads as stale to its own compiler.
Authored files only, which is the boundary just_dna_format.verification argues at length:
the derived sidecars carry per-run noise (fetched_at) that would invalidate an attestation on a
re-enrichment that changed nothing anyone claimed.
The bytes are newline-normalized, and manifest.inputs[] is deliberately not (RM82). This
docstring said the compiler hashes these into manifest.inputs; it never did — that field is
filled independently, by file_entries(spec_dir, _INPUT_FILES) over the raw bytes, and the two
were only ever equal by coincidence of computing the same thing. Since 0.6 they are two different
questions asked of one file set. manifest.inputs[] and artifact.digest answer are these the
exact bytes, so they follow every byte, line endings included. The binding answers is this still
the module those checks were put against, and an editor rewriting \r\n as \n changes no
value, so it must not un-close a module. The asymmetry is the decision, not an inconsistency to
tidy: see integrity.newline_normalized_file_entry for the transform and where it stops.
Source code in compiler/src/just_dna_compiler/compiler.py
load_spec ¶
Load and validate a module_spec.yaml, raising on anything wrong (S74).
The public route to a ModuleSpecConfig. The model has always been exported from
just_dna_format.spec and the only thing that produced one was _load_yaml, underscored — so a
consumer wanting weighting: or authorship: had to yaml.safe_load the file and read a raw
dict, losing the authority-key handling and every diagnosis below, and carrying PyYAML for no
reason except that ours was unreachable. Sibling of just_dna_format.read_manifest and
read_verification: same shape, same contract, one per file a module carries.
It lives here rather than in the format tier because loading it needs pyyaml, and the format
tier is pydantic + cryptography by charter. A consumer already depending on the compiler —
which every caller of validate_spec is — can drop their own PyYAML with this.
authority_keys is inject-only and unchanged: pass
just_dna_format.normalize.IDENTITY_AUTHORITY_KEYS so a registry-stamped
namespace:/owner:/canonical_id: is stripped before validation rather than tripping
extra="forbid". The format applies none by default, and a key outside the injected set still
trips. Which keys were dropped is not reported here — that is validate_spec's .info, and a
caller who needs it wants the validator rather than the loader.
Raises SpecError with every diagnosis joined, where _load_yaml returns them for accumulation.
That difference is the whole reason both exist: validate_spec collects errors from a dozen
sources and reports them together, which is right for a validator and wrong for a loader — a
caller who just wants the object should not have to check a tuple's second element to find out
it got None.
Source code in compiler/src/just_dna_compiler/compiler.py
load_csv_rows ¶
load_csv_rows(
path: Path,
row_model: type,
file_label: str,
genome_build: str = DEFAULT_GENOME_BUILD,
) -> tuple[list[Any], list[str], list[str]]
Load a CSV and validate each row against a Pydantic model. Returns (rows, errors, warnings).
Public as of 0.5.1 (RM41), and it was public in practice long before. This is the only correct
way to turn an authored CSV into row models, just-dna-enricher consumes it across a package
boundary in a dozen places, and a downstream consumer wiring the pipeline server-side had the
choice of reaching for a private symbol or re-implementing it. Re-implementing is a trap rather
than a chore, because it is not csv.DictReader plus Model(**row) — it carries the two rules
below, each of which this workspace has already had to fix once:
- an empty cell becomes
None, and the key is kept.MeasureBinRow.measure_kindhas a default, sois_required()isFalse, but the model then receivesNonerather than its default and fails on type. A""where this would have putNoneis a different failure again. genome_buildis told to each row (below), so a loader that omits it mints GRCh38 identities for a GRCh37 module.
_load_csv_rows remains as an alias so no caller breaks.
genome_build is told to each row, not read from it. A coordinate is not absolute, so a row
deriving an identity from one needs the module's assembly — and a pydantic model built from a CSV
dict has no module_spec.yaml in scope. Injecting it here keeps the build declared exactly once
(the yaml) while reaching every row: it is a private attribute, so it is not a column, reaches no
parquet, and moves no digest. See AuthoredModel._genome_build for why per-row or per-CSV
declaration was rejected. Callers that load a build-independent table — the resolution and fact
sidecars, which carry their own genome_build column — leave it at the default and are unaffected.
Source code in compiler/src/just_dna_compiler/compiler.py
load_spec_variants ¶
A spec directory's variants.csv, loaded and re-stamped for the build the module declares.
The other half of RM41. Two enricher checks take rows rather than a spec_dir —
acmg.verify_acmg_sf and identifiers.check_identifiers — unlike every other pass, so a caller
has to do this itself, and doing it right means three steps rather than one: read the declared
build out of module_spec.yaml, inject it into every row, and then re-stamp the identities, since
VariantRow._freeze_identity runs at construction where the yaml is not in scope.
Missing or unreadable yaml falls back to DEFAULT_GENOME_BUILD, matching what compiling that
directory would assume — this is a read-only check helper, not the enrichment path, which refuses
rather than choose a build for a module whose declaration cannot be read (it writes facts back).
Returns (variants, errors, warnings); an absent variants.csv is an empty list and one error,
exactly as load_csv_rows reports it.
Source code in compiler/src/just_dna_compiler/compiler.py
positional_placement ¶
(rows, placed) across the three positional table kinds — the counts the manifest publishes.
The structured half of _check_positional_joinability, which reports the same facts as prose per
table (S31). Public because the number is what a catalog wants: until 0.6 the only record of how
much of a PGx table joins to a VCF was UNJOINABLE_PHRASE inside manifest.compilation.warnings,
and a downstream registry substring-matched it. Anything a consumer can only learn from a warning
string is an unversioned interface (RM44).
Call it after _apply_positional_resolution, so placed counts what the artifact actually
carries rather than what the author typed — the fill is where an rsid-authored PGx module gets its
coordinates, and before it the answer is the pre-RM43 one.
A module carrying no positional table returns (0, 0), which is a real answer and distinct from
the None the manifest holds for a compile that never counted; see Compilation.positional_rows.
Source code in compiler/src/just_dna_compiler/compiler.py
load_citing_rows ¶
Every citing table present beside a spec, keyed by CSV name — the annotation rows that
ground their own claim with a pmid (MeasureBinRow.pmid RM47, PharmVariantRow.pmid RM132).
Public because a second tier needs it: the enricher's literature pass has to check these pointers
alongside studies.csv, and its two alternatives were importing a private symbol or hand-keeping a
parallel list of the citing kinds — the RM40/RM41 shape exactly, and the list would go stale on the
next kind that declares the column.
Supersedes load_binning_rows, which stays and still means what it always did: it reads the
binning kinds only, so a caller wanting the citations a module makes wants this one.
Source code in compiler/src/just_dna_compiler/compiler.py
load_binning_rows ¶
Every binning table present beside a spec, keyed by CSV name (RM47).
Narrower than load_citing_rows since 0.7 and deliberately kept: a caller asking for the binning
kinds is asking about thresholds, not about citations, and quietly widening what it returns would
hand such a caller PharmVariantRows where it expects MeasureBinRows.
Source code in compiler/src/just_dna_compiler/compiler.py
table_citations ¶
Digit-only PMIDs the module's annotation tables cite, de-duplicated — every citing kind
(MeasureBinRow.pmid RM47, PharmVariantRow.pmid RM132).
Takes the whole table-kind map a caller already holds and reads only the citing kinds out of it, so
a caller cannot accidentally hand over a haplotypes.csv (no pmid column) and get an attribute
error. The kind set is derived from the models, never hand-listed, so a kind that gains the column
is read here without an edit.
First-occurrence order rather than sorted, because it feeds emission order downstream (P7), and
normalization goes through extract_pmids so a table pointer and studies.csv cannot drift into
two spellings of one citation.
Source code in compiler/src/just_dna_compiler/compiler.py
binning_citations ¶
Digit-only PMIDs the binning tables cite (MeasureBinRow.pmid), de-duplicated.
Narrowed by table_citations since 0.7 the same way load_binning_rows is by load_citing_rows,
and kept for the same reason. Nothing inside the compiler calls it any more: the literature
cross-check must read every citation site, and this one answers a question about thresholds.
Source code in compiler/src/just_dna_compiler/compiler.py
load_overlay ¶
Read overrides.csv and answer with (rows, errors, warnings) (RM124).
Public since RM136, because the enricher needs it: an author who corrects a derived cell
through the overlay must not go on being told the same finding by the tier that writes the file.
Private, it would have been reached into the way load_spec was before S74, or — worse —
reimplemented in the enricher, which is the drift the overlay's own design refuses.
One loader for both public entry points, because both have to read it and a second copy is where
validate and compile learn to disagree — the parity rule this module keeps re-learning
(@validate-refuses-all). Everything it reports is structural: the rows parse or they do not, the
key is duplicated or it is not, a key group carries one operation or several. Nothing here
consults a derived table, which is what keeps every finding identical on both laps of a round
trip.
A file that is present with no rows is an error, the same answer _TABLE_KINDS gives: an
empty authored table is a header somebody meant to fill.
Source code in compiler/src/just_dna_compiler/compiler.py
validate_spec ¶
validate_spec(
spec_dir: Path,
authority_keys: Iterable[str] | None = None,
*,
strict: bool = False,
resolve_with_ensembl: bool = True,
) -> ValidationResult
Validate a module spec directory without producing output.
authority_keys (inject-only) is the set of consumer/registry-owned identity keys to strip from
the authored module: block before validation — pass just_dna_format.normalize.
IDENTITY_AUTHORITY_KEYS (or your own set) so a legacy spec carrying namespace:/owner:/
canonical_id: validates; the format applies none by default. Stripped keys are surfaced on
.info. Everything else still trips extra="forbid".
strict mirrors compile_module's flag and exists for one reason: several checks are a mode
ladder (warning in best_effort, error in strict), so without a mode here the pre-flight
could not answer the question the author actually asked — the documented order is validate then
compile --strict, and a modeless validate is a pre-flight for the other compile.
It mirrors compile_module's severities exactly, which is not the same as changing severity
only — this said the latter, in both of the two docstrings carrying it, and it was false
(RM218). Two findings are aggregates with no best_effort counterpart sentence: the
unresolved-position refusal (strict compile: N variant(s) …) and build_disagreement_error.
Their best_effort rung is a different sentence — the per-subject rsid_unresolved warning,
which fires in both modes — so under strict the aggregate is genuinely added rather than
promoted. The contract the two commands share is that validate(strict=x) and
compile(strict=x) reach the same verdict, not that the two modes of validate differ by a
severity column.
resolve_with_ensembl mirrors it for the same reason and is passed through by compile_module.
The pre-flight applies the injected table to the positional 0.4 tables (RM43), and that decides
whether their rows are reported as unjoinable — so a validate that ignored the master resolution
switch would be more optimistic than the compile it precedes, which is the disagreement
direction the parity rule exists to prevent.
Stats include genes/categories as lists (filtering None) plus variant_count,
gene_count, study_count, and the ClinVar quality counts
(clinvar_count/pathogenic_count/benign_count) — the fields the manifest needs. See
ValidationResult.stats for the full key contract.
warnings is the complete list, unchanged; carried names the subset an author cannot clear and
warnings_summary counts them by code (RM131). A caller that wants to keep building on this
run's findings — as compile_module and close_module do — calls _validate_spec instead, for
the reason given there.
Source code in compiler/src/just_dna_compiler/compiler.py
variant_stats ¶
The variants.csv-derived facets of ValidationResult.stats / manifest.stats.
Its own function because it now has two callers, and the second is why: compile_module may
discard a row for carrying an unusable symbolic allele (RM5), and the stats were computed by
validate_spec before that happened. weights_rows counts the parquet and so is post-drop, so a
published manifest claimed a variant_count one higher than the artifact contained — the RM44
class of defect exactly, a manifest number a catalog keys on and cannot check.
Source code in compiler/src/just_dna_compiler/compiler.py
module_stats ¶
module_stats(
variants: list[VariantRow],
kind_rows: dict[str, list[Any]] | None = None,
) -> dict[str, Any]
variant_stats plus the gene facets taken over every authored table, not just variants.
PUBLIC, and it exists rather than a second parameter on variant_stats because that function's
name is a promise about which table it reads and renaming it would be a major (S14's rule). What
the two return differs in exactly two keys.
stats describes the module, and Stats has always said so — "card/detail stats derived from
the spec", not from one table of it. variant_stats nevertheless derived genes from
variants.csv alone, so a module whose lead table is diplotypes.csv, allele_function.csv,
copynumbers.csv or any other gene-bearing kind published gene_count: 0, genes: [] however many
of its rows named a gene — and a registry's gene index is fed from that field, so the module was
unreachable by a gene search (S57). Measured on reference_examples/cyp2c19_star_alleles/: 1,332
rows carrying gene=CYP2C19 across three tables, and genes: [].
The honest workaround an author was left with was prose in the README, and the dishonest one —
inventing an empty variants.csv to be discoverable — is what makes this ours to fix rather than a
documentation note.
Only authored kinds count. _GENE_BEARING_TABLE_KINDS derives from _TABLE_KINDS, which
deliberately excludes the derived fact sidecars, so a gene reaching gene_metrics.csv because a
pass looked it up never becomes a gene the module claims to be about.
Source code in compiler/src/just_dna_compiler/compiler.py
spec_tables ¶
The parsed, defaults-folded authored rows content_signature hashes, and the declared build.
PUBLIC, and the reason is that everything finer than a whole-module hash needs these rows and had
no way to get them (S53). content_signature returned only the digest, so a tool answering what
moved between two versions of this module — per table, per row — had to rebuild the mapping
outside, and rebuilding it meant restating two private things: the table roster (_TABLE_KINDS)
and the defaults: fold (_resolve_spec_defaults, _DEFAULTED_VARIANT_FIELDS).
The fold is the part that silently produces a wrong answer, which is why this returns the
finished mapping rather than exporting the pieces. A caller hashing load_csv_rows output directly
gets a different digest from content_signature for the same module: measured on
reference_examples/hfe_hemochromatosis, writing one curator value on every variant row in one
copy and the identical value under defaults: in another, content_signature agrees across the
pair (correct — RM37) while the raw-rows build disagrees, so a per-table comparison built the
obvious way reports twelve changed rows where there are none. Exporting _TABLE_KINDS and
_resolve_spec_defaults separately would hand out three pieces that must be assembled in one
order — load with the declared build injected, fold, then hash — and the order is the easy half to
get wrong. One function that returns the finished mapping cannot be assembled wrongly.
The roster is authored tables only, so the licensing table is outside it: sources.csv /
licensing.csv is hashed by integrity.source_signature instead, and neither renaming it nor
editing a cell in it moves content_signature. Both verified on the same example. That is correct
— the licence layer is its own identity — and it is stated here because it is the one authored,
hand-editable table a licence audit sends an author looking for.
Raises ValueError if a present data CSV fails validation, exactly as content_signature does:
the contract carries over unchanged, because that function is now this one plus the hash.
Source code in compiler/src/just_dna_compiler/compiler.py
4761 4762 4763 4764 4765 4766 4767 4768 4769 4770 4771 4772 4773 4774 4775 4776 4777 4778 4779 4780 4781 4782 4783 4784 4785 4786 4787 4788 4789 4790 4791 4792 4793 4794 4795 4796 4797 4798 4799 4800 4801 4802 4803 4804 4805 4806 4807 4808 4809 4810 4811 4812 4813 4814 4815 4816 4817 4818 4819 4820 4821 4822 | |
content_signature ¶
Stable content identity over the raw authored data CSVs — name- and Ensembl-independent.
Reads variants.csv, studies.csv, and any present 0.4 table CSVs, validates each row, and
hashes the normalized + deterministically-sorted rows via
just_dna_format.integrity.content_signature. The data is read as authored (no Ensembl
resolution, no parquet build), so this is cheap and reference-independent — a client can compute
it without recompiling and dedup against a registry, surviving both metadata-strip and a recompile
against a different reference. Raises ValueError if a present data CSV fails validation.
"As authored" means the rows, not the spelling: module_spec.yaml's defaults: block is
folded into each variant row first (_resolve_spec_defaults, RM37), because a value written once
under defaults: and the same value written on every row are the same content.
This is spec_tables plus the hash and nothing else, so a consumer wanting the rows behind the
digest — per-table or per-row work — calls that instead of restating the roster and the fold (S53).
Source code in compiler/src/just_dna_compiler/compiler.py
compile_module ¶
compile_module(
spec_dir: Path,
output_dir: Path,
compression: str = "zstd",
resolve_with_ensembl: bool = True,
ensembl_cache: Path | None = None,
compiled_by: str | None = None,
ensembl_reference: str | None = None,
log_files: list[Path] | None = None,
provenance_file: Path | None = None,
logo_file: Path | None = None,
readme_file: Path | None = None,
authority_keys: Iterable[str] | None = None,
strict: bool = False,
ba1_threshold: float = BA1_ALLELE_FREQUENCY_THRESHOLD,
) -> CompilationResult
Compile a module spec directory into parquet files plus a manifest.json.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
spec_dir
|
Path
|
Path to the module spec directory. |
required |
output_dir
|
Path
|
Directory for output parquet files + manifest.json. |
required |
compression
|
str
|
Parquet compression codec. |
'zstd'
|
resolve_with_ensembl
|
bool
|
Master switch for resolution — of every kind, despite the name.
With a |
True
|
ensembl_cache
|
Path | None
|
Deprecated (removed at 1.0). Path to a prebuilt Ensembl DuckDB or parquet
cache dir. In-compiler DuckDB resolution has moved to |
None
|
compiled_by
|
str | None
|
Provenance tag for the manifest (the marketplace passes "marketplace-server"; a local compile leaves it None, so downloaders treat it as untrusted). |
None
|
ensembl_reference
|
str | None
|
Pinned reference id recorded in the manifest for reproducibility. |
None
|
log_files
|
list[Path] | None
|
Explicit run/provenance log files to record. If None, auto-discovers a top-level
|
None
|
provenance_file
|
Path | None
|
Explicit structured-provenance document. If None, auto-discovers
|
None
|
logo_file
|
Path | None
|
Explicit module logo image. If None, auto-discovers |
None
|
readme_file
|
Path | None
|
Explicit module readme. If None, auto-discovers the first of
|
None
|
authority_keys
|
Iterable[str] | None
|
Inject-only set of consumer/registry-owned identity keys to strip from the
authored |
None
|
strict
|
bool
|
All-or-nothing compile. When True, fail (rather than emit a partial artifact) if any
variant still lacks a resolved genomic position ( |
False
|
ba1_threshold
|
float
|
Allele frequency above which a |
BA1_ALLELE_FREQUENCY_THRESHOLD
|
Source code in compiler/src/just_dna_compiler/compiler.py
4845 4846 4847 4848 4849 4850 4851 4852 4853 4854 4855 4856 4857 4858 4859 4860 4861 4862 4863 4864 4865 4866 4867 4868 4869 4870 4871 4872 4873 4874 4875 4876 4877 4878 4879 4880 4881 4882 4883 4884 4885 4886 4887 4888 4889 4890 4891 4892 4893 4894 4895 4896 4897 4898 4899 4900 4901 4902 4903 4904 4905 4906 4907 4908 4909 4910 4911 4912 4913 4914 4915 4916 4917 4918 4919 4920 4921 4922 4923 4924 4925 4926 4927 4928 4929 4930 4931 4932 4933 4934 4935 4936 4937 4938 4939 4940 4941 4942 4943 4944 4945 4946 4947 4948 4949 4950 4951 4952 4953 4954 4955 4956 4957 4958 4959 4960 4961 4962 4963 4964 4965 4966 4967 4968 4969 4970 4971 4972 4973 4974 4975 4976 4977 4978 4979 4980 4981 4982 4983 4984 4985 4986 4987 4988 4989 4990 4991 4992 4993 4994 4995 4996 4997 4998 4999 5000 5001 5002 5003 5004 5005 5006 5007 5008 5009 5010 5011 5012 5013 5014 5015 5016 5017 5018 5019 5020 5021 5022 5023 5024 5025 5026 5027 5028 5029 5030 5031 5032 5033 5034 5035 5036 5037 5038 5039 5040 5041 5042 5043 5044 5045 5046 5047 5048 5049 5050 5051 5052 5053 5054 5055 5056 5057 5058 5059 5060 5061 5062 5063 5064 5065 5066 5067 5068 5069 5070 5071 5072 5073 5074 5075 5076 5077 5078 5079 5080 5081 5082 5083 5084 5085 5086 5087 5088 5089 5090 5091 5092 5093 5094 5095 5096 5097 5098 5099 5100 5101 5102 5103 5104 5105 5106 5107 5108 5109 5110 5111 5112 5113 5114 5115 5116 5117 5118 5119 5120 5121 5122 5123 5124 5125 5126 5127 5128 5129 5130 5131 5132 5133 5134 5135 5136 5137 5138 5139 5140 5141 5142 5143 5144 5145 5146 5147 5148 5149 5150 5151 5152 5153 5154 5155 5156 5157 5158 5159 5160 5161 5162 5163 5164 5165 5166 5167 5168 5169 5170 5171 5172 5173 5174 5175 5176 5177 5178 5179 5180 5181 5182 5183 5184 5185 5186 5187 5188 5189 5190 5191 5192 5193 5194 5195 5196 5197 5198 5199 5200 5201 5202 5203 5204 5205 5206 5207 5208 5209 5210 5211 5212 5213 5214 5215 5216 5217 5218 5219 5220 5221 5222 5223 5224 5225 5226 5227 5228 5229 5230 5231 5232 5233 5234 5235 5236 5237 5238 5239 5240 5241 5242 5243 5244 5245 5246 5247 5248 5249 5250 5251 5252 5253 5254 5255 5256 5257 5258 5259 5260 5261 5262 5263 5264 5265 5266 5267 5268 5269 5270 5271 5272 5273 5274 5275 5276 5277 5278 5279 5280 5281 5282 5283 5284 5285 5286 5287 5288 5289 5290 5291 5292 5293 5294 5295 5296 5297 5298 5299 5300 5301 5302 5303 5304 5305 5306 5307 5308 5309 5310 5311 5312 5313 5314 5315 5316 5317 5318 5319 5320 5321 5322 5323 5324 5325 5326 5327 5328 5329 5330 5331 5332 5333 5334 5335 5336 5337 5338 5339 5340 5341 5342 5343 5344 5345 5346 5347 5348 5349 5350 5351 5352 5353 5354 5355 5356 5357 5358 5359 5360 5361 5362 5363 5364 5365 5366 5367 5368 5369 5370 5371 5372 5373 5374 5375 5376 5377 5378 5379 5380 5381 5382 5383 5384 5385 5386 5387 5388 5389 5390 5391 5392 5393 5394 5395 5396 5397 5398 5399 5400 5401 5402 5403 5404 5405 5406 5407 5408 5409 5410 5411 5412 5413 5414 5415 5416 5417 5418 5419 5420 5421 5422 5423 5424 5425 5426 5427 5428 5429 5430 5431 5432 5433 5434 5435 5436 5437 5438 5439 5440 5441 5442 5443 5444 5445 5446 5447 5448 5449 5450 5451 5452 5453 5454 5455 5456 5457 5458 5459 5460 5461 5462 5463 5464 5465 5466 5467 5468 5469 5470 5471 5472 5473 5474 5475 5476 5477 5478 5479 5480 5481 5482 5483 5484 5485 5486 5487 5488 5489 5490 5491 5492 5493 5494 5495 5496 5497 5498 5499 5500 5501 5502 5503 5504 5505 5506 5507 5508 5509 5510 5511 5512 5513 5514 5515 5516 5517 5518 5519 5520 5521 5522 5523 5524 5525 5526 5527 5528 5529 5530 5531 5532 5533 5534 5535 5536 5537 5538 5539 5540 5541 5542 5543 5544 5545 5546 5547 5548 5549 5550 5551 5552 5553 5554 5555 5556 5557 5558 5559 5560 5561 5562 5563 5564 5565 5566 5567 5568 5569 5570 5571 5572 5573 5574 5575 5576 5577 5578 5579 5580 5581 5582 5583 5584 5585 5586 5587 5588 5589 5590 5591 5592 5593 5594 5595 5596 5597 5598 5599 5600 5601 5602 5603 5604 5605 5606 5607 5608 5609 5610 5611 5612 5613 5614 5615 5616 5617 5618 5619 5620 5621 5622 5623 5624 5625 5626 5627 5628 5629 5630 5631 5632 5633 5634 5635 5636 5637 5638 5639 5640 5641 5642 5643 5644 5645 5646 5647 5648 5649 5650 5651 5652 5653 5654 5655 5656 5657 5658 5659 5660 5661 5662 5663 5664 5665 5666 5667 5668 5669 5670 5671 5672 5673 5674 5675 5676 5677 5678 5679 5680 5681 5682 5683 5684 5685 5686 5687 5688 5689 5690 5691 5692 5693 5694 5695 5696 5697 5698 5699 5700 5701 5702 5703 5704 5705 5706 5707 5708 5709 5710 5711 5712 5713 5714 5715 5716 5717 5718 5719 5720 5721 5722 5723 5724 5725 5726 5727 | |
close_module ¶
close_module(
spec_dir: Path,
*,
closed_by: str | None = None,
private_key_pem: bytes | None = None,
now: str | None = None,
difficulty: int | None = None,
) -> ClosureResult
Declare a module's authoring phase finished, binding the statement to its authored bytes (RM73).
A flat CSV row records nothing about how it came to be, so authoring had no end and every check
that needed one guessed. This is the end: a closure block inside the module's verification.json
naming the hash of the authored files as they stand. A later edit moves that hash, the compiler
recomputes it, and the closure is dropped along with the rest of the attestation — which is why
there is no second file and no second binding here to keep in step.
Deliberate, never a side effect. validate_spec stays read-only and nothing stamps this on a
passing run: a record written by whatever happened to execute says only someone ran a tool,
which is the exact defect RM73 levels at an attestation produced as a by-product. So this is its
own function behind its own command, and --private-key makes the act attributable rather than
merely evident.
It refuses on an invalid spec and not on a warning. Declaring a set finished that the compiler
will not accept is a contradiction; declaring one finished that carries an unresolvable rsID or an
ungrounded threshold is ordinary, and refusing there would make closure unreachable for every
module whose findings no authored edit can clear (P5, the not_covered class).
Existing check records survive only while they describe these bytes, and then the whole
document is kept verbatim rather than rebuilt — producer names who put the checks, so stamping
this tier's label over it would have the compiler claim an enricher's cross-checks. Records that no
longer hold are dropped and named in dropped_checks: carrying them across would re-bind a claim
to rows the check never saw, which is the failure module_hash exists to catch, committed by the
tool instead of by an edit.
Source code in compiler/src/just_dna_compiler/compiler.py
5730 5731 5732 5733 5734 5735 5736 5737 5738 5739 5740 5741 5742 5743 5744 5745 5746 5747 5748 5749 5750 5751 5752 5753 5754 5755 5756 5757 5758 5759 5760 5761 5762 5763 5764 5765 5766 5767 5768 5769 5770 5771 5772 5773 5774 5775 5776 5777 5778 5779 5780 5781 5782 5783 5784 5785 5786 5787 5788 5789 5790 5791 5792 5793 5794 5795 5796 5797 5798 5799 5800 5801 5802 5803 5804 5805 5806 5807 5808 5809 5810 5811 5812 5813 5814 5815 5816 5817 5818 5819 5820 5821 5822 5823 5824 5825 5826 5827 5828 5829 5830 5831 5832 5833 | |
build_disagreement_error ¶
The one recorded finding strict refuses on, or None (S78, RM143).
This does not move the strict line, and the distinction is the whole item. strict means
reproducible, never right — the compiler has no reference, so it cannot check a coordinate, and
a whole file shifted by one base still passes. genome_build_agreement is the exception on
internal-consistency grounds rather than correctness ones: a recorded finding there says the
module's rows are on a different assembly than the genome_build it declares, which is one
authored file contradicting another. Every other recorded finding is a disagreement between the
module and an outside archive, where the archive is the stale side often enough that failing a
build would have the format arbitrate someone else's dispute — that reasoning is unchanged and
covers clinical_significance, reference_allele and the rest.
A fact the toolchain already established, not a check re-run here. The judgement is the
enricher's, made against the GRCh37 service the compiler may never call (Principle 2); what changed
is that it stops being discarded at the boundary. So the gate keys on a record the enricher
wrote — findings > 0 on that one check — and the compiler adds no reference, no network and no
opinion of its own.
Silent when no attestation exists, deliberately, and that is not a hole this leaves open: an
unverified module is the ordinary case, _read_verification_block says nothing about it on purpose,
and refusing there would fail every module that has never been enriched. What this closes is the
case where the answer was obtained and thrown away.
A stale attestation is dropped before this sees it, which is the correct order: bytes that moved since the check ran are bytes the check did not judge.
Source code in compiler/src/just_dna_compiler/compiler.py
split_cited_literature ¶
split_cited_literature(
rows: list[LiteratureRow],
studies: list[StudyRow],
kind_rows: dict[str, list[Any]] | None = None,
) -> tuple[list[LiteratureRow], list[LiteratureRow]]
(kept, dropped) — the literature rows this module actually cites, and the rest (RM79).
The compiler discards the rest; literature.csv keeps them. A row describing a citation no
study and no citing table row names is dead weight in the artifact: nothing joins to it, and it
is only there
because literature.csv is merge-not-clobber, so a citation the author has since deleted from
studies.csv leaves its row behind. Keeping the row in the CSV is the point of that rule — it is
the pin that makes a re-run cheap — and carrying it into the parquet and the manifest is a
separate decision that nobody had taken deliberately.
What this settles. manifest.literature.missing_count counted exists is False over every
row in the table while the citation_existence verification record counted over the module's
current citations, so the two disagreed in a published manifest with nothing wrong in the
module. Both were honest about their own subject, which is what made it a decision rather than a
bug. Filtering here makes them the same subject by construction, rather than documenting a
discrepancy a reader would have to reconcile.
cited empty means discard nothing, deliberately, and it is not the degenerate case it looks
like: a module that cites nothing at all cannot distinguish "the sidecar is stale" from "the
citations are not authored yet", and emptying its whole table on that reading would delete an
enrichment pass's entire output. The if cited guard the orphan check already had is kept for the
same reason it existed.
On the round trip. reverse_module rebuilds literature.csv from the parquet, so a reversed
copy carries the kept rows only. That is a deterministic narrowing rather than a P7 breach —
literature.csv is a machine-written derived sidecar, not an authored value (the RM69 reading of
Principle 7's letter) — and it converges: everything in the parquet is cited by construction,
so lap two discards nothing and the signatures are a fixed point. The rows are recoverable the way
every derived sidecar's are, by re-running the pass.
Source code in compiler/src/just_dna_compiler/compiler.py
cited_pmids ¶
Every PMID this module cites, from both citation sites, through the one normalizer.
Extracted so RM137's reachability predicate asks the same question split_cited_literature
answers, rather than a second statement of it. The two would drift silently and in the worst
direction: the predicate would call a row unreachable that the drop had kept, so a healthy overlay
would report a finding forever.
The empty case is the caller's to interpret, and it is not "nothing is cited". A module citing
nothing cannot distinguish a stale sidecar from citations not yet authored, so split_cited_literature
discards nothing there — and the predicate must mirror that or every literature update on such a
module reads as unreachable. literature_target_survives builds the mirror; nothing should test
this set for emptiness on its own.
Source code in compiler/src/just_dna_compiler/compiler.py
literature_target_survives ¶
literature_target_survives(
studies: list[StudyRow],
kind_rows: dict[str, list[Any]] | None = None,
) -> Callable[[str], bool]
Can an artifact of this module carry a literature.csv row for this PMID? (RM137)
True when the PMID is cited, because split_cited_literature keeps exactly the cited rows — and
True for everything when the module cites nothing at all, which is that function's own guard
reproduced rather than restated. Without the guard a module with no citations would mark every
literature correction unreachable: a stable false positive, which is worse than the unstable true
one RM137 is about.
Source code in compiler/src/just_dna_compiler/compiler.py
resolution_target_survives ¶
resolution_target_survives(
variants: list[VariantRow],
resolution_rows: list[ResolutionRow],
) -> Callable[[str], bool]
Can an artifact of this module carry a resolution.csv row for this variant_key? (RM137)
resolution.csv has no parquet, so reverse_module rebuilds it from the SNP core and
_write_resolution_csv skips a row with no resolved position — "rows without a resolved position
carry no fact and are skipped". So the surviving set is the subjects this module can place.
Computed from the authored coordinates and the injected table together, which is what makes it answer the same on both laps: on lap 1 the unpositioned row is present and its own cells say it is unpositioned; on lap 2 the row is gone and the authored side still says the same thing. Neither reading depends on the row being there to be matched.
Source code in compiler/src/just_dna_compiler/compiler.py
reverse_module ¶
reverse_module(
parquet_dir: Path,
output_dir: Path,
module_name: str | None = None,
title: str | None = None,
description: str | None = None,
report_title: str | None = None,
icon: str = "database",
color: str = "#6435c9",
version: str | None = None,
write_resolution: bool = True,
genome_build: str | None = None,
) -> Path
Reverse-engineer a parquet module back into the spec DSL (yaml + csv). Returns output_dir.
version (like title/description) is authored module: metadata, out of artifact.digest
and so not materialized into any parquet. Since RM103 it is recovered from the artifact's own
manifest.json when the caller supplies none, rather than dropped: an explicit argument still
wins, and a bare parquet directory with no manifest still leaves the key out of the block. What is
recovered is identity.version_coerced_from where the compile recorded one, falling back to
identity.version — the pre-coercion string, because re-emitting the coerced one gives the next
compile nothing to coerce and version_coerced_from then goes absent on lap 2, which is a module
disagreeing with its own round trip on a published field.
genome_build is not in that class, even though it reaches the artifact the same way (the
manifest, never a parquet column). A wrong title is cosmetic; a wrong build relocates every
coordinate in the module. This used to be hardcoded "GRCh38", so
compile → reverse → compile on a genome_build: GRCh37 module re-emitted it as GRCh38 and the
recompile minted ga4gh:VA.… ids — GRCh38 allele identities for GRCh37 positions — moving
artifact.digest and asserting a variant at a base the module never named. Resolution order is
therefore: this argument, else the artifact's own manifest.json, else "GRCh38" for a bare
parquet directory that records nothing.
write_resolution (default True) also emits resolution.csv — the resolved facts recovered from
the artifact — so reverse → compile reproduces the identical artifact.digest with no network
and no Ensembl reference (Principle 7 hardened from reference-dependent to self-contained). A
coord-keyed row's resolved rsid, dropped from variants.csv, is carried here and restored on
recompile via resolution.resolve_from_table.
The authored tables are written under their one legal name at the root. The machine-written
sidecars go through layout.sidecar_write_path, so a fresh tree gets the preferred spelling
(licensing.csv, not the deprecated sources.csv _FACT_TABLES still names for its parquet) and
an output directory that already carries a copy has that copy overwritten rather than joined by a
second one. Reversing into a directory that already holds two copies of one sidecar raises
layout.SidecarCollision: which of two hand-editable claims to overwrite is not something this
function may decide silently.
Source code in compiler/src/just_dna_compiler/compiler.py
7812 7813 7814 7815 7816 7817 7818 7819 7820 7821 7822 7823 7824 7825 7826 7827 7828 7829 7830 7831 7832 7833 7834 7835 7836 7837 7838 7839 7840 7841 7842 7843 7844 7845 7846 7847 7848 7849 7850 7851 7852 7853 7854 7855 7856 7857 7858 7859 7860 7861 7862 7863 7864 7865 7866 7867 7868 7869 7870 7871 7872 7873 7874 7875 7876 7877 7878 7879 7880 7881 7882 7883 7884 7885 7886 7887 7888 7889 7890 7891 7892 7893 7894 7895 7896 7897 7898 7899 7900 7901 7902 7903 7904 7905 7906 7907 7908 7909 7910 7911 7912 7913 7914 7915 7916 7917 7918 7919 7920 7921 7922 7923 7924 7925 7926 7927 7928 7929 7930 7931 7932 7933 7934 7935 7936 7937 7938 7939 7940 7941 7942 7943 7944 7945 7946 7947 7948 7949 7950 7951 7952 7953 7954 7955 7956 7957 7958 7959 7960 7961 7962 7963 7964 7965 7966 7967 7968 7969 7970 7971 7972 7973 7974 7975 7976 7977 7978 7979 7980 7981 7982 7983 7984 7985 7986 7987 7988 7989 7990 7991 7992 7993 7994 7995 7996 7997 7998 7999 8000 8001 8002 8003 8004 8005 8006 8007 8008 8009 8010 8011 8012 8013 8014 8015 8016 8017 8018 8019 8020 8021 8022 8023 8024 8025 8026 8027 8028 8029 8030 8031 8032 8033 8034 8035 8036 8037 8038 8039 8040 8041 8042 8043 | |