Skip to content

just_dna_compiler.sweep

just_dna_compiler.sweep

The instrument behind the release record (RM126): measure what a release changed about compiled output, and fail a release whose measurement carries no declaration.

just_dna_format.release_records holds the table and the pure function a consumer reads. This module holds the half that cannot live in the format tier, because producing a record means compiling.

Why a measurement and not a hand-kept map. The map was the first thing everybody proposed and the first thing rejected: it is the defect wearing a public name. Five of the six RM104–RM111 fixes were a derived value restated by hand, and a per-release map of what a release changed is that shape exactly. So the sweep measures, the gate refuses a measured change nobody declared, and the declaration is forced by the measurement rather than remembered by the author.

What it compares. Two trees of compiled output — one produced by the previous release, one by this one, from the same spec inputs. Feeding each side its own tree's reference_examples/ would measure spec drift as compiler drift; the discipline is one spec root, two compilers. The comparison itself needs only this tier: a manifest is JSON and a parquet schema is a polars read.

Scope, and it is narrower than "what a release changed". Compiler-derived outputs only. The enricher's sidecars are unmeasured, which is not the same claim as unchanged, and release_records says so in the record it publishes.

Run it in the release sequence, not as an ordinary test: it needs the previous release actually installed. COMPILER.md § The release-record sweep carries the exact command sequence.

ModuleOutput dataclass

ModuleOutput(
    name: str,
    manifest: dict[str, Any],
    parquet_schemas: dict[str, dict[str, str]],
)

One compiled module as the sweep reads it: the manifest, and every parquet's column types.

ModuleDelta dataclass

ModuleDelta(
    name: str,
    axes: dict[str, bool],
    manifest_fields: tuple[str, ...],
    warnings_added: tuple[str, ...],
    warnings_removed: tuple[str, ...],
    carried_added: tuple[str, ...] = (),
    actionable_added: tuple[str, ...] = (),
)

What moved for one module across the interval, per axis.

SweepMeasurement dataclass

SweepMeasurement(
    before: str,
    after: str,
    axes: dict[str, bool],
    manifest_fields: tuple[str, ...],
    modules: tuple[str, ...],
    only_before: tuple[str, ...],
    only_after: tuple[str, ...],
    per_module: tuple[ModuleDelta, ...],
    build_failures: dict[str, str],
)

The union across every module the sweep could measure, plus the ones it could not.

A module present on one side only is a module this sweep says nothing about, and rolling it silently into an all-False result would be the silence the record exists to replace. Which side it is missing from is a different fact about a different release (RM139), so the two are kept apart and unmeasured is their union rather than a third stored field:

  • only_before — the previous release compiled it and this one does not. That is a regression in the release being cut, and it is fatal however it is declared. A stale reused BEFORE directory holding a module the spec root no longer has lands here too, which is the fail-safe direction; the runbook's answer is a fresh tree every time.
  • only_after — the previous release produced no output for it, so no before state exists to compare against. A module added since the last release, or one whose spec now uses a column the previous release refuses under extra="forbid", which happens in every minor that adds an authored column and exercises it in the corpus. Nothing failed and nothing is measurable, so the record names it in unmeasured and the gate checks that list for equality.

unmeasured property

unmeasured: tuple[str, ...]

Every module measured on one side only, derived rather than stored beside the two halves.

moved_counts property

moved_counts: dict[str, int]

How many of the measured modules moved on each axis — the denominator is modules.

Computed and published, not computed and discarded: an axis that reads False on the record is a measured zero, and a zero with no denominator beside it is the shape a consumer cannot check.

evidence property

evidence: str

The sentence a ReleaseRecord.evidence carries — what was compiled, and what was not.

as_record

as_record(version: str, previous: str) -> ReleaseRecord

The measured half of a record, ready for the declared half to be written onto it.

Produces a record with an empty declared list on purpose: the gate then refuses it until somebody says whether each moved value was wrong or merely absent, which is the whole mechanism that keeps this from becoming a map somebody maintains by memory.

Refuses to stamp an interval this measurement did not take. A record whose version and previous disagree with its own evidence sentence is the hand-kept map again, one field smaller.

Source code in compiler/src/just_dna_compiler/sweep.py
def as_record(self, version: str, previous: str) -> ReleaseRecord:
    """The measured half of a record, ready for the declared half to be written onto it.

    Produces a record with an **empty** `declared` list on purpose: the gate then refuses it
    until somebody says whether each moved value was wrong or merely absent, which is the whole
    mechanism that keeps this from becoming a map somebody maintains by memory.

    Refuses to stamp an interval this measurement did not take. A record whose `version` and
    `previous` disagree with its own `evidence` sentence is the hand-kept map again, one field
    smaller.
    """
    if (
        release_version(version) != self.after
        or release_version(previous) != self.before
        or self.before == self.after
    ):
        raise ValueError(
            f"this sweep measured {self.before} -> {self.after}, so it cannot mint a record for "
            f"{previous} -> {version}"
        )
    if self.only_before:
        # `only_after` becomes a declared denominator below; `only_before` never does. A module
        # the previous release compiled and this one cannot is a regression, and minting a record
        # over its surviving neighbours would publish the false green the gate exists to refuse.
        raise ValueError(
            f"{', '.join(self.only_before)} compiled under {self.before} and not under "
            f"{self.after}, so this measurement cannot mint a record for {version}"
        )
    return ReleaseRecord(
        version=version,
        previous=previous,
        axes=dict(self.axes),
        manifest_fields=list(self.manifest_fields),
        declared=[],
        unmeasured=list(self.only_after),
        evidence=self.evidence,
    )

read_output

read_output(module_dir: Path) -> ModuleOutput

Read one compiled module directory — manifest.json plus every parquet beside it.

Source code in compiler/src/just_dna_compiler/sweep.py
def read_output(module_dir: Path) -> ModuleOutput:
    """Read one compiled module directory — `manifest.json` plus every parquet beside it."""
    manifest = json.loads((module_dir / "manifest.json").read_text(encoding="utf-8"))
    schemas = {path.name: _parquet_schema(path) for path in sorted(module_dir.glob("*.parquet"))}
    return ModuleOutput(name=module_dir.name, manifest=manifest, parquet_schemas=schemas)

read_outputs

read_outputs(root: Path) -> dict[str, ModuleOutput]

Every compiled module under root, keyed by directory name.

Source code in compiler/src/just_dna_compiler/sweep.py
def read_outputs(root: Path) -> dict[str, ModuleOutput]:
    """Every compiled module under `root`, keyed by directory name."""
    found = {
        child.name: read_output(child)
        for child in sorted(root.iterdir())
        if child.is_dir() and (child / "manifest.json").is_file()
    }
    if not found:
        raise ValueError(f"no compiled modules found under {root}")
    return found

build_outputs

build_outputs(
    spec_root: Path, out_root: Path
) -> tuple[dict[str, ModuleOutput], dict[str, str]]

Compile every spec under spec_root into out_root/<name>/ with THIS compiler, and what broke.

Discovery rather than a list, the same rule test_reference_examples_roundtrip follows: a spec added without being added here is a spec nobody sweeps, and adding one is precisely when a new shape arrives.

The second half of the pair is the compiler's own errors per spec that did not compile. Returned rather than only logged (RM139): the module lands in only_before, which is fatal, and a fatal finding that names the error is the difference between a release an operator can fix and one they have to go looking for.

Source code in compiler/src/just_dna_compiler/sweep.py
def build_outputs(spec_root: Path, out_root: Path) -> tuple[dict[str, ModuleOutput], dict[str, str]]:
    """Compile every spec under `spec_root` into `out_root/<name>/` with THIS compiler, and what broke.

    Discovery rather than a list, the same rule `test_reference_examples_roundtrip` follows: a spec
    added without being added here is a spec nobody sweeps, and adding one is precisely when a new
    shape arrives.

    The second half of the pair is the compiler's own errors per spec that did not compile. Returned
    rather than only logged (RM139): the module lands in `only_before`, which is fatal, and a fatal
    finding that names the error is the difference between a release an operator can fix and one they
    have to go looking for.
    """
    specs = sorted(d for d in spec_root.iterdir() if (d / "module_spec.yaml").is_file())
    if not specs:
        raise ValueError(f"no module specs found under {spec_root}")
    built: dict[str, ModuleOutput] = {}
    failures: dict[str, str] = {}
    for spec in specs:
        result = compile_module(spec, out_root / spec.name)
        if not result.success:
            # Not silently skipped: the module lands outside `built`, so it reaches `only_before` and
            # the gate refuses the release. A sweep that quietly measured the modules that still work
            # is the silence this whole surface exists to replace.
            failures[spec.name] = "; ".join(result.errors)
            logger.warning("sweep: %s did not compile: %s", spec.name, failures[spec.name])
            continue
        built[spec.name] = read_output(out_root / spec.name)
    if not built:
        raise ValueError(
            f"no spec under {spec_root} compiled — nothing to measure (this is a broken compiler, "
            "not a release that changed nothing)"
        )
    return built, failures

changed_manifest_fields

changed_manifest_fields(
    before: dict[str, Any], after: dict[str, Any]
) -> tuple[str, ...]

The published manifest paths that moved, with EXCLUDED_MANIFEST_FIELDS never among them.

Source code in compiler/src/just_dna_compiler/sweep.py
def changed_manifest_fields(before: dict[str, Any], after: dict[str, Any]) -> tuple[str, ...]:
    """The published manifest paths that moved, with `EXCLUDED_MANIFEST_FIELDS` never among them."""
    flat_before = _flatten(before)
    flat_after = _flatten(after)
    paths = set(flat_before) | set(flat_after)
    moved = {
        path
        for path in paths
        if not _is_excluded(path) and flat_before.get(path, _ABSENT) != flat_after.get(path, _ABSENT)
    }
    return tuple(sorted(moved))

compare_module

compare_module(
    before: ModuleOutput, after: ModuleOutput
) -> ModuleDelta

The per-axis movement for one module.

The warnings axis is computed apart and reported apart, and it is not folded into manifest_fields however published compilation.warnings is. A release that reworks the warning channel would otherwise report a manifest field changed on every module in a catalogue, and a registry acting on that mints an immutable PATCH for a message change.

RM131's carried split is now that discriminator, and it is reported rather than acted on. axes["warnings"] still fires on any change to the set — narrowing it here would make a published axis mean something different from what every record already written claims about it, and the axis is outside RECOMPILE_DRIVING_AXES anyway, so nothing keys a rebuild on it. What the split buys is the reading: carried_added is a finding no author can clear appearing, usually this repository saying more about a limit it always had; actionable_added is work arriving at somebody's door. A manifest with no carried field reports every addition as actionable, which is the safe direction — it never tells a reader that a finding they could fix is unfixable.

Source code in compiler/src/just_dna_compiler/sweep.py
def compare_module(before: ModuleOutput, after: ModuleOutput) -> ModuleDelta:
    """The per-axis movement for one module.

    **The `warnings` axis is computed apart and reported apart**, and it is not folded into
    `manifest_fields` however published `compilation.warnings` is. A release that reworks the warning
    channel would otherwise report *a manifest field changed* on every module in a catalogue, and a
    registry acting on that mints an immutable PATCH for a message change.

    **RM131's `carried` split is now that discriminator**, and it is reported rather than acted on.
    `axes["warnings"]` still fires on any change to the set — narrowing it here would make a
    published axis mean something different from what every record already written claims about it,
    and the axis is outside `RECOMPILE_DRIVING_AXES` anyway, so nothing keys a rebuild on it. What the
    split buys is the reading: `carried_added` is a finding no author can clear appearing, usually
    this repository saying more about a limit it always had; `actionable_added` is work arriving at
    somebody's door. A manifest with no `carried` field reports every addition as actionable, which
    is the safe direction — it never tells a reader that a finding they could fix is unfixable.
    """
    before_warnings = set(_warning_lists(before.manifest))
    after_warnings = set(_warning_lists(after.manifest))
    carried_after = _carried_list(after.manifest)
    fields = changed_manifest_fields(before.manifest, after.manifest)
    axes = {
        "parquet_schema": before.parquet_schemas != after.parquet_schemas,
        "parquet_bytes": (
            before.manifest.get("artifact", {}).get("digest")
            != after.manifest.get("artifact", {}).get("digest")
        ),
        "content_signature": (
            before.manifest.get("content_signature") != after.manifest.get("content_signature")
        ),
        "manifest_fields": bool(fields),
        "warnings": before_warnings != after_warnings,
    }
    added = tuple(sorted(after_warnings - before_warnings))
    return ModuleDelta(
        name=after.name,
        axes=axes,
        manifest_fields=fields,
        warnings_added=added,
        warnings_removed=tuple(sorted(before_warnings - after_warnings)),
        carried_added=tuple(w for w in added if w in carried_after),
        actionable_added=tuple(w for w in added if w not in carried_after),
    )

compare_outputs

compare_outputs(
    before: dict[str, ModuleOutput],
    after: dict[str, ModuleOutput],
    build_failures: dict[str, str] | None = None,
) -> SweepMeasurement

Union the per-module deltas into the measurement one release record carries.

build_failures is what build_outputs saw refusing to compile under THIS release, so a module in only_before can be reported with the compiler's own error rather than as a bare absence.

Source code in compiler/src/just_dna_compiler/sweep.py
def compare_outputs(
    before: dict[str, ModuleOutput],
    after: dict[str, ModuleOutput],
    build_failures: dict[str, str] | None = None,
) -> SweepMeasurement:
    """Union the per-module deltas into the measurement one release record carries.

    `build_failures` is what `build_outputs` saw refusing to compile under THIS release, so a module
    in `only_before` can be reported with the compiler's own error rather than as a bare absence.
    """
    shared = sorted(set(before) & set(after))
    deltas = tuple(compare_module(before[name], after[name]) for name in shared)
    axes = {axis: any(delta.axes[axis] for delta in deltas) for axis in sorted(VALID_RELEASE_OUTPUT_AXES)}
    fields: set[str] = set()
    for delta in deltas:
        fields.update(delta.manifest_fields)
    return SweepMeasurement(
        before=_release_of(before, "before"),
        after=_release_of(after, "after"),
        axes=axes,
        manifest_fields=tuple(sorted(fields)),
        modules=tuple(shared),
        only_before=tuple(sorted(set(before) - set(after))),
        only_after=tuple(sorted(set(after) - set(before))),
        per_module=deltas,
        build_failures=dict(build_failures or {}),
    )

gate_findings

gate_findings(
    measurement: SweepMeasurement,
    version: str,
    records: dict[str, ReleaseRecord] | None = None,
) -> tuple[list[str], list[str]]

(findings, notes) — the release gate. A non-empty findings fails the release.

A finding is a measured movement no declaration covers, or a sweep that did not measure what the caller thinks it did. The second half is not decoration: the likeliest operator error in the documented sequence is running the sweep before uv sync has propagated the version bump, which builds the AFTER tree with the previous release still installed. Both sides are then one compiler, every axis reads False, and a gate checking only the lower end of the interval passes a release it never measured — a false green in the one mechanism the whole item rests on. So the interval's upper end is checked against the version being gated, a degenerate interval is refused outright, and a module that compiled on one side only fails rather than vanishing into an all-False result over its surviving neighbours.

Which side a module is missing from decides which of those it is (RM139). One side only was read as a compile failed until the first real cut, where it was wrong: RM70 put an optional column on pharm_variants.csv, one reference example uses it, and 0.6.6 refuses that spec under extra="forbid". Nothing failed — the previous release simply cannot produce a before state, and that recurs in every minor that adds an authored column and exercises it in the corpus, as it does for any example added since the last release. So:

  • a module in the BEFORE tree and not the AFTER one is a regression in the release being gated and fails unconditionally, with the compiler's own error beside it where the AFTER side was built here;
  • a module in the AFTER tree and not the BEFORE one fails unless the record's unmeasured names it, and a record naming one the sweep did measure is reported as a note.

That is the same forcing shape as declared, not an exemption bolted beside it: as_record mints the measured half and the gate refuses until the author commits it to the published record. It cannot silence anything, because it reaches neither direction that carries a measurement — a movement on a measured module still gates however unmeasured reads, and a regression is fatal however it is listed.

A note is the other direction — a declaration this sweep did not see move — and it is deliberately not a failure: the reference corpus is sixteen modules and a real correction can land on a shape none of them has.

The gate runs in the bump → uv sync → tag sequence rather than as an ordinary test, because it needs the previous release actually installed.

Source code in compiler/src/just_dna_compiler/sweep.py
def gate_findings(
    measurement: SweepMeasurement,
    version: str,
    records: dict[str, ReleaseRecord] | None = None,
) -> tuple[list[str], list[str]]:
    """`(findings, notes)` — the release gate. A non-empty `findings` fails the release.

    A finding is a measured movement no declaration covers, **or a sweep that did not measure what
    the caller thinks it did**. The second half is not decoration: the likeliest operator error in
    the documented sequence is running the sweep before `uv sync` has propagated the version bump,
    which builds the AFTER tree with the *previous* release still installed. Both sides are then one
    compiler, every axis reads `False`, and a gate checking only the lower end of the interval passes
    a release it never measured — a false green in the one mechanism the whole item rests on. So the
    interval's **upper** end is checked against the version being gated, a degenerate interval is
    refused outright, and a module that compiled on one side only fails rather than vanishing into an
    all-`False` result over its surviving neighbours.

    **Which side a module is missing from decides which of those it is (RM139).** *One side only* was
    read as *a compile failed* until the first real cut, where it was wrong: RM70 put an optional
    column on `pharm_variants.csv`, one reference example uses it, and 0.6.6 refuses that spec under
    `extra="forbid"`. Nothing failed — the previous release simply cannot produce a before state, and
    that recurs in every minor that adds an authored column and exercises it in the corpus, as it does
    for any example added since the last release. So:

    * a module in the BEFORE tree and not the AFTER one is a **regression in the release being
      gated** and fails unconditionally, with the compiler's own error beside it where the AFTER side
      was built here;
    * a module in the AFTER tree and not the BEFORE one fails **unless the record's `unmeasured`
      names it**, and a record naming one the sweep did measure is reported as a note.

    That is the same forcing shape as `declared`, not an exemption bolted beside it: `as_record`
    mints the measured half and the gate refuses until the author commits it to the published record.
    It cannot silence anything, because it reaches neither direction that carries a measurement — a
    movement on a measured module still gates however `unmeasured` reads, and a regression is fatal
    however it is listed.

    A **note** is the other direction — a declaration this sweep did not see move — and it is
    deliberately not a failure: the reference corpus is sixteen modules and a real correction can land
    on a shape none of them has.

    The gate runs in the bump → `uv sync` → tag sequence rather than as an ordinary test, because it
    needs the previous release actually installed.
    """
    table = RELEASE_RECORDS if records is None else records
    version = release_version(version)
    findings: list[str] = []
    notes: list[str] = []

    # The interval checks come first and are unconditional, because every one of them describes a
    # sweep that did not measure what the caller thinks it measured — and a record cannot cover a
    # measurement that was never taken.
    if measurement.before == measurement.after:
        findings.append(
            f"the sweep {DEGENERATE_INTERVAL_PHRASE}: both trees are stamped {measurement.after}. "
            "The likeliest cause is the documented sequence run before `uv sync` propagated the "
            "version bump, so the AFTER tree was built by the previous release"
        )
    if measurement.after != version:
        findings.append(
            f"the sweep {WRONG_VERSION_PHRASE}: measured {measurement.before} → "
            f"{measurement.after}, gating {version}"
        )
    if not measurement.modules:
        findings.append(f"the sweep {NO_MODULES_PHRASE}: the two trees share no module")
    for name in measurement.only_before:
        reason = measurement.build_failures.get(name)
        findings.append(
            f"{name} {UNMEASURED_MODULE_PHRASE} — it {REGRESSED_MODULE_PHRASE}"
            + (f": {reason}" if reason else "")
        )

    record = table.get(version)
    # Read before the `record is None` return so an unmeasured module is reported alongside the
    # missing record rather than behind it — an author minting the record needs both in one pass.
    declared_unmeasured = set(record.unmeasured) if record is not None else set()
    for name in measurement.only_after:
        if name not in declared_unmeasured:
            findings.append(
                f"{name} {UNMEASURED_MODULE_PHRASE} — it {UNDECLARED_UNMEASURED_PHRASE} as "
                "unmeasured. Nothing failed if its spec uses a column the previous release refuses, "
                "or if it is new since that release, but the record must say so"
            )
    for name in sorted(declared_unmeasured - set(measurement.only_after)):
        notes.append(f"{name} {OVERDECLARED_UNMEASURED_NOTE_PHRASE}")

    if record is None:
        findings.append(
            f"{version} {NO_RECORD_PHRASE}: a release that changed compiled output must declare "
            "what it changed, and a release where nothing moved must record the measured zero "
            "with its evidence"
        )
        return findings, notes

    if record.previous != measurement.before:
        findings.append(
            f"the sweep {WRONG_PREVIOUS_PHRASE}: measured {measurement.before} → "
            f"{measurement.after}, record {version} names {record.previous}"
        )

    declared_axes = {change.axis for change in record.declared}
    for axis in sorted(VALID_RELEASE_OUTPUT_AXES):
        if measurement.axes[axis] and record.axes.get(axis) is not True:
            findings.append(f"axis {axis!r} {UNDECLARED_AXIS_PHRASE}")
        if measurement.axes[axis] and axis not in declared_axes:
            findings.append(f"axis {axis!r} {UNDECLARED_KIND_PHRASE}")
        if record.axes.get(axis) is True and not measurement.axes[axis]:
            notes.append(f"axis {axis!r} {OVERDECLARED_NOTE_PHRASE}")

    listed = set(record.manifest_fields)
    for field in measurement.manifest_fields:
        if field not in listed:
            findings.append(f"manifest field {field!r} {UNDECLARED_FIELD_PHRASE}")
    for field in sorted(listed - set(measurement.manifest_fields)):
        notes.append(f"manifest field {field!r} {OVERDECLARED_NOTE_PHRASE}")
    return findings, notes

measurement_json

measurement_json(
    measurement: SweepMeasurement,
) -> dict[str, Any]

The measurement as plain JSON, so a release script can diff or archive it.

Source code in compiler/src/just_dna_compiler/sweep.py
def measurement_json(measurement: SweepMeasurement) -> dict[str, Any]:
    """The measurement as plain JSON, so a release script can diff or archive it."""
    return {
        "before": measurement.before,
        "after": measurement.after,
        "axes": dict(measurement.axes),
        "manifest_fields": list(measurement.manifest_fields),
        "modules": list(measurement.modules),
        # The union stays under its original key — a release script piping this to `jq` predates the
        # split — with the two halves beside it rather than in place of it.
        "unmeasured": list(measurement.unmeasured),
        "only_before": list(measurement.only_before),
        "only_after": list(measurement.only_after),
        "build_failures": dict(measurement.build_failures),
        "evidence": measurement.evidence,
        "per_module": [
            {
                "name": delta.name,
                "axes": dict(delta.axes),
                "manifest_fields": list(delta.manifest_fields),
                "warnings_added": list(delta.warnings_added),
                "warnings_removed": list(delta.warnings_removed),
                "carried_added": list(delta.carried_added),
                "actionable_added": list(delta.actionable_added),
            }
            for delta in measurement.per_module
        ],
    }