Skip to content

just_dna_compiler.scaffold

just_dna_compiler.scaffold

Scaffolding — create a module's authored files from nothing, refusing to touch what exists (0.5).

The de-novo half of the authoring story, beside draft's append-into-an-existing-table half. Nothing generated module_spec.yaml before this: every example and test in the repo hand-wrote one, which is the same "copy a shape out of the docs" that blank_template exists to abolish for CSVs.

Pure and offline (Principle 2) — it writes only into the spec directory it is given, and only files that are not already there.

Refusal is file-level here, and row-level in draft — the difference is derivable, not stipulated. draft's docstring rejects file-level refusal because it "self-defuses after the first gene": you re-run an append per gene, so refusing a file that exists would make a multi-gene module unbuildable. Scaffolding has no such re-run: you create a module's tables once. And row-level refusal presupposes a natural key that decides sameness, which a stub row does not have — its key columns are the placeholder. What it does share with draft is the definition of absent: a zero-byte file (a bare touch) counts as not there, so the two never disagree about whether a file exists.

Refusal is also per file, not per run: an existing variants.csv must not stop a pgs.csv from being scaffolded into the same module, or a module could never gain a second table kind.

ScaffoldPlan dataclass

ScaffoldPlan(
    spec_dir: Path,
    created: list[Path] = list(),
    refused: list[tuple[Path, str]] = list(),
    warnings: list[str] = list(),
    written: bool = False,
)

What a scaffold run created, and what it refused to touch.

companions_for

companions_for(kinds: Sequence[str]) -> list[str]

The kinds to scaffold beside those asked for, with COMPANION_KINDS' conditional half applied.

Public because a consumer passing the raw mapping through answers a question it cannot answer alone, and would then contradict its own composition advice: never add an empty table to keep another company (S49).

variants.csv still pulls studies.csv unconditionally — that direction has no condition, since the compiler requires grounding evidence for a variant claim however the module is composed. studies.csv pulls variants.csv only when nothing else recognised was requested, which is the "alone" its own justification always named: beside a binning table the studies row grounds the bins through pmid and the module needs no variant at all.

Source code in compiler/src/just_dna_compiler/scaffold.py
def companions_for(kinds: Sequence[str]) -> list[str]:
    """The kinds to scaffold *beside* those asked for, with `COMPANION_KINDS`' conditional half applied.

    Public because a consumer passing the raw mapping through answers a question it cannot answer
    alone, and would then contradict its own composition advice: never add an empty table to keep
    another company (S49).

    `variants.csv` still pulls `studies.csv` unconditionally — that direction has no condition, since
    the compiler requires grounding evidence for a variant claim however the module is composed.
    `studies.csv` pulls `variants.csv` **only when nothing else recognised was requested**, which is
    the "alone" its own justification always named: beside a binning table the studies row grounds the
    bins through `pmid` and the module needs no variant at all.
    """
    requested = list(dict.fromkeys(kinds))
    grounded_by_another_table = bool(_RECOGNIZED_TABLES & set(requested))
    companions: list[str] = []
    for kind in requested:
        for companion in COMPANION_KINDS.get(kind, ()):
            if companion in requested or companion in companions:
                continue
            if kind == "studies.csv" and grounded_by_another_table:
                continue
            companions.append(companion)
    return companions

module_spec_template

module_spec_template(*, name: str | None = None) -> str

A module_spec.yaml skeleton generated from the live models.

Required scalars carry the placeholder, so the file cannot compile until a human has filled them; defaulted ones carry their real default, so the author sees the values that will apply. Optional blocks (panel, license) are left out entirely rather than emitted empty — panel: null would fail typing, and an empty authorship: [] would quietly claim the module has no authors.

Source code in compiler/src/just_dna_compiler/scaffold.py
def module_spec_template(*, name: str | None = None) -> str:
    """A `module_spec.yaml` skeleton generated from the live models.

    Required scalars carry the placeholder, so the file cannot compile until a human has filled them;
    defaulted ones carry their real default, so the author sees the values that will apply. Optional
    blocks (`panel`, `license`) are left out entirely rather than emitted empty — `panel: null` would
    fail typing, and an empty `authorship: []` would quietly claim the module has no authors.
    """
    spec: dict[str, Any] = {}
    for field_name in ModuleSpecConfig.model_fields:
        category = field_category(ModuleSpecConfig, field_name)
        if category == "optional":
            continue
        annotation = ModuleSpecConfig.model_fields[field_name].annotation
        if isinstance(annotation, type) and issubclass(annotation, BaseModel):
            spec[field_name] = {
                inner: _stub_value(annotation, inner)
                for inner in authored_field_names(annotation)
                if field_category(annotation, inner) != "optional"
            }
        else:
            value = _stub_value(ModuleSpecConfig, field_name)
            # An empty collection default is dropped rather than emitted: `authorship: []` is valid
            # but states "this module has no contributors", which is a claim, not a starting point.
            # The author adds the block when there is someone to name (RM14).
            if isinstance(value, (list, dict, tuple)) and not value:
                continue
            spec[field_name] = value
    if name is not None and isinstance(spec.get("module"), dict):
        spec["module"]["name"] = name
    return yaml.safe_dump(spec, sort_keys=False, allow_unicode=True, default_flow_style=False)

scaffold_module

scaffold_module(
    spec_dir: Path,
    *,
    kinds: Sequence[str] = (),
    name: str | None = None,
    rows: int = 1,
    dry_run: bool = False,
) -> ScaffoldPlan

Create module_spec.yaml plus a stub CSV per requested kind, skipping whatever exists.

Never overwrites, never deletes, never edits: an existing file is reported and left byte-for-byte alone. dry_run reports the same plan and writes nothing.

Source code in compiler/src/just_dna_compiler/scaffold.py
def scaffold_module(
    spec_dir: Path,
    *,
    kinds: Sequence[str] = (),
    name: str | None = None,
    rows: int = 1,
    dry_run: bool = False,
) -> ScaffoldPlan:
    """Create `module_spec.yaml` plus a stub CSV per requested kind, skipping whatever exists.

    Never overwrites, never deletes, never edits: an existing file is reported and left byte-for-byte
    alone. `dry_run` reports the same plan and writes nothing.
    """
    requested = list(dict.fromkeys(kinds))
    for kind in requested:
        model_for(kind)  # raises DraftError for a kind this format does not define

    companions = companions_for(requested)
    plan = ScaffoldPlan(spec_dir=spec_dir)
    for companion in companions:
        plan.warnings.append(
            f"{companion} added: a module carrying {', '.join(k for k in requested if companion in COMPANION_KINDS.get(k, ()))}"
            f" does not compile without it."
        )

    targets: list[tuple[Path, str]] = [(spec_dir / MODULE_SPEC, module_spec_template(name=name))]
    targets.extend((spec_dir / kind, stub_template(kind, rows=rows)) for kind in requested + companions)

    for path, content in targets:
        if _exists(path):
            plan.refused.append((path, "already exists — scaffolding never overwrites"))
            continue
        plan.created.append(path)
        if not dry_run:
            path.parent.mkdir(parents=True, exist_ok=True)
            path.write_text(content, encoding="utf-8")
    plan.written = bool(plan.created) and not dry_run
    return plan