Skip to content

just_dna_compiler.draft

just_dna_compiler.draft

Drafting — append validated rows into an authored CSV without ever clobbering one (0.5).

The mechanism half of the drafting story. A source that publishes a table (CPIC's star-allele definitions, a ClinVar gene slice) can hand its rows to append_rows and they land in the authored CSV a human then owns; the fetching half lives in just-dna-enricher, which is the only tier allowed to go to the network. This module is pure: rows in, CSV appended, report out.

It lives in the compiler because the compiler already does this. reverse_module writes authored CSVs from parquet, and _TABLE_DUPE_KEYS already defines what makes two rows the same row. Building a second, parallel notion of either in the enricher is how they drift.

Append-only, at row granularity. A file-level "refuse if it exists" rule self-defuses after the first gene and makes a multi-gene module unbuildable, so the granularity is the row:

  • a row whose natural key is not present is appended;
  • a row whose key IS present is never rewritten — it is reported as already_present, or as differs when the incoming cells disagree with the authored ones, and the run continues.

That is the line between this and the enricher-co-authoring idea the roadmap parks: appending rows a source publishes leaves content_signature a function of the authored bytes exactly as before, while mutating a cell a human wrote would make the content identity depend on a network fetch. Drift on rows that already exist is the cross-check passes' job to report, not this module's to fix.

Existing cells are never touched. New rows are appended to the open file; the table is rewritten whole only when the header must grow or when placement is delegated, and a rewrite re-reads existing rows as text (_render_existing), so it can never reformat 1.0 into 1. Rows gain an empty cell in a new column — value-neutral, and content_signature ignores unset optional columns by construction.

Where a row lands: the end by default, or its group when asked (group_by, 0.5.1). Appending at the end makes a re-drafted file chronological rather than logical, so a gene's rows scatter as a module grows; delegated insertion puts a row after the last member of its block instead. The tool chooses where, never what — there is deliberately no at=N, because an index is the caller deciding and buys nothing a text editor does not. Moving a row's line number is safe: a pure reorder moves artifact.digest but leaves content_signature untouched (it is order-independent), the round-trip fixed point still holds, and duplicate keys are rejected so order can disambiguate nothing. DraftReport.shifted names every row whose line moved.

DraftError

Bases: RuntimeError

A draft could not be attempted — an unknown table kind, or an unreadable existing file.

RowOutcome dataclass

RowOutcome(
    key: tuple | None,
    status: str,
    differences: dict[str, tuple[Any, Any]] = dict(),
)

What happened to one incoming row.

PartialRow dataclass

PartialRow(
    model: type[BaseModel],
    cells: dict[str, Any],
    stubbed: tuple[str, ...],
    match_on: tuple[str, ...],
)

Cells a source published, plus the columns only a human can decide.

match_on is what makes a re-draft safe. A partial row cannot be keyed the usual way — its natural key runs through a column that is still a placeholder — so sameness is decided on the columns that are filled. Once the human replaces the stub, a re-run matches on those same columns and reports already_present instead of appending the stub again.

rendered

rendered(
    fieldnames: list[str], list_fields: set[str]
) -> dict[str, str]

The CSV cells for this row: what the source gave, the placeholder where it could not.

Source code in compiler/src/just_dna_compiler/draft.py
def rendered(self, fieldnames: list[str], list_fields: set[str]) -> dict[str, str]:
    """The CSV cells for this row: what the source gave, the placeholder where it could not."""
    out: dict[str, str] = {}
    for name in fieldnames:
        if name in self.stubbed:
            out[name] = TEMPLATE_PLACEHOLDER
            continue
        value = self.cells.get(name)
        out[name] = _list_cell(value) if name in list_fields else _scalar_cell(value)
    return out

validation_errors

validation_errors() -> list[str]

What is wrong with the cells the source did publish.

The stubbed columns are validated by omission: the row is built without them and any error located on one is discarded. That avoids the alternative — a per-column table of plausible dummy values — which would be exactly the hand-kept list this module keeps abolishing.

Source code in compiler/src/just_dna_compiler/draft.py
def validation_errors(self) -> list[str]:
    """What is wrong with the cells the source *did* publish.

    The stubbed columns are validated by omission: the row is built without them and any error
    located on one is discarded. That avoids the alternative — a per-column table of plausible
    dummy values — which would be exactly the hand-kept list this module keeps abolishing.
    """
    payload = {
        name: value
        for name, value in self.cells.items()
        if name not in self.stubbed and value is not None and value != ""
    }
    try:
        self.model.model_validate(payload)
    except Exception as exc:  # pydantic ValidationError
        return [
            f"{(error.get('loc') or ('?',))[0]}: {error.get('msg')}"
            for error in getattr(exc, "errors", list)()
            if (error.get("loc") or ("?",))[0] not in self.stubbed
        ]
    return []

model_for

model_for(csv_name: str) -> type[BaseModel]

The row model for an authored CSV name, or DraftError for one this format does not define.

Source code in compiler/src/just_dna_compiler/draft.py
def model_for(csv_name: str) -> type[BaseModel]:
    """The row model for an authored CSV name, or `DraftError` for one this format does not define."""
    model = DRAFTABLE.get(csv_name)
    if model is None:
        raise DraftError(f"{csv_name!r} is not an authored table of this format. Known: {sorted(DRAFTABLE)}")
    return model

natural_key

natural_key(row: BaseModel) -> tuple | None

The identity that decides whether two rows are the same row, or None when the kind has none.

Reuses the compiler's own _TABLE_DUPE_KEYS so an append can never produce a row the compiler would then reject as a duplicate. The binning kinds return None on purpose: their duplicate rule is overlap, not equality (validate_bins), and two bins can conflict while sharing no key — so they are appended and the overlap is caught at compile, where it can actually be judged.

Source code in compiler/src/just_dna_compiler/draft.py
def natural_key(row: BaseModel) -> tuple | None:
    """The identity that decides whether two rows are the same row, or `None` when the kind has none.

    Reuses the compiler's own `_TABLE_DUPE_KEYS` so an append can never produce a row the compiler
    would then reject as a duplicate. The binning kinds return `None` on purpose: their duplicate rule
    is *overlap*, not equality (`validate_bins`), and two bins can conflict while sharing no key — so
    they are appended and the overlap is caught at compile, where it can actually be judged."""
    for table, key_of in (*_CORE_DUPE_KEYS.items(), *_TABLE_DUPE_KEYS.items()):
        if isinstance(row, table):
            return key_of(row)
    return None

blank_template

blank_template(csv_name: str) -> str

A header-only CSV for one table kind, with columns in the model's own field order.

Generated from the model, never a hand-kept list — the same drift-proof route reference.authoring_reference() takes. Starting a new table currently means copying a header out of the docs, which is exactly the thing that goes stale.

Compiler-managed columns are left out (base.authored_field_names): variant_key and authored_ident are stamped at load and never written back by reverse_module, so offering them would invite an author to fill a column the compiler overwrites — and authored_ident is a list, so a rendered cell would not even reload.

Source code in compiler/src/just_dna_compiler/draft.py
def blank_template(csv_name: str) -> str:
    """A header-only CSV for one table kind, with columns in the model's own field order.

    Generated from the model, never a hand-kept list — the same drift-proof route
    `reference.authoring_reference()` takes. Starting a new table currently means copying a header out
    of the docs, which is exactly the thing that goes stale.

    Compiler-managed columns are left out (`base.authored_field_names`): `variant_key` and
    `authored_ident` are stamped at load and never written back by `reverse_module`, so offering them
    would invite an author to fill a column the compiler overwrites — and `authored_ident` is a list,
    so a rendered cell would not even reload."""
    return ",".join(authored_field_names(model_for(csv_name))) + "\n"

required_fields

required_fields(csv_name: str) -> list[str]

The columns an author must fill for this kind — everything without a default.

Field-local only, and deliberately unchanged (it is shipped API). It does not answer "what must I fill to get a valid row": see authoring_requirements, which adds the columns that have a default but reject an empty cell, and the alternative identity groups a model validator enforces.

Source code in compiler/src/just_dna_compiler/draft.py
def required_fields(csv_name: str) -> list[str]:
    """The columns an author must fill for this kind — everything without a default.

    Field-local only, and deliberately unchanged (it is shipped API). It does **not** answer "what
    must I fill to get a valid row": see `authoring_requirements`, which adds the columns that have a
    default but reject an empty cell, and the alternative identity groups a model validator enforces."""
    model = model_for(csv_name)
    authored = set(authored_field_names(model))
    return [name for name, f in model.model_fields.items() if f.is_required() and name in authored]

authoring_requirements

authoring_requirements(csv_name: str) -> dict[str, Any]

What an author actually has to supply for one table kind, machine-readably.

Three parts, because requiredness has three shapes here and only the first is visible to pydantic's is_required():

  • always — columns with no default;
  • any_of — alternative identity groups, any ONE of which satisfies the row (rsid, or chrom+start). Enforced by a model validator, so no per-field flag can express it;
  • defaulted — {column: rendered default} for columns that have a default but reject an empty cell. A template writes these out; see field_category.
Source code in compiler/src/just_dna_compiler/draft.py
def authoring_requirements(csv_name: str) -> dict[str, Any]:
    """What an author actually has to supply for one table kind, machine-readably.

    Three parts, because requiredness has three shapes here and only the first is visible to
    pydantic's `is_required()`:

    * `always`   — columns with no default;
    * `any_of`   — alternative identity groups, any ONE of which satisfies the row (`rsid`, or
      `chrom`+`start`). Enforced by a model validator, so no per-field flag can express it;
    * `defaulted` — `{column: rendered default}` for columns that have a default but reject an empty
      cell. A template writes these out; see `field_category`.
    """
    model = model_for(csv_name)
    categories = {name: field_category(model, name) for name in authored_field_names(model)}
    return {
        "csv": csv_name,
        "always": [n for n, c in categories.items() if c == "required"],
        "any_of": [sorted(group) for group in getattr(model, "REQUIRED_ANY_OF", ())],
        "defaulted": {
            n: _scalar_cell(model.model_fields[n].get_default(call_default_factory=True))
            for n, c in categories.items()
            if c == "defaulted"
        },
        "optional": [n for n, c in categories.items() if c == "optional"],
    }

stub_template

stub_template(csv_name: str, *, rows: int = 1) -> str

A header plus rows stub rows: the sentinel where a human must decide, defaults written out.

The point of the sentinel over a blank cell is that an unreplaced stub cannot compile — vocab.reject_template_placeholders refuses it by name and row, in both modes. A blank required cell would also fail, but a blank optional one would silently become a real row asserting nothing, and the binning kinds' unresolved sentinel could not be used for this at all: it is real data designed to compile.

For a binning kind the mandatory unresolved companion row is emitted too, because a binning table without one is incomplete by contract and the author would otherwise meet that rule as a compile error about a row they never wrote.

Source code in compiler/src/just_dna_compiler/draft.py
def stub_template(csv_name: str, *, rows: int = 1) -> str:
    """A header plus `rows` stub rows: the sentinel where a human must decide, defaults written out.

    The point of the sentinel over a blank cell is that an unreplaced stub **cannot compile** —
    `vocab.reject_template_placeholders` refuses it by name and row, in both modes. A blank required
    cell would also fail, but a blank *optional* one would silently become a real row asserting
    nothing, and the binning kinds' `unresolved` sentinel could not be used for this at all: it is
    real data designed to compile.

    For a binning kind the mandatory `unresolved` companion row is emitted too, because a binning
    table without one is incomplete by contract and the author would otherwise meet that rule as a
    compile error about a row they never wrote."""
    model = model_for(csv_name)
    fieldnames = authored_field_names(model)
    requirements = authoring_requirements(csv_name)
    # Only the FIRST identity group, not the union: the groups are alternatives, so stubbing all of
    # them would tell an author to supply an rsid *and* a coordinate. First is the declaration order,
    # which puts the cheapest identity (`rsid`) in front.
    identity = list(requirements["any_of"][0]) if requirements["any_of"] else []
    stubbed = set(requirements["always"]) | set(identity)
    defaults = requirements["defaulted"]

    def cell(name: str) -> str:
        if name in defaults:
            return defaults[name]
        return TEMPLATE_PLACEHOLDER if name in stubbed else ""

    lines = [",".join(fieldnames)]
    lines.extend(",".join(cell(f) for f in fieldnames) for _ in range(rows))
    if issubclass(model, MeasureBinRow):
        lines.append(",".join(_unresolved_cell(model, f, defaults) for f in fieldnames))
    return "\n".join(lines) + "\n"

append_rows

append_rows(
    spec_dir: Path,
    csv_name: str,
    rows: list[BaseModel],
    *,
    group_by: Sequence[str] = (),
    dry_run: bool = False,
    before_commit: Callable[[], None] | None = None,
) -> DraftReport

Append rows to spec_dir/csv_name, skipping any whose natural key is already there.

Creates the file when absent. Returns a DraftReport describing every incoming row, so a caller can show what a run would do (dry_run=True writes nothing and reports the same thing).

group_by turns the append into a delegated insertion: each new row lands after the last existing row sharing those columns (its gene, its haplotype) instead of at the end of the file. The tool picks the position; the caller never supplies an index. Existing cells are still never rewritten — only their line number can move, and DraftReport.shifted names every row it did.

before_commit is handed straight to layout.atomic_writer, so it runs after this table's bytes are down and before the rename (RM232). It is what lets a drafter's licence row land inside the commit of the table it licenses, the way an enrichment pass's already does (S98, RM231): the enricher records the row only once a table has rows in it, so without this the drafted rows were on disk first and a refused merge left them with nothing to say what licensed them.

It fires only when this call actually writes, which is the grain the caller wants: a run whose every row is already_present or differs adds nothing to this table and reaches no writer, so no licence row is claimed for a table this run did not change. A drafter appending several tables passes the same callable to each — the merge is never-clobber, so N firings record one row, and binding it to only the first or only the last would leave a table committed unlicensed whenever that particular one was the no-op.

Source code in compiler/src/just_dna_compiler/draft.py
def append_rows(
    spec_dir: Path,
    csv_name: str,
    rows: list[BaseModel],
    *,
    group_by: Sequence[str] = (),
    dry_run: bool = False,
    before_commit: Callable[[], None] | None = None,
) -> DraftReport:
    """Append `rows` to `spec_dir/csv_name`, skipping any whose natural key is already there.

    Creates the file when absent. Returns a `DraftReport` describing every incoming row, so a caller
    can show what a run would do (`dry_run=True` writes nothing and reports the same thing).

    `group_by` turns the append into a **delegated insertion**: each new row lands after the last
    existing row sharing those columns (its gene, its haplotype) instead of at the end of the file.
    The tool picks the position; the caller never supplies an index. Existing cells are still never
    rewritten — only their line number can move, and `DraftReport.shifted` names every row it did.

    `before_commit` is handed straight to `layout.atomic_writer`, so it runs after this table's bytes
    are down and **before** the rename (RM232). It is what lets a drafter's licence row land inside
    the commit of the table it licenses, the way an enrichment pass's already does (S98, RM231): the
    enricher records the row only once a table has rows in it, so without this the drafted rows were
    on disk first and a refused merge left them with nothing to say what licensed them.

    **It fires only when this call actually writes**, which is the grain the caller wants: a run whose
    every row is `already_present` or `differs` adds nothing to this table and reaches no writer, so
    no licence row is claimed for a table this run did not change. A drafter appending several tables
    passes the same callable to each — the merge is never-clobber, so N firings record one row, and
    binding it to only the first or only the last would leave a table committed unlicensed whenever
    that particular one was the no-op.
    """
    spec_dir = Path(spec_dir)
    path = _draft_path(spec_dir, csv_name)
    model = model_for(csv_name)

    existing_rows: list[BaseModel] = []
    existing_header: list[str] = []
    # A zero-byte file (a bare `touch`) is treated as absent rather than as an invalid table: it has
    # no header to key against and no rows to preserve, so refusing it would only mean the author has
    # to delete a file that says nothing.
    if path.exists() and path.stat().st_size > 0:
        existing_rows, errors, _ = _load_csv_rows(path, model, csv_name)
        if errors:
            stub_lines = _placeholder_lines(path)
            if stub_lines:
                # The scaffold's own stub (S103): `scaffold --kind haplotypes.csv` then `draft` refused
                # on the template row the first command wrote, with a sentence that read as a broken
                # file. A stub row has no natural key — its key cells *are* the placeholder — so it is
                # never merged over; it is deleted, or never made. Diagnosed, not applied: the drafter
                # appends and does not rewrite an existing row, however machine-written.
                lines = ", ".join(str(n) for n in stub_lines)
                raise DraftError(
                    f"existing {csv_name} still carries the scaffold's template row (line {lines}), so a "
                    f"draft cannot be keyed against it. Delete that row — a drafter writes real rows and "
                    f"the stub is only a column guide — or scaffold without `--kind {csv_name}` when a "
                    f"drafter will write the table. ({errors[0]})"
                )
            raise DraftError(
                f"existing {csv_name} does not validate, so a draft cannot be keyed against it: {errors[0]}"
            )
        with open(path, encoding="utf-8", newline="") as handle:
            existing_header = next(csv.reader(handle), [])

    by_key: dict[tuple, BaseModel] = {}
    for row in existing_rows:
        key = natural_key(row)
        if key is not None:
            by_key.setdefault(key, row)

    outcomes: list[RowOutcome] = []
    to_write: list[BaseModel] = []
    for row in rows:
        key = natural_key(row)
        if key is None:
            outcomes.append(RowOutcome(key=None, status="appended_unkeyed"))
            to_write.append(row)
            continue
        authored = by_key.get(key)
        if authored is None:
            outcomes.append(RowOutcome(key=key, status="added"))
            by_key[key] = row  # so a duplicate inside `rows` itself is caught too
            to_write.append(row)
            continue
        differences = _compare(authored, row)
        outcomes.append(
            RowOutcome(
                key=key,
                status="differs" if differences else "already_present",
                differences=differences,
            )
        )

    # Which columns the file needs: whatever it already has, plus every column the new rows actually
    # set. A column no incoming row fills is not added — a scaffold should not widen a hand-authored
    # table with empties it has nothing to put in.
    field_order = authored_field_names(model)
    filled = {
        name
        for row in to_write
        for name, value in _authored_dump(row).items()
        if value is not None and value != []
    }
    header = existing_header or [f for f in field_order if f in filled or f in set(required_fields(csv_name))]
    extended = [f for f in field_order if f in filled and f not in header]
    fieldnames = header + extended

    if dry_run or not to_write:
        return DraftReport(csv_name, path, outcomes, written=False, header_extended=extended)

    list_fields = _list_fields(model)
    rendered = [_render(row, fieldnames, list_fields) for row in to_write]
    shifted: list[int] = []
    if extended or not existing_header or group_by:
        # The table is written whole when the header has to grow, when there is none yet (the file is
        # absent, or present but empty — a bare `touch`), or when placement is delegated and a row may
        # land mid-file. Existing rows keep their **cells**: `_render_existing` re-reads them as text,
        # so a rewrite can never reformat `1.0` to `1`; only their line number can change.
        previous = _render_existing(path, fieldnames) if existing_header else []
        merged, shifted = place_rows(previous, rendered, group_by)
        # Atomic, on both branches: this is the author's own file, the one class of file this module
        # promises never to damage, and a kill between `open(.., "w")` and the last `writerows` left
        # it truncated to a valid short CSV that nothing downstream could tell from a shorter table.
        with atomic_writer(path, newline="", before_commit=before_commit) as handle:
            writer = csv.DictWriter(handle, fieldnames=fieldnames)
            writer.writeheader()
            writer.writerows(merged)
    else:
        # The common case: the header already fits, so existing bytes are literally untouched — bar a
        # final newline the file may be missing. A hand-authored CSV often has none (plenty of editors
        # do not add one), and appending straight onto it would glue the first new row to the author's
        # last one, corrupting a row this module promises never to touch. The existing bytes are
        # copied through verbatim rather than re-rendered, so the atomic rewrite changes none of them.
        existing_text = path.read_text(encoding="utf-8", newline="")
        with atomic_writer(path, newline="", before_commit=before_commit) as handle:
            handle.write(existing_text)
            if not _ends_with_newline(path):
                handle.write(csv.excel.lineterminator)
            csv.DictWriter(handle, fieldnames=fieldnames).writerows(rendered)

    return DraftReport(csv_name, path, outcomes, written=True, header_extended=extended, shifted=shifted)

append_partial_rows

append_partial_rows(
    spec_dir: Path,
    csv_name: str,
    partials: list[PartialRow],
    *,
    group_by: Sequence[str] = (),
    dry_run: bool = False,
    before_commit: Callable[[], None] | None = None,
) -> DraftReport

Append rows a source could only partly fill, leaving the rest as stubs a human must replace.

Same promises as append_rows: never rewrites a cell, never removes a row, and a row already covered is reported rather than duplicated. The difference is only how sameness is decided — see PartialRow.match_on — because a row whose key column is a placeholder has no usable key yet.

before_commit behaves exactly as it does in append_rows, and for the same reason (RM232).

Source code in compiler/src/just_dna_compiler/draft.py
def append_partial_rows(
    spec_dir: Path,
    csv_name: str,
    partials: list[PartialRow],
    *,
    group_by: Sequence[str] = (),
    dry_run: bool = False,
    before_commit: Callable[[], None] | None = None,
) -> DraftReport:
    """Append rows a source could only partly fill, leaving the rest as stubs a human must replace.

    Same promises as `append_rows`: never rewrites a cell, never removes a row, and a row already
    covered is reported rather than duplicated. The difference is only how sameness is decided — see
    `PartialRow.match_on` — because a row whose key column is a placeholder has no usable key yet.

    `before_commit` behaves exactly as it does in `append_rows`, and for the same reason (RM232).
    """
    spec_dir = Path(spec_dir)
    path = _draft_path(spec_dir, csv_name)
    model = model_for(csv_name)
    fieldnames = authored_field_names(model)
    list_fields = _list_fields(model)

    existing_header: list[str] = []
    if path.exists() and path.stat().st_size > 0:
        with open(path, encoding="utf-8", newline="") as handle:
            existing_header = next(csv.reader(handle), [])

    # The header grows to fit what this batch fills, the same way `append_rows` does — and, as there,
    # what the batch fills is decided over the rows that will actually be WRITTEN, after the loop
    # below has rejected the invalid ones and the ones already covered. Computed over every partial
    # up front, a column an invalid row filled, or an already-present row filled, widened the
    # author's header with a column no written row had anything to put in; and a raw `""` counted
    # as filled where `append_rows`' `_authored_dump` already reads a blank as `None`. The header the
    # existing rows are re-rendered against is settled before that rewrite for the reason
    # `@specific-rejection` records: rendered against the model's full field list and written under
    # the file's narrower header, `csv.DictWriter` raised a raw `ValueError: dict contains fields not
    # in fieldnames`, on every shipped example whose CSV predates a column the model has since gained.
    header = existing_header or fieldnames
    model_fieldnames = fieldnames
    previous = _render_existing(path, header) if existing_header else []

    # The covered-set is built from ONE `match_on`, so every partial in a batch must share it. A
    # provider computing `match_on` per row from whichever identity cells that row happened to carry
    # produces signatures of different arities, none of which can match the covered-set — and the
    # symptom appears only on the second lap, as a file that grows by the same rows every run. Caught
    # here rather than documented, because the silent version is indistinguishable from working.
    distinct_match_on = {partial.match_on for partial in partials}
    if len(distinct_match_on) > 1:
        raise ValueError(
            f"every PartialRow in one batch must share a match_on, because sameness is decided "
            f"against a single covered-set; got {sorted(distinct_match_on)}. Use one constant tuple "
            f"and let absent cells compare as empty."
        )

    outcomes: list[RowOutcome] = []
    accepted: list[PartialRow] = []
    covered = (
        [tuple((row.get(column) or "").strip() for column in partials[0].match_on) for row in previous]
        if partials
        else []
    )

    for partial in partials:
        errors = partial.validation_errors()
        signature = tuple(str(partial.cells.get(column) or "").strip() for column in partial.match_on)
        if errors:
            outcomes.append(RowOutcome(signature, "invalid", {"errors": (None, "; ".join(errors))}))
            continue
        if signature in covered:
            outcomes.append(RowOutcome(signature, "already_present"))
            continue
        covered.append(signature)
        outcomes.append(RowOutcome(signature, "added"))
        accepted.append(partial)

    filled = {
        name
        for partial in accepted
        for name in (*partial.cells, *partial.stubbed)
        if name in partial.stubbed or partial.cells.get(name) not in (None, "", [])
    }
    extended = [name for name in model_fieldnames if name in filled and name not in header]
    fieldnames = header + extended
    to_write = [partial.rendered(fieldnames, list_fields) for partial in accepted]

    if dry_run or not to_write:
        return DraftReport(csv_name, path, outcomes, written=False, header_extended=extended)

    if extended:
        previous = _render_existing(path, fieldnames) if existing_header else []
    merged, shifted = place_rows(previous, to_write, group_by)
    path.parent.mkdir(parents=True, exist_ok=True)
    # Atomic: the author's own file, see `append_rows`.
    with atomic_writer(path, newline="", before_commit=before_commit) as handle:
        writer = csv.DictWriter(handle, fieldnames=fieldnames)
        writer.writeheader()
        writer.writerows(merged)
    return DraftReport(csv_name, path, outcomes, written=True, header_extended=extended, shifted=shifted)

group_of

group_of(
    cells: dict[str, str], columns: Sequence[str]
) -> tuple | None

The block a row belongs to, or None when it declares no group (then it goes to the end).

Source code in compiler/src/just_dna_compiler/draft.py
def group_of(cells: dict[str, str], columns: Sequence[str]) -> tuple | None:
    """The block a row belongs to, or `None` when it declares no group (then it goes to the end)."""
    values = tuple((cells.get(column) or "").strip() for column in columns)
    return values if any(values) else None

place_rows

place_rows(
    existing: list[dict[str, str]],
    incoming: list[dict[str, str]],
    group_by: Sequence[str],
) -> tuple[list[dict[str, str]], list[int]]

Merge incoming into existing, each row after the last member of its group.

Returns the final row list and the original indices of existing rows whose line moved. With no group_by this is a plain append and nothing shifts. Incoming rows keep their relative order within a group, so a provider's own ordering survives.

Source code in compiler/src/just_dna_compiler/draft.py
def place_rows(
    existing: list[dict[str, str]], incoming: list[dict[str, str]], group_by: Sequence[str]
) -> tuple[list[dict[str, str]], list[int]]:
    """Merge `incoming` into `existing`, each row after the last member of its group.

    Returns the final row list and the **original indices** of existing rows whose line moved. With no
    `group_by` this is a plain append and nothing shifts. Incoming rows keep their relative order
    within a group, so a provider's own ordering survives.
    """
    if not group_by:
        return existing + incoming, []
    merged = list(existing)
    original_index = {id(row): position for position, row in enumerate(existing)}
    for row in incoming:
        group = group_of(row, group_by)
        insert_at = len(merged)
        if group is not None:
            for position in range(len(merged) - 1, -1, -1):
                if group_of(merged[position], group_by) == group:
                    insert_at = position + 1
                    break
        merged.insert(insert_at, row)
    shifted = [
        original_index[id(row)]
        for position, row in enumerate(merged)
        if id(row) in original_index and original_index[id(row)] != position
    ]
    return merged, shifted