Skip to content

just_dna_format.layout

just_dna_format.layout

Where a module's machine-written sidecars live, and what they may be called.

The compiler reads these files, the enricher writes them, a publisher uploads them and a registry splits them back apart — four parties, one layout. just_dna_enricher.locations exists because every disagreement about a snapshot's layout so far has been silent, and this is the same class of fact one tier down, so it lives in the schema tier where both consumers import it rather than copy it.

Nothing here fetches, parses or validates. It is the standard library and two tuples of names, which is why it can sit in the dependency-light tier at all — pathlib for the places, and since S66 os plus tempfile for the writing, because how a sidecar is put on disk turned out to be the same class of fact as where: four parties agreeing on the name is worth nothing if a killed run leaves half a table under it.

Scope is the machine-written sidecars only — resolution.csv and the fact tables. The authored DSL (module_spec.yaml, variants.csv, studies.csv, the table kinds) has exactly one legal name in exactly one legal place, and that asymmetry is deliberate: a second legal home for variants.csv means a module can carry two copies and the one the compiler ignores is invisible, which is the silent-success shape this codebase treats as the worst kind of mistake. Only the machine-written tables move, because only they have a machine that knows where to put them.

SidecarCollision

Bases: ValueError

Two files claim to be the same sidecar, and neither may be silently preferred.

Not a merge and not newest-wins. These tables are fact-hashed and human-overridable — a curator edits a row the enricher wrote and that edit is the point — so two copies are two legitimate claims and picking one discards somebody's work without saying so. The only honest answer names both paths and stops.

sidecar_key

sidecar_key(name: str) -> str

The table key behind any of its spellings — sources.csv for licensing.csv — else name.

The key is the spelling sources.parquet and manifest.sources keep, and it is what the map is keyed on. A caller holding bytes and a filename — a tar member, an upload part — has the preferred spelling and not the key, and until S96 the helpers below read that as a table they had never heard of, answering the one-tuple ('licensing.csv',): sidecar_write_path then created the preferred copy beside the deprecated one, which is the collision the docstring below says it exists to prevent. Either spelling is a key now.

Source code in schema/src/just_dna_format/layout.py
def sidecar_key(name: str) -> str:
    """The table key behind any of its spellings — `sources.csv` for `licensing.csv` — else `name`.

    The key is the spelling `sources.parquet` and `manifest.sources` keep, and it is what the map is
    keyed on. A caller holding bytes and a filename — a tar member, an upload part — has the *preferred*
    spelling and not the key, and until S96 the helpers below read that as a table they had never heard
    of, answering the one-tuple `('licensing.csv',)`: `sidecar_write_path` then created the preferred
    copy beside the deprecated one, which is the collision the docstring below says it exists to
    prevent. Either spelling is a key now.
    """
    return _KEY_FOR_SPELLING.get(name, name)

sidecar_spellings

sidecar_spellings(name: str) -> tuple[str, ...]

Every accepted filename for name, deprecated first, preferred last.

name may be the table key or any of its spellings — sidecar_spellings("licensing.csv") answers the same pair as sidecar_spellings("sources.csv"). A name the map does not know has exactly one spelling, its own.

Source code in schema/src/just_dna_format/layout.py
def sidecar_spellings(name: str) -> tuple[str, ...]:
    """Every accepted filename for `name`, deprecated first, preferred last.

    `name` may be the table key **or any of its spellings** — `sidecar_spellings("licensing.csv")`
    answers the same pair as `sidecar_spellings("sources.csv")`. A name the map does not know has
    exactly one spelling, its own.
    """
    return SIDECAR_SPELLINGS.get(sidecar_key(name), (name,))

preferred_spelling

preferred_spelling(name: str) -> str

The spelling a fresh file should be created under.

Source code in schema/src/just_dna_format/layout.py
def preferred_spelling(name: str) -> str:
    """The spelling a fresh file should be created under."""
    return sidecar_spellings(name)[-1]

is_deprecated_spelling

is_deprecated_spelling(filename: str) -> bool

Whether a filename is a spelling that still reads but should no longer be written.

Source code in schema/src/just_dna_format/layout.py
def is_deprecated_spelling(filename: str) -> bool:
    """Whether a filename is a spelling that still reads but should no longer be written."""
    return filename in DEPRECATED_SPELLINGS

sidecar_candidates

sidecar_candidates(spec_dir: Path, name: str) -> list[Path]

Every place name may legally be — each spelling, at the root and under derived/.

Order is the tie-break nothing else can supply, so it is fixed rather than incidental: a caller that wants "the one that exists" gets a deterministic answer, and a caller that wants "where do I create it" reads the same list from the front. The root comes first because it is the layout every existing module and every reverse_module output uses.

Source code in schema/src/just_dna_format/layout.py
def sidecar_candidates(spec_dir: Path, name: str) -> list[Path]:
    """Every place `name` may legally be — each spelling, at the root and under `derived/`.

    Order is the tie-break nothing else can supply, so it is fixed rather than incidental: a caller
    that wants "the one that exists" gets a deterministic answer, and a caller that wants "where do I
    create it" reads the same list from the front. The root comes first because it is the layout every
    existing module and every `reverse_module` output uses.
    """
    root = Path(spec_dir)
    return [
        root / directory / spelling
        for directory in ("", DERIVED_SUBDIR)
        for spelling in sidecar_spellings(name)
    ]

resolve_sidecar

resolve_sidecar(spec_dir: Path, name: str) -> Path | None

The single existing copy of a sidecar, or None when the module carries none.

Raises SidecarCollision when more than one candidate exists. That is an error rather than a warning because there is no correct silent behaviour available — see the exception's own docstring.

Source code in schema/src/just_dna_format/layout.py
def resolve_sidecar(spec_dir: Path, name: str) -> Path | None:
    """The single existing copy of a sidecar, or `None` when the module carries none.

    Raises `SidecarCollision` when more than one candidate exists. That is an error rather than a
    warning because there is no correct silent behaviour available — see the exception's own docstring.
    """
    present = [path for path in sidecar_candidates(spec_dir, name) if path.is_file()]
    if not present:
        return None
    if len(present) > 1:
        listed = " and ".join(str(path) for path in present)
        raise SidecarCollision(
            f"{listed} are the same table in two places, and both are present. These tables are "
            f"fact-hashed and may be edited by hand, so two copies are two claims and neither can be "
            f"preferred without discarding the other — keep one and delete the other. "
            f"{preferred_spelling(name)!r} is the spelling to keep; either the spec root or "
            f"{DERIVED_SUBDIR}/ is a fine place for it."
        )
    return present[0]

sidecar_relative_names

sidecar_relative_names(name: str) -> list[str]

Every legal spelling-and-location of name, as paths relative to the spec directory.

For hashing rather than reading: integrity.file_entries skips the names that are not there, so handing it this list records whichever one the module actually carries — and FileEntry.name then carries the relative path, which is what a registry re-splitting a downloaded tree needs.

Source code in schema/src/just_dna_format/layout.py
def sidecar_relative_names(name: str) -> list[str]:
    """Every legal spelling-and-location of `name`, as paths relative to the spec directory.

    For hashing rather than reading: `integrity.file_entries` skips the names that are not there, so
    handing it this list records whichever one the module actually carries — and `FileEntry.name` then
    carries the relative path, which is what a registry re-splitting a downloaded tree needs.
    """
    return [str(path) for path in sidecar_candidates(Path(), name)]

sidecar_write_path

sidecar_write_path(spec_dir: Path, name: str) -> Path

Where a pass should write a sidecar: the copy that exists, else the preferred spelling.

Write to the file you read. A pass that always created the preferred spelling at the root would, on a module carrying the deprecated one or a derived/ tree, leave two copies behind — the collision above, produced by following the documented workflow rather than by misusing it. Running enrich on a downloaded split module is exactly that path.

With nothing to follow it creates the flat layout, because derived/ is tolerated rather than canonical: a module only ends up split because somebody chose to split it.

Source code in schema/src/just_dna_format/layout.py
def sidecar_write_path(spec_dir: Path, name: str) -> Path:
    """Where a pass should write a sidecar: the copy that exists, else the preferred spelling.

    **Write to the file you read.** A pass that always created the preferred spelling at the root
    would, on a module carrying the deprecated one or a `derived/` tree, leave two copies behind — the
    collision above, produced by following the documented workflow rather than by misusing it. Running
    `enrich` on a downloaded split module is exactly that path.

    With nothing to follow it creates the flat layout, because `derived/` is *tolerated* rather than
    canonical: a module only ends up split because somebody chose to split it.
    """
    found = resolve_sidecar(spec_dir, name)
    return found if found is not None else Path(spec_dir) / preferred_spelling(name)

deprecation_notice

deprecation_notice(
    path: Path, name: str, *, shown_as: str | None = None
) -> str | None

The warning for reading a deprecated spelling, or None when the name is current.

Actionable by construction, which is what the 0.6 amendment requires of a deprecation in a minor: the replacement exists, the old name is not mandatory, and the migration is git mv.

shown_as is what the message calls the file — a caller that knows the spec directory passes the path relative to it, so on a split tree the notice says which copy rather than a bare filename that could be either.

Source code in schema/src/just_dna_format/layout.py
def deprecation_notice(path: Path, name: str, *, shown_as: str | None = None) -> str | None:
    """The warning for reading a deprecated spelling, or `None` when the name is current.

    Actionable by construction, which is what the 0.6 amendment requires of a deprecation in a minor:
    the replacement exists, the old name is not mandatory, and the migration is `git mv`.

    `shown_as` is what the message calls the file — a caller that knows the spec directory passes the
    path relative to it, so on a split tree the notice says *which* copy rather than a bare filename
    that could be either.
    """
    if not is_deprecated_spelling(path.name):
        return None
    return (
        f"{shown_as or path.name} is the deprecated spelling of this table and will be removed at 1.0 "
        f"— rename it to {preferred_spelling(name)!r}. It is read exactly as before until then; the "
        f"compiled parquet and the manifest key keep their current names, which only a major may change."
    )

atomic_write_text

atomic_write_text(
    path: Path, text: str, *, encoding: str = "utf-8"
) -> Path

Write text to path so a reader ever only sees the whole file or the previous one.

The naive write_text truncates in place, so a process killed between the truncate and the last byte leaves a file that is syntactically valid and simply short — the one failure mode a table cannot report about itself, because a resolution table with 162 of 330 rows looks exactly like a module whose author resolved less (S66). Every sidecar here is read back and merged by the next run, so a short file is not merely lost work: it is silently believed.

Same directory for the temp file, because os.replace is only atomic within a filesystem and a tempdir may be on another one. fsync before the replace so the bytes are on disk and not merely in the page cache — a machine that loses power after the rename would otherwise expose an empty file where a complete one used to be, which is strictly worse than what we started with.

Source code in schema/src/just_dna_format/layout.py
def atomic_write_text(path: Path, text: str, *, encoding: str = "utf-8") -> Path:
    """Write `text` to `path` so a reader ever only sees the whole file or the previous one.

    The naive `write_text` truncates in place, so a process killed between the truncate and the last
    byte leaves a file that is **syntactically valid and simply short** — the one failure mode a table
    cannot report about itself, because a resolution table with 162 of 330 rows looks exactly like a
    module whose author resolved less (S66). Every sidecar here is read back and merged by the next
    run, so a short file is not merely lost work: it is silently believed.

    Same directory for the temp file, because `os.replace` is only atomic within a filesystem and a
    tempdir may be on another one. `fsync` before the replace so the bytes are on disk and not merely
    in the page cache — a machine that loses power after the rename would otherwise expose an empty
    file where a complete one used to be, which is strictly worse than what we started with.
    """
    with atomic_writer(path, encoding=encoding) as handle:
        handle.write(text)
    return Path(path)

atomic_writer

atomic_writer(
    path: Path,
    *,
    encoding: str = "utf-8",
    newline: str | None = None,
    before_commit: Callable[[], None] | None = None,
) -> Iterator[TextIO]

A text handle whose writes land at path only if the block completes — with one stated residual when before_commit is given, below.

The csv.DictWriter half of the same guarantee — the writers pass newline="" exactly as they do to open, so routing one through this changes the emitted bytes not at all.

On any exception the partial temp file is removed and path is left untouched, which is the property that makes this safe to wrap around a writer that can raise mid-table: today a row whose cell fails to serialize leaves the previous table truncated at that row.

before_commit runs after the temp file is closed and fsynced and before the rename (S98, RM231). It is for the write that has to land with this one or not at all — the licence row a pass records for the table it is writing. Eight passes used to write the table and then merge the row, so a placeholder licensing.csv split the two: a data table on disk, FAILED on the screen, and no licence record anywhere. Inside the commit, a failure on either side leaves nothing: the callback raising removes the temp, and a table that fails to serialize never reaches the callback. The one residual: the callback's own commit has returned by the time this rename runs, so a rename that fails leaves the callback's side effect without the table — conservative for a licence row, and the OSError raised then says so rather than leaving it to be found. Two files are two renames; no ordering makes them one.

Source code in schema/src/just_dna_format/layout.py
@contextmanager
def atomic_writer(
    path: Path,
    *,
    encoding: str = "utf-8",
    newline: str | None = None,
    before_commit: Callable[[], None] | None = None,
) -> Iterator[TextIO]:
    """A text handle whose writes land at `path` only if the block completes — with one stated residual
    when `before_commit` is given, below.

    The `csv.DictWriter` half of the same guarantee — the writers pass `newline=""` exactly as they do
    to `open`, so routing one through this changes the emitted bytes not at all.

    On any exception the partial temp file is removed and `path` is left untouched, which is the
    property that makes this safe to wrap around a writer that can raise mid-table: today a row whose
    cell fails to serialize leaves the previous table truncated at that row.

    `before_commit` runs after the temp file is closed and fsynced and **before** the rename (S98,
    RM231). It is for the write that has to land with this one or not at all — the licence row a
    pass records for the table it is writing. Eight passes used to write the table and then merge the
    row, so a placeholder `licensing.csv` split the two: a data table on disk, `FAILED` on the screen,
    and no licence record anywhere. Inside the commit, a failure on either side leaves nothing: the
    callback raising removes the temp, and a table that fails to serialize never reaches the callback.
    **The one residual**: the callback's own commit has returned by the time this rename runs, so a
    rename that fails leaves the callback's side effect without the table — conservative for a
    licence row, and the `OSError` raised then says so rather than leaving it to be found. Two files
    are two renames; no ordering makes them one.
    """
    path = Path(path)
    path.parent.mkdir(parents=True, exist_ok=True)
    # `delete=False` because the file has to survive `close()` to be renamed; the `finally` below is
    # what actually deletes it, on every path that does not reach the replace.
    tmp = tempfile.NamedTemporaryFile(  # noqa: SIM115 — the `with tmp:` below is the context manager
        mode="w",
        encoding=encoding,
        newline=newline,
        dir=path.parent,
        prefix=f".{path.name}.",
        suffix=".tmp",
        delete=False,
    )
    tmp_path = Path(tmp.name)
    try:
        with tmp:
            yield tmp
            tmp.flush()
            os.fsync(tmp.fileno())
        if before_commit is not None:
            before_commit()
        try:
            os.replace(tmp_path, path)
        except OSError as exc:
            if before_commit is None:
                raise
            # Two files are two renames, and no ordering makes them one. The callback's own commit
            # has returned, so what is on disk is its side effect without this table — conservative
            # (a licence row for data that never arrived), and SAID rather than left to be found.
            raise OSError(
                f"{path} was not committed ({exc}); a before_commit side effect for it has already "
                f"landed — re-run, or remove what the callback wrote"
            ) from exc
    finally:
        # Reached with the temp already gone on the success path; `missing_ok` is what makes the one
        # `finally` serve both, rather than an `except` that has to re-raise.
        tmp_path.unlink(missing_ok=True)