Where a module's machine-written sidecars live, and what they may be called.
The compiler reads these files, the enricher writes them, a publisher uploads them and a registry
splits them back apart — four parties, one layout. just_dna_enricher.locations exists because every
disagreement about a snapshot's layout so far has been silent, and this is the same class of fact one
tier down, so it lives in the schema tier where both consumers import it rather than copy it.
Nothing here fetches, parses or validates. It is the standard library and two tuples of names, which is
why it can sit in the dependency-light tier at all — pathlib for the places, and since S66 os plus
tempfile for the writing, because how a sidecar is put on disk turned out to be the same class of
fact as where: four parties agreeing on the name is worth nothing if a killed run leaves half a table
under it.
Scope is the machine-written sidecars only — resolution.csv and the fact tables. The authored DSL
(module_spec.yaml, variants.csv, studies.csv, the table kinds) has exactly one legal name in
exactly one legal place, and that asymmetry is deliberate: a second legal home for variants.csv means
a module can carry two copies and the one the compiler ignores is invisible, which is the
silent-success shape this codebase treats as the worst kind of mistake. Only the machine-written tables
move, because only they have a machine that knows where to put them.
Bases: ValueError
Two files claim to be the same sidecar, and neither may be silently preferred.
Not a merge and not newest-wins. These tables are fact-hashed and human-overridable — a curator
edits a row the enricher wrote and that edit is the point — so two copies are two legitimate claims
and picking one discards somebody's work without saying so. The only honest answer names both paths
and stops.
sidecar_key(name: str) -> str
The table key behind any of its spellings — sources.csv for licensing.csv — else name.
The key is the spelling sources.parquet and manifest.sources keep, and it is what the map is
keyed on. A caller holding bytes and a filename — a tar member, an upload part — has the preferred
spelling and not the key, and until S96 the helpers below read that as a table they had never heard
of, answering the one-tuple ('licensing.csv',): sidecar_write_path then created the preferred
copy beside the deprecated one, which is the collision the docstring below says it exists to
prevent. Either spelling is a key now.
Source code in schema/src/just_dna_format/layout.py
| def sidecar_key(name: str) -> str:
"""The table key behind any of its spellings — `sources.csv` for `licensing.csv` — else `name`.
The key is the spelling `sources.parquet` and `manifest.sources` keep, and it is what the map is
keyed on. A caller holding bytes and a filename — a tar member, an upload part — has the *preferred*
spelling and not the key, and until S96 the helpers below read that as a table they had never heard
of, answering the one-tuple `('licensing.csv',)`: `sidecar_write_path` then created the preferred
copy beside the deprecated one, which is the collision the docstring below says it exists to
prevent. Either spelling is a key now.
"""
return _KEY_FOR_SPELLING.get(name, name)
|
sidecar_spellings(name: str) -> tuple[str, ...]
Every accepted filename for name, deprecated first, preferred last.
name may be the table key or any of its spellings — sidecar_spellings("licensing.csv")
answers the same pair as sidecar_spellings("sources.csv"). A name the map does not know has
exactly one spelling, its own.
Source code in schema/src/just_dna_format/layout.py
| def sidecar_spellings(name: str) -> tuple[str, ...]:
"""Every accepted filename for `name`, deprecated first, preferred last.
`name` may be the table key **or any of its spellings** — `sidecar_spellings("licensing.csv")`
answers the same pair as `sidecar_spellings("sources.csv")`. A name the map does not know has
exactly one spelling, its own.
"""
return SIDECAR_SPELLINGS.get(sidecar_key(name), (name,))
|
preferred_spelling(name: str) -> str
The spelling a fresh file should be created under.
Source code in schema/src/just_dna_format/layout.py
| def preferred_spelling(name: str) -> str:
"""The spelling a fresh file should be created under."""
return sidecar_spellings(name)[-1]
|
is_deprecated_spelling(filename: str) -> bool
Whether a filename is a spelling that still reads but should no longer be written.
Source code in schema/src/just_dna_format/layout.py
| def is_deprecated_spelling(filename: str) -> bool:
"""Whether a filename is a spelling that still reads but should no longer be written."""
return filename in DEPRECATED_SPELLINGS
|
sidecar_candidates(spec_dir: Path, name: str) -> list[Path]
Every place name may legally be — each spelling, at the root and under derived/.
Order is the tie-break nothing else can supply, so it is fixed rather than incidental: a caller
that wants "the one that exists" gets a deterministic answer, and a caller that wants "where do I
create it" reads the same list from the front. The root comes first because it is the layout every
existing module and every reverse_module output uses.
Source code in schema/src/just_dna_format/layout.py
| def sidecar_candidates(spec_dir: Path, name: str) -> list[Path]:
"""Every place `name` may legally be — each spelling, at the root and under `derived/`.
Order is the tie-break nothing else can supply, so it is fixed rather than incidental: a caller
that wants "the one that exists" gets a deterministic answer, and a caller that wants "where do I
create it" reads the same list from the front. The root comes first because it is the layout every
existing module and every `reverse_module` output uses.
"""
root = Path(spec_dir)
return [
root / directory / spelling
for directory in ("", DERIVED_SUBDIR)
for spelling in sidecar_spellings(name)
]
|
resolve_sidecar(spec_dir: Path, name: str) -> Path | None
The single existing copy of a sidecar, or None when the module carries none.
Raises SidecarCollision when more than one candidate exists. That is an error rather than a
warning because there is no correct silent behaviour available — see the exception's own docstring.
Source code in schema/src/just_dna_format/layout.py
| def resolve_sidecar(spec_dir: Path, name: str) -> Path | None:
"""The single existing copy of a sidecar, or `None` when the module carries none.
Raises `SidecarCollision` when more than one candidate exists. That is an error rather than a
warning because there is no correct silent behaviour available — see the exception's own docstring.
"""
present = [path for path in sidecar_candidates(spec_dir, name) if path.is_file()]
if not present:
return None
if len(present) > 1:
listed = " and ".join(str(path) for path in present)
raise SidecarCollision(
f"{listed} are the same table in two places, and both are present. These tables are "
f"fact-hashed and may be edited by hand, so two copies are two claims and neither can be "
f"preferred without discarding the other — keep one and delete the other. "
f"{preferred_spelling(name)!r} is the spelling to keep; either the spec root or "
f"{DERIVED_SUBDIR}/ is a fine place for it."
)
return present[0]
|
sidecar_relative_names(name: str) -> list[str]
Every legal spelling-and-location of name, as paths relative to the spec directory.
For hashing rather than reading: integrity.file_entries skips the names that are not there, so
handing it this list records whichever one the module actually carries — and FileEntry.name then
carries the relative path, which is what a registry re-splitting a downloaded tree needs.
Source code in schema/src/just_dna_format/layout.py
| def sidecar_relative_names(name: str) -> list[str]:
"""Every legal spelling-and-location of `name`, as paths relative to the spec directory.
For hashing rather than reading: `integrity.file_entries` skips the names that are not there, so
handing it this list records whichever one the module actually carries — and `FileEntry.name` then
carries the relative path, which is what a registry re-splitting a downloaded tree needs.
"""
return [str(path) for path in sidecar_candidates(Path(), name)]
|
sidecar_write_path(spec_dir: Path, name: str) -> Path
Where a pass should write a sidecar: the copy that exists, else the preferred spelling.
Write to the file you read. A pass that always created the preferred spelling at the root
would, on a module carrying the deprecated one or a derived/ tree, leave two copies behind — the
collision above, produced by following the documented workflow rather than by misusing it. Running
enrich on a downloaded split module is exactly that path.
With nothing to follow it creates the flat layout, because derived/ is tolerated rather than
canonical: a module only ends up split because somebody chose to split it.
Source code in schema/src/just_dna_format/layout.py
| def sidecar_write_path(spec_dir: Path, name: str) -> Path:
"""Where a pass should write a sidecar: the copy that exists, else the preferred spelling.
**Write to the file you read.** A pass that always created the preferred spelling at the root
would, on a module carrying the deprecated one or a `derived/` tree, leave two copies behind — the
collision above, produced by following the documented workflow rather than by misusing it. Running
`enrich` on a downloaded split module is exactly that path.
With nothing to follow it creates the flat layout, because `derived/` is *tolerated* rather than
canonical: a module only ends up split because somebody chose to split it.
"""
found = resolve_sidecar(spec_dir, name)
return found if found is not None else Path(spec_dir) / preferred_spelling(name)
|
deprecation_notice(
path: Path, name: str, *, shown_as: str | None = None
) -> str | None
The warning for reading a deprecated spelling, or None when the name is current.
Actionable by construction, which is what the 0.6 amendment requires of a deprecation in a minor:
the replacement exists, the old name is not mandatory, and the migration is git mv.
shown_as is what the message calls the file — a caller that knows the spec directory passes the
path relative to it, so on a split tree the notice says which copy rather than a bare filename
that could be either.
Source code in schema/src/just_dna_format/layout.py
| def deprecation_notice(path: Path, name: str, *, shown_as: str | None = None) -> str | None:
"""The warning for reading a deprecated spelling, or `None` when the name is current.
Actionable by construction, which is what the 0.6 amendment requires of a deprecation in a minor:
the replacement exists, the old name is not mandatory, and the migration is `git mv`.
`shown_as` is what the message calls the file — a caller that knows the spec directory passes the
path relative to it, so on a split tree the notice says *which* copy rather than a bare filename
that could be either.
"""
if not is_deprecated_spelling(path.name):
return None
return (
f"{shown_as or path.name} is the deprecated spelling of this table and will be removed at 1.0 "
f"— rename it to {preferred_spelling(name)!r}. It is read exactly as before until then; the "
f"compiled parquet and the manifest key keep their current names, which only a major may change."
)
|
atomic_write_text
atomic_write_text(
path: Path, text: str, *, encoding: str = "utf-8"
) -> Path
Write text to path so a reader ever only sees the whole file or the previous one.
The naive write_text truncates in place, so a process killed between the truncate and the last
byte leaves a file that is syntactically valid and simply short — the one failure mode a table
cannot report about itself, because a resolution table with 162 of 330 rows looks exactly like a
module whose author resolved less (S66). Every sidecar here is read back and merged by the next
run, so a short file is not merely lost work: it is silently believed.
Same directory for the temp file, because os.replace is only atomic within a filesystem and a
tempdir may be on another one. fsync before the replace so the bytes are on disk and not merely
in the page cache — a machine that loses power after the rename would otherwise expose an empty
file where a complete one used to be, which is strictly worse than what we started with.
Source code in schema/src/just_dna_format/layout.py
| def atomic_write_text(path: Path, text: str, *, encoding: str = "utf-8") -> Path:
"""Write `text` to `path` so a reader ever only sees the whole file or the previous one.
The naive `write_text` truncates in place, so a process killed between the truncate and the last
byte leaves a file that is **syntactically valid and simply short** — the one failure mode a table
cannot report about itself, because a resolution table with 162 of 330 rows looks exactly like a
module whose author resolved less (S66). Every sidecar here is read back and merged by the next
run, so a short file is not merely lost work: it is silently believed.
Same directory for the temp file, because `os.replace` is only atomic within a filesystem and a
tempdir may be on another one. `fsync` before the replace so the bytes are on disk and not merely
in the page cache — a machine that loses power after the rename would otherwise expose an empty
file where a complete one used to be, which is strictly worse than what we started with.
"""
with atomic_writer(path, encoding=encoding) as handle:
handle.write(text)
return Path(path)
|
atomic_writer(
path: Path,
*,
encoding: str = "utf-8",
newline: str | None = None,
before_commit: Callable[[], None] | None = None,
) -> Iterator[TextIO]
A text handle whose writes land at path only if the block completes — with one stated residual
when before_commit is given, below.
The csv.DictWriter half of the same guarantee — the writers pass newline="" exactly as they do
to open, so routing one through this changes the emitted bytes not at all.
On any exception the partial temp file is removed and path is left untouched, which is the
property that makes this safe to wrap around a writer that can raise mid-table: today a row whose
cell fails to serialize leaves the previous table truncated at that row.
before_commit runs after the temp file is closed and fsynced and before the rename (S98,
RM231). It is for the write that has to land with this one or not at all — the licence row a
pass records for the table it is writing. Eight passes used to write the table and then merge the
row, so a placeholder licensing.csv split the two: a data table on disk, FAILED on the screen,
and no licence record anywhere. Inside the commit, a failure on either side leaves nothing: the
callback raising removes the temp, and a table that fails to serialize never reaches the callback.
The one residual: the callback's own commit has returned by the time this rename runs, so a
rename that fails leaves the callback's side effect without the table — conservative for a
licence row, and the OSError raised then says so rather than leaving it to be found. Two files
are two renames; no ordering makes them one.
Source code in schema/src/just_dna_format/layout.py
| @contextmanager
def atomic_writer(
path: Path,
*,
encoding: str = "utf-8",
newline: str | None = None,
before_commit: Callable[[], None] | None = None,
) -> Iterator[TextIO]:
"""A text handle whose writes land at `path` only if the block completes — with one stated residual
when `before_commit` is given, below.
The `csv.DictWriter` half of the same guarantee — the writers pass `newline=""` exactly as they do
to `open`, so routing one through this changes the emitted bytes not at all.
On any exception the partial temp file is removed and `path` is left untouched, which is the
property that makes this safe to wrap around a writer that can raise mid-table: today a row whose
cell fails to serialize leaves the previous table truncated at that row.
`before_commit` runs after the temp file is closed and fsynced and **before** the rename (S98,
RM231). It is for the write that has to land with this one or not at all — the licence row a
pass records for the table it is writing. Eight passes used to write the table and then merge the
row, so a placeholder `licensing.csv` split the two: a data table on disk, `FAILED` on the screen,
and no licence record anywhere. Inside the commit, a failure on either side leaves nothing: the
callback raising removes the temp, and a table that fails to serialize never reaches the callback.
**The one residual**: the callback's own commit has returned by the time this rename runs, so a
rename that fails leaves the callback's side effect without the table — conservative for a
licence row, and the `OSError` raised then says so rather than leaving it to be found. Two files
are two renames; no ordering makes them one.
"""
path = Path(path)
path.parent.mkdir(parents=True, exist_ok=True)
# `delete=False` because the file has to survive `close()` to be renamed; the `finally` below is
# what actually deletes it, on every path that does not reach the replace.
tmp = tempfile.NamedTemporaryFile( # noqa: SIM115 — the `with tmp:` below is the context manager
mode="w",
encoding=encoding,
newline=newline,
dir=path.parent,
prefix=f".{path.name}.",
suffix=".tmp",
delete=False,
)
tmp_path = Path(tmp.name)
try:
with tmp:
yield tmp
tmp.flush()
os.fsync(tmp.fileno())
if before_commit is not None:
before_commit()
try:
os.replace(tmp_path, path)
except OSError as exc:
if before_commit is None:
raise
# Two files are two renames, and no ordering makes them one. The callback's own commit
# has returned, so what is on disk is its side effect without this table — conservative
# (a licence row for data that never arrived), and SAID rather than left to be found.
raise OSError(
f"{path} was not committed ({exc}); a before_commit side effect for it has already "
f"landed — re-run, or remove what the callback wrote"
) from exc
finally:
# Reached with the temp already gone on the success path; `missing_ok` is what makes the one
# `finally` serve both, rather than an `except` that has to re-raise.
tmp_path.unlink(missing_ok=True)
|