Skip to content

just_dna_enricher.licensing

just_dna_enricher.licensing

Data-source terms, and the declared-use gate (0.5).

The enricher is the only tier that fetches, so it is the only tier that knows where a fact came from and on what terms. This module owns both halves of that: the terms for each service it can reach, and the refusal that happens at the moment of acquisition.

Why the refusal lives here and not in the compiler. Under a data-usage policy, the terms are accepted when the data is taken. Refusing here means nothing is fetched; refusing at compile would only mean nothing is written, after the copy already exists on disk. The compiler still has a gate, but it is a different one — it enforces that whatever a module carries is accompanied by a declaration, computed purely from the injected sources.csv.

Why the constants sit beside the client that uses them, not in a shared registry. The pair (endpoint, terms) is one fact about one service; separating them is how they drift. TERMS here is the small residue that cannot be read from the payload: a service that ships no licence file with its data has to be described somewhere. Where a source does ship its terms — ClinPGx bundles a LICENSE.txt inside every archive — the pass reads them out of the same bytes it took the data from and overrides the constant, which makes the recorded licence provably contemporaneous with the recorded data rather than a lookup that was true once.

That distinction is not theoretical. Both halves of the static table went stale inside a single release: api.pharmgkb.org was retired on 2026-07-20, and CPIC's licence page moved to the ClinPGx policy when the two merged. A recorded license_sha256 turns the next such change into a finding.

LicenseRefusal

Bases: RuntimeError

Raised when a declared use is incompatible with a source's terms.

Fatal in both modes, unlike most enricher findings. The mode ladder grades how confident we are in a finding; this is not a finding about the data, it is a statement that the fetch is not permitted. best_effort means "resolve what you can", never "take what you may not".

SourceTerms dataclass

SourceTerms(
    source: str,
    license: str | None = None,
    license_url: str | None = None,
    attribution: str | None = None,
    notice: str | None = None,
    share_alike: bool | None = None,
    commercial_use: bool | None = None,
    redistribution: bool | None = None,
)

The terms a service publishes, as far as they can be established without the payload.

row

row(
    layer: str,
    *,
    declared_use: str,
    dataset: str | None = None,
    license_text: str | None = None,
) -> SourceRow

A SourceRow for this source at layer.

license_text, when the pass could read the terms out of the payload, is hashed into license_sha256 — pinning the terms to the same moment as the data.

Blank is absent, and the normalization lives here rather than at the four call sites. A licence file that exists and says nothing is not terms, and hashing it produces sha256:e3b0c442…b855 — a definite answer to a question nobody answered, indistinguishable from a real pin once it is in sources.csv. Whitespace-only counts as blank: the readers upstream guard on is_file(), so an empty file reached this far, and the tri-state rule says an unknown withholds rather than takes a default that is itself an answer. One normalizer because there are four sinks (two archive readers, the ClinPGx drafter, and a registry's status field) and a rule restated per caller is a rule three callers will drift from.

Source code in enricher/src/just_dna_enricher/licensing.py
def row(
    self,
    layer: str,
    *,
    declared_use: str,
    dataset: str | None = None,
    license_text: str | None = None,
) -> SourceRow:
    """A `SourceRow` for this source at `layer`.

    `license_text`, when the pass could read the terms out of the payload, is hashed into
    `license_sha256` — pinning the terms to the same moment as the data.

    **Blank is absent, and the normalization lives here rather than at the four call sites.** A
    licence file that exists and says nothing is not terms, and hashing it produces
    `sha256:e3b0c442…b855` — a definite answer to a question nobody answered, indistinguishable
    from a real pin once it is in `sources.csv`. Whitespace-only counts as blank: the readers
    upstream guard on `is_file()`, so an empty file reached this far, and the tri-state rule says
    an unknown withholds rather than takes a default that is itself an answer. One normalizer
    because there are four sinks (two archive readers, the ClinPGx drafter, and a registry's
    status field) and a rule restated per caller is a rule three callers will drift from.
    """
    pinned = license_text if (license_text or "").strip() else None
    return SourceRow(
        source=self.source,
        layer=layer,
        license=self.license,
        license_url=self.license_url,
        license_sha256=(
            "sha256:" + hashlib.sha256(pinned.encode("utf-8")).hexdigest() if pinned is not None else None
        ),
        attribution=self.attribution,
        notice=self.notice,
        share_alike=self.share_alike,
        commercial_use=self.commercial_use,
        redistribution=self.redistribution,
        declared_use=declared_use,
        dataset=dataset,
        fetched_at=now_utc_iso(),
    )

ArticleTerms dataclass

ArticleTerms(
    share_alike: bool | None = None,
    commercial_use: bool | None = None,
    redistribution: bool | None = None,
)

The three rights a cited article carries, read off its own licence (RM46).

A separate shape from SourceTerms because it answers a different question about a different thing. SourceTerms describes a service and produces a SourceRow; this describes one paper and lands on the LiteratureRow for it. There is deliberately no pubmed entry in TERMS_BY_SOURCE and there will not be one: a literature source's terms are per article, not per source — PubMed's metadata is one thing and the publisher's article is another, and Europe PMC's open subset spans CC-BY, CC-BY-NC and bronze. One pubmed row would be right for a module citing only ids and a false all-clear for one carrying a provenance_quote lifted from a CC-BY-NC article, since that quote is publisher text in the module's own annotation layer.

ScoreRights dataclass

ScoreRights(
    share_alike: bool | None = None,
    commercial_use: bool | None = None,
    redistribution: bool | None = None,
)

The three rights one PGS Catalog score's own license string grants.

A separate shape from SourceTerms for ArticleTerms' reason: this describes one licensed work inside a hosting service rather than the service. pgs_score_terms folds it back into a SourceTerms so the row is written by the one constructor every other pass uses.

article_terms

article_terms(license_name: str | None) -> ArticleTerms

The rights a licence string grants, or all-unknown when it names nothing this tier knows.

Case- and whitespace-insensitive, and tolerant of the CC-BY-NC spelling as well as Europe PMC's own cc by-nc, because the same value reaches this function from a hand-edited sidecar. Nothing is inferred from a substring: a licence this tier has not read is unknown, and an unknown right is withheld rather than guessed in either direction.

Source code in enricher/src/just_dna_enricher/licensing.py
def article_terms(license_name: str | None) -> ArticleTerms:
    """The rights a licence string grants, or all-unknown when it names nothing this tier knows.

    Case- and whitespace-insensitive, and tolerant of the `CC-BY-NC` spelling as well as Europe PMC's
    own `cc by-nc`, because the same value reaches this function from a hand-edited sidecar. Nothing
    is inferred from a substring: a licence this tier has not read is unknown, and an unknown right is
    withheld rather than guessed in either direction.
    """
    if not license_name or not license_name.strip():
        return ArticleTerms()
    key = " ".join(license_name.strip().lower().replace("_", "-").split())
    return ARTICLE_TERMS_BY_LICENSE.get(key) or ARTICLE_TERMS_BY_LICENSE.get(
        key.replace("cc-", "cc ", 1), ArticleTerms()
    )

pgs_license_class

pgs_license_class(
    license_text: str | None,
) -> tuple[str | None, str | None, ScoreRights]

(class name, the licence name to record, the rights) for one score's license string.

All three are empty for a string this tier has not read — unknown on every axis, withheld rather than guessed in either direction, and logged so a new class becomes visible instead of being absorbed into the permissive-looking default.

Source code in enricher/src/just_dna_enricher/licensing.py
def pgs_license_class(license_text: str | None) -> tuple[str | None, str | None, ScoreRights]:
    """`(class name, the licence name to record, the rights)` for one score's `license` string.

    All three are empty for a string this tier has not read — unknown on every axis, withheld rather
    than guessed in either direction, and logged so a new class becomes visible instead of being
    absorbed into the permissive-looking default.
    """
    text = (license_text or "").strip()
    if not text:
        return None, None, ScoreRights()
    for name, pattern, licence, rights in PGS_LICENSE_CLASSES:
        if pattern.search(text):
            return name, licence, rights
    logger.warning(
        "A PGS Catalog score states a licence this tier has not read (%r); its terms are recorded "
        "verbatim with every right left unknown. Unknown is not permission.",
        text,
    )
    return None, None, ScoreRights()

pgs_score_terms

pgs_score_terms(
    pgs_id: str, license_text: str | None
) -> SourceTerms

The terms for ONE score: PGS_TERMS as the floor, the score's own license on top.

The row's source namespaces the accession under the service (pgs_catalog:PGS000013) because SourceRow is keyed (source, layer) and the terms genuinely differ per score — one row per service would have to pick one of them, and picking the majority is picking the permissive answer for the minority. Nothing joins this value: PgsRow carries no source column, and the annotation layer is structurally exempt from the compiler's orphan check (@orphan-check-exempt), so the namespaced name costs nothing and says which score it is about.

The published string goes in notice; a short name goes in license. license_sha256 is computed over the verbatim text by the caller, and that is what pins the terms to the moment they were read — putting the sentence itself into license would push a paragraph into manifest.sources.licenses and into the compiler's declared-licence comparison, which is an equality test the author would then have no way to satisfy.

Source code in enricher/src/just_dna_enricher/licensing.py
def pgs_score_terms(pgs_id: str, license_text: str | None) -> SourceTerms:
    """The terms for ONE score: `PGS_TERMS` as the floor, the score's own `license` on top.

    The row's `source` namespaces the accession under the service (`pgs_catalog:PGS000013`) because
    `SourceRow` is keyed `(source, layer)` and the terms genuinely differ per score — one row per
    service would have to pick one of them, and picking the majority is picking the permissive answer
    for the minority. Nothing joins this value: `PgsRow` carries no `source` column, and the
    `annotation` layer is structurally exempt from the compiler's orphan check
    (`@orphan-check-exempt`), so the namespaced name costs nothing and says which score it is about.

    **The published string goes in `notice`; a short name goes in `license`.** `license_sha256` is
    computed over the verbatim text by the caller, and that is what pins the terms to the moment they
    were read — putting the sentence itself into `license` would push a paragraph into
    `manifest.sources.licenses` and into the compiler's declared-licence comparison, which is an
    equality test the author would then have no way to satisfy.
    """
    name, licence, rights = pgs_license_class(license_text)
    text = " ".join((license_text or "").split()) or None
    return SourceTerms(
        source=f"{PGS_TERMS.source}:{pgs_id}",
        license=licence,
        license_url=PGS_TERMS.license_url,
        attribution=PGS_TERMS.attribution,
        notice=(
            f"Terms for PGS Catalog score {pgs_id}, read from the score record's own `license` field "
            f"and classified as {name or 'unrecognised'}. The Catalog hosts scores licensed by their "
            f"authors, so these terms are the score's and not the service's. As published: "
            f"{text or '(the record states no licence)'}"
        ),
        share_alike=rights.share_alike,
        commercial_use=rights.commercial_use,
        redistribution=rights.redistribution,
    )

resolution_authority

resolution_authority(link: str | None) -> str | None

The licensed source a resolution link speaks for, or None when there is no external one.

Source code in enricher/src/just_dna_enricher/licensing.py
def resolution_authority(link: str | None) -> str | None:
    """The licensed source a resolution link speaks for, or `None` when there is no external one."""
    return RESOLUTION_AUTHORITY_BY_LINK.get(link or "")

record_source_terms

record_source_terms(
    source_names: Iterable[str],
    layer: str,
    spec_dir: Path,
    *,
    error: type[Exception],
    declared_use: str = "unstated",
    datasets: Mapping[str, str] | None = None,
    license_texts: Mapping[str, str] | None = None,
) -> list[SourceRow]

Record the terms of every licensed source a pass consulted, at layer.

A pass that consults a source must write its SourceRow — the rule clingen.py and then pgx_draft.py each shipped without. The compile gate and manifest.sources read sources.csv and nothing else, so a source that is only used is a source the module cannot account for. The three machine-fact passes (resolution, frequency, gene metrics) all skipped it, which is why VALID_SOURCE_LAYERS has reserved members nothing ever wrote.

The fact layers cannot taint a module — taints_commercial_use requires the annotation layer, because a coordinate or an AC/AN is a fact the source reports rather than expression it owns. For those passes what this records is attribution — which gnomAD, Ensembl and ClinVar all request and none of them enforces — and that is precisely the case the table exists to carry, not only prohibitions. declared_use defaults to unstated for the same reason: no fact-layer source here forbids sale, so those passes never have to ask the author for a declaration.

But annotation callers exist and this said they did not (RM222). civic_draft records at annotation with an explicit declared_use, so the paragraph above — which read "None of these layers can taint" about all of them — sent a reader to the wrong conclusion about whether a drafting pass can taint. It can: that layer is exactly the one the gate reads.

datasets maps a source name to the release its rows came from, for a caller that knows one. A drafting pass does; a fact pass usually does not, and an absent entry leaves dataset unset rather than guessing. It matters because SourceRow.dataset is what --verify-datasets compares and what withdraw_stale_dataset withdraws — a row without one puts the module outside the currency check altogether, which is where every CIViC-drafted module was.

license_texts is the same shape one axis over, added for the same reason (RM228): a pass that read the terms out of the payload can pin license_sha256 to the same moment as the data, and SourceTerms.row has always accepted one — this function simply had no way to pass it, so the two PGx drafters that do extract a licence file had to build their rows by hand and were therefore outside every other guarantee this function gives.

A name with no terms constant is skipped rather than guessed at: TERMS_BY_SOURCE is what this tier can state, and inventing a row for the rest would be worse than the compiler's honest warning that the terms are unrecorded. Existing rows are never clobbered (merge_sources_file), so a human's hand-written terms survive a re-run.

Source code in enricher/src/just_dna_enricher/licensing.py
def record_source_terms(
    source_names: Iterable[str],
    layer: str,
    spec_dir: Path,
    *,
    error: type[Exception],
    declared_use: str = "unstated",
    datasets: Mapping[str, str] | None = None,
    license_texts: Mapping[str, str] | None = None,
) -> list[SourceRow]:
    """Record the terms of every licensed source a pass consulted, at `layer`.

    **A pass that consults a source must write its `SourceRow`** — the rule `clingen.py` and then
    `pgx_draft.py` each shipped without. The compile gate and `manifest.sources` read `sources.csv` and
    nothing else, so a source that is only *used* is a source the module cannot account for. The three
    machine-fact passes (resolution, frequency, gene metrics) all skipped it, which is why
    `VALID_SOURCE_LAYERS` has reserved members nothing ever wrote.

    **The fact layers cannot taint a module** — `taints_commercial_use` requires the `annotation`
    layer, because a coordinate or an AC/AN is a fact the source *reports* rather than expression it
    *owns*. For those passes what this records is **attribution** — which gnomAD, Ensembl and ClinVar
    all request and none of them enforces — and that is precisely the case the table exists to carry,
    not only prohibitions. `declared_use` defaults to `unstated` for the same reason: no fact-layer
    source here forbids sale, so those passes never have to ask the author for a declaration.

    **But `annotation` callers exist and this said they did not** (RM222). `civic_draft` records at
    `annotation` with an explicit `declared_use`, so the paragraph above — which read *"None of these
    layers can taint"* about all of them — sent a reader to the wrong conclusion about whether a
    drafting pass can taint. It can: that layer is exactly the one the gate reads.

    `datasets` maps a source name to the release its rows came from, for a caller that knows one. A
    drafting pass does; a fact pass usually does not, and an absent entry leaves `dataset` unset rather
    than guessing. It matters because `SourceRow.dataset` is what `--verify-datasets` compares and what
    `withdraw_stale_dataset` withdraws — a row without one puts the module outside the currency check
    altogether, which is where every CIViC-drafted module was.

    `license_texts` is the same shape one axis over, added for the same reason (RM228): a pass that
    read the terms out of the payload can pin `license_sha256` to the same moment as the data, and
    `SourceTerms.row` has always accepted one — this function simply had no way to pass it, so the two
    PGx drafters that *do* extract a licence file had to build their rows by hand and were therefore
    outside every other guarantee this function gives.

    A name with no terms constant is skipped rather than guessed at: `TERMS_BY_SOURCE` is what this tier
    can state, and inventing a row for the rest would be worse than the compiler's honest warning that
    the terms are unrecorded. Existing rows are never clobbered (`merge_sources_file`), so a human's
    hand-written terms survive a re-run.
    """
    terms = [TERMS_BY_SOURCE[name] for name in sorted(set(source_names)) if name in TERMS_BY_SOURCE]
    if not terms:
        return []
    labels = datasets or {}
    texts = license_texts or {}
    return merge_sources_file(
        [
            t.row(
                layer,
                declared_use=declared_use,
                dataset=labels.get(t.source, ""),
                license_text=texts.get(t.source),
            )
            for t in terms
        ],
        spec_dir,
        error=error,
    )

check_declared_use

check_declared_use(
    terms: SourceTerms, declared_use: str
) -> str | None

Decide whether a fetch may proceed. Returns a skip reason, or raises, or returns None to go.

Three outcomes rather than two, and the middle one is the point:

  • raise — the source forbids sale and the caller declared commercial. A direct contradiction; refuse in both modes rather than take the data.
  • skip (a reason string) — either the caller declared nothing (unstated) and the source forbids sale, or the source's terms are unknown. Conservative by default: the tool must not assert a purpose on the user's behalf, and "we could not establish the terms" is not permission. Mirrors --offline making a pass a no-op with a warning rather than a failure.
  • None — proceed.
Source code in enricher/src/just_dna_enricher/licensing.py
def check_declared_use(terms: SourceTerms, declared_use: str) -> str | None:
    """Decide whether a fetch may proceed. Returns a skip reason, or raises, or returns None to go.

    Three outcomes rather than two, and the middle one is the point:

    * **raise** — the source forbids sale and the caller declared `commercial`. A direct
      contradiction; refuse in both modes rather than take the data.
    * **skip (a reason string)** — either the caller declared nothing (`unstated`) and the source
      forbids sale, or the source's terms are *unknown*. Conservative by default: the tool must not
      assert a purpose on the user's behalf, and "we could not establish the terms" is not permission.
      Mirrors `--offline` making a pass a no-op with a warning rather than a failure.
    * **None** — proceed.
    """
    # Through the shared checker, so a `-`/`_` slip is canonicalized here exactly as it is in the cell
    # this gate later reads: a caller passing `non-commercial` must not get a different verdict from
    # the same string written into the file.
    declared_use = check_vocab(declared_use, VALID_DECLARED_USE, "declared_use") or declared_use
    if terms.commercial_use is None:
        # RM289, from S99: the sentence used to stop at "could not be established", which read as a
        # constant nobody had filled in yet. Every source held at `None` here is a probed absence
        # whose reading is recorded in `notice`, so the skip quotes that reading and its URL rather
        # than leaving the reader to hunt for missing configuration. The leading phrase is pinned.
        established = terms.notice or "this tier records no reading of the source's terms."
        where = f" ({terms.license_url})" if terms.license_url else ""
        return (
            f"{terms.source}: terms could not be established, so the data is not used, and no declared "
            f"use changes that: there are no stated terms to judge a declaration against. What is "
            f"established: {established}{where} Unknown is not a finding that it is forbidden — it is "
            f"the absence of a finding either way."
        )
    if terms.commercial_use is True:
        return None
    if declared_use == "commercial":
        raise LicenseRefusal(
            f"{terms.source} is {terms.license} and its terms forbid offering the data for sale "
            f"({terms.license_url}). A commercial declaration cannot be reconciled with that, so "
            f"nothing was fetched. Use --use non-commercial if that describes your use."
        )
    if declared_use == "non_commercial":
        return None
    return (
        f"{terms.source} forbids sale and no use was declared, so it was skipped. Re-run with "
        f"--use non-commercial to record a declaration ({terms.license_url})."
    )

effective_declared_use

effective_declared_use(
    spec_dir: Path,
    terms: SourceTerms,
    declared_use: str,
    layer: str = "annotation",
) -> tuple[str, str | None]

The declaration a gate should judge, and where it came from (S105, RM252).

The caller's flag when it states one. Otherwise the module's own recorded declaration for this source at this layer — the row an earlier run wrote into the licence table, which is the same author, the same module, and the very file the compile gate keys on. pgx on a module drafted with --use non-commercial was saying "no use was declared" about CPIC while the declaration sat in the file it had just read, and asking the author to assert the same position a second time — which is the fabrication risk the draft's own message warns about, arriving at check time.

Three things this is not. It is not a default: with nothing recorded the answer is still unstated, and the tool asserts no purpose (@declared-use-third-axis). It is not per module: a declaration for CPIC says nothing about PharmVar, so a leg with no row still asks. And it is not a reading of the flag's meaning, which is unchanged — an explicit --use outranks the file, in both directions, and commercial against a forbidding source still refuses.

Returns (declaration, origin): origin is None when the flag decided, else the licence file's name, for the sentence that tells the author where the declaration was read from.

Source code in enricher/src/just_dna_enricher/licensing.py
def effective_declared_use(
    spec_dir: Path, terms: SourceTerms, declared_use: str, layer: str = "annotation"
) -> tuple[str, str | None]:
    """The declaration a gate should judge, and where it came from (S105, RM252).

    The caller's flag when it states one. Otherwise the module's **own recorded declaration** for this
    source at this layer — the row an earlier run wrote into the licence table, which is the same
    author, the same module, and the very file the compile gate keys on. `pgx` on a module drafted
    with `--use non-commercial` was saying *"no use was declared"* about CPIC while the declaration sat
    in the file it had just read, and asking the author to assert the same position a second time —
    which is the fabrication risk the draft's own message warns about, arriving at check time.

    Three things this is not. It is not a default: with nothing recorded the answer is still
    `unstated`, and the tool asserts no purpose (`@declared-use-third-axis`). It is not per module: a
    declaration for CPIC says nothing about PharmVar, so a leg with no row still asks. And it is not a
    reading of the *flag's* meaning, which is unchanged — an explicit `--use` outranks the file, in
    both directions, and `commercial` against a forbidding source still refuses.

    Returns `(declaration, origin)`: `origin` is `None` when the flag decided, else the licence file's
    name, for the sentence that tells the author where the declaration was read from.
    """
    stated = check_vocab(declared_use, VALID_DECLARED_USE, "declared_use") or declared_use
    if stated != "unstated":
        return stated, None
    for row in read_sources_file(spec_dir):
        recorded = row.declared_use or "unstated"
        if row.source == terms.source and row.layer == layer and recorded != "unstated":
            try:
                origin = sidecar_write_path(spec_dir, SOURCES_CSV).name
            except SidecarCollision:  # pragma: no cover - read_sources_file already returned [] here
                origin = SOURCES_CSV
            return recorded, origin
    return "unstated", None

write_sources_csv

write_sources_csv(
    rows: list[SourceRow], path: Path
) -> None

Write sources.csv in a fixed column order (normalized, like every reverse writer).

Source code in enricher/src/just_dna_enricher/licensing.py
def write_sources_csv(rows: list[SourceRow], path: Path) -> None:
    """Write `sources.csv` in a fixed column order (normalized, like every reverse writer)."""
    with atomic_writer(Path(path), newline="") as handle:
        writer = csv.DictWriter(handle, fieldnames=SOURCES_FIELDNAMES)
        writer.writeheader()
        for row in rows:
            dumped = row.model_dump()
            writer.writerow({name: _cell(dumped.get(name)) for name in SOURCES_FIELDNAMES})

merge_sources_csv

merge_sources_csv(
    rows: list[SourceRow],
    path: Path,
    existing: list[SourceRow],
) -> list[SourceRow]

Merge emitted rows into whatever is already recorded, never clobbering (as enrich() does).

Sorted by (source, layer) so the emitted order is deterministic (Principle 7).

Source code in enricher/src/just_dna_enricher/licensing.py
def merge_sources_csv(rows: list[SourceRow], path: Path, existing: list[SourceRow]) -> list[SourceRow]:
    """Merge emitted rows into whatever is already recorded, never clobbering (as `enrich()` does).

    Sorted by (source, layer) so the emitted order is deterministic (Principle 7).
    """
    merged: dict[tuple[str, str], SourceRow] = {(r.source, r.layer): r for r in existing}
    for row in rows:
        merged.setdefault((row.source, row.layer), row)
    out = [merged[key] for key in sorted(merged)]
    write_sources_csv(out, path)
    return out

sidecar_path

sidecar_path(
    spec_dir: Path, name: str, *, error: type[Exception]
) -> Path

Where a machine-written sidecar lives for this module — the copy it has, else a fresh one.

The enricher's single entry point into just_dna_format.layout, so a pass never joins a filename onto a spec directory itself. That matters twice over: the licence table has two accepted spellings (RM51), and any of them may sit under derived/ (RM49). A pass keeping its own literal would read the split copy and write the flat one, leaving the module with two — which is the collision, arrived at by following the documented workflow rather than by misuse.

SidecarCollision is re-raised as the caller's own error, so a pass still fails as itself rather than as a schema-tier ValueError nobody up the stack is catching.

Source code in enricher/src/just_dna_enricher/licensing.py
def sidecar_path(spec_dir: Path, name: str, *, error: type[Exception]) -> Path:
    """Where a machine-written sidecar lives for this module — the copy it has, else a fresh one.

    The enricher's single entry point into `just_dna_format.layout`, so a pass never joins a filename
    onto a spec directory itself. That matters twice over: the licence table has two accepted
    spellings (RM51), and any of them may sit under `derived/` (RM49). A pass keeping its own literal
    would read the split copy and write the flat one, leaving the module with two — which is the
    collision, arrived at by following the documented workflow rather than by misuse.

    `SidecarCollision` is re-raised as the caller's own error, so a pass still fails as itself rather
    than as a schema-tier `ValueError` nobody up the stack is catching.
    """
    try:
        return sidecar_write_path(spec_dir, name)
    except SidecarCollision as exc:
        raise error(str(exc)) from exc

sources_path

sources_path(
    spec_dir: Path, *, error: type[Exception]
) -> Path

Where this module's licence table lives — the file it already has, else the current spelling.

Nine passes used to write spec_dir / "sources.csv" by hand; they all come through here now.

Source code in enricher/src/just_dna_enricher/licensing.py
def sources_path(spec_dir: Path, *, error: type[Exception]) -> Path:
    """Where this module's licence table lives — the file it already has, else the current spelling.

    Nine passes used to write `spec_dir / "sources.csv"` by hand; they all come through here now.
    """
    return sidecar_path(spec_dir, SOURCES_CSV, error=error)

require_sources_file

require_sources_file(
    spec_dir: Path, *, error: type[Exception]
) -> list[SourceRow]

The module's licence table as it stands, [] when it has none — refusing one that does not load.

The read half of merge_sources_file, published on its own so a pass can run it before the fetch it is about to pay for (S98, RM231): a scaffold's placeholder row used to be found only at the merge, after a 47-minute query and after the data table was already on disk. Same refusal, same error type, moved to where it costs a second. The gentle counterpart for a reader is read_sources_file below, which withholds instead of refusing.

Source code in enricher/src/just_dna_enricher/licensing.py
def require_sources_file(spec_dir: Path, *, error: type[Exception]) -> list[SourceRow]:
    """The module's licence table as it stands, `[]` when it has none — refusing one that does not load.

    The read half of `merge_sources_file`, published on its own so a pass can run it **before** the
    fetch it is about to pay for (S98, RM231): a scaffold's placeholder row used to be found only at
    the merge, after a 47-minute query and after the data table was already on disk. Same refusal,
    same error type, moved to where it costs a second. The gentle counterpart for a *reader* is
    `read_sources_file` below, which withholds instead of refusing.
    """
    path = sources_path(spec_dir, error=error)
    if not path.exists():
        return []
    parsed, errors, _ = load_csv_rows(path, SourceRow, path.name)
    if errors:
        raise error(f"existing {path.name} is invalid: {errors[0]}")
    return parsed

merge_sources_file

merge_sources_file(
    rows: list[SourceRow],
    spec_dir: Path,
    *,
    error: type[Exception],
) -> list[SourceRow]

Read the module's licence table if it is there, merge rows in, and write it back.

The read-merge-write every terms-emitting pass performs, in one place: a pass that consulted a source has to record it, and each of them was otherwise growing its own copy of these nine lines. An unparseable existing file raises rather than being overwritten — merging into a table that did not load would silently drop the rows already recorded. error is the caller's own exception type, so a failure still surfaces as that pass's error rather than as a licensing one.

Takes the spec directory, not a path: resolving the filename here is what stops a caller naming a spelling the module does not use.

Source code in enricher/src/just_dna_enricher/licensing.py
def merge_sources_file(rows: list[SourceRow], spec_dir: Path, *, error: type[Exception]) -> list[SourceRow]:
    """Read the module's licence table if it is there, merge `rows` in, and write it back.

    The read-merge-write every terms-emitting pass performs, in one place: a pass that consulted a
    source has to record it, and each of them was otherwise growing its own copy of these nine lines.
    An unparseable existing file raises rather than being overwritten — merging into a table that did
    not load would silently drop the rows already recorded. `error` is the caller's own exception
    type, so a failure still surfaces as that pass's error rather than as a licensing one.

    Takes the **spec directory**, not a path: resolving the filename here is what stops a caller
    naming a spelling the module does not use.
    """
    return merge_sources_csv(
        rows, sources_path(spec_dir, error=error), require_sources_file(spec_dir, error=error)
    )

withdraw_stale_dataset

withdraw_stale_dataset(
    spec_dir: Path,
    source: str,
    layer: str,
    dataset: str | None,
    *,
    error: type[Exception],
) -> str | None

Blank a recorded dataset that this run's rows did not come from. Returns what it withdrew.

The one place anything overwrites a cell merge_sources_file would have kept, and it is narrow on purpose: merge_sources_csv is never-clobber so a curator's hand-written terms survive a re-run, which is right, and dataset inherited that protection at the moment RM4 made it load-bearing. A module drafted from one release and then widened from a newer one kept the older label — a licence row asserting a release half its rows did not come from, in the column a published manifest.sources carries and the clinical cross-check keys on.

It only ever withdraws, never re-labels, because the honest value for a module carrying rows from two releases is not the newer label either — one column cannot name two releases, so the answer is unknown and unknown is withheld (the house rule). That is also the safe direction for everything downstream: an empty dataset skips nothing, so the cross-check simply runs.

None when there was nothing to withdraw — no row, or a row already naming this run's release. The caller decides whether its rows even changed the module's provenance; a re-draft that added nothing must not reach this at all.

Source code in enricher/src/just_dna_enricher/licensing.py
def withdraw_stale_dataset(
    spec_dir: Path, source: str, layer: str, dataset: str | None, *, error: type[Exception]
) -> str | None:
    """Blank a recorded `dataset` that this run's rows did not come from. Returns what it withdrew.

    The one place anything overwrites a cell `merge_sources_file` would have kept, and it is narrow on
    purpose: `merge_sources_csv` is never-clobber so a curator's hand-written **terms** survive a
    re-run, which is right, and `dataset` inherited that protection at the moment RM4 made it
    load-bearing. A module drafted from one release and then widened from a newer one kept the older
    label — a licence row asserting a release half its rows did not come from, in the column a
    published `manifest.sources` carries and the clinical cross-check keys on.

    It only ever **withdraws**, never re-labels, because the honest value for a module carrying rows
    from two releases is not the newer label either — one column cannot name two releases, so the
    answer is unknown and unknown is withheld (the house rule). That is also the safe direction for
    everything downstream: an empty `dataset` skips nothing, so the cross-check simply runs.

    `None` when there was nothing to withdraw — no row, or a row already naming this run's release.
    The caller decides whether its rows even changed the module's provenance; a re-draft that added
    nothing must not reach this at all.
    """
    path = sources_path(spec_dir, error=error)
    if not path.exists():
        return None
    rows, errors, _ = load_csv_rows(path, SourceRow, path.name)
    if errors:
        raise error(f"existing {path.name} is invalid: {errors[0]}")
    recorded = next((r for r in rows if r.source == source and r.layer == layer), None)
    if recorded is None or (recorded.dataset or None) == (dataset or None):
        return None
    withdrawn = recorded.dataset
    recorded.dataset = None
    write_sources_csv(rows, path)
    return withdrawn

read_sources_file

read_sources_file(spec_dir: Path) -> list[SourceRow]

The module's licence rows as recorded, or [] when there are none that can be read.

The gentle counterpart to the strict load inside merge_sources_file, for a reader whose only power is to let a check be skipped — clinical.tautology_reason, which asks whether the licence row says these annotation rows were drafted from the snapshot the check is about to read (RM4).

Gentle deliberately, and in the same direction the rest of this codebase withholds: a table that could not be read has established nothing, so [] leaves every check running. A pass that writes must still fail loudly on an unreadable table — merging into one that did not load would drop rows already recorded — and merge_sources_file does.

Source code in enricher/src/just_dna_enricher/licensing.py
def read_sources_file(spec_dir: Path) -> list[SourceRow]:
    """The module's licence rows as recorded, or `[]` when there are none that can be read.

    The gentle counterpart to the strict load inside `merge_sources_file`, for a *reader* whose only
    power is to let a check be skipped — `clinical.tautology_reason`, which asks whether the licence
    row says these annotation rows were drafted from the snapshot the check is about to read (RM4).

    Gentle deliberately, and in the same direction the rest of this codebase withholds: a table that
    could not be read has established nothing, so `[]` leaves every check running. A pass that
    *writes* must still fail loudly on an unreadable table — merging into one that did not load would
    drop rows already recorded — and `merge_sources_file` does.
    """
    try:
        path = sidecar_write_path(spec_dir, SOURCES_CSV)
    except SidecarCollision as exc:
        logger.warning("Cannot read this module's licence table (%s); treating it as unrecorded.", exc)
        return []
    if not path.exists():
        return []
    rows, errors, _ = load_csv_rows(path, SourceRow, path.name)
    if errors:
        logger.warning("%s is invalid (%s); treating it as unrecorded.", path.name, errors[0])
        return []
    return rows

overlaid_input_rows

overlaid_input_rows(
    spec_dir: Path,
    table: str,
    rows: list,
    *,
    error: type[Exception],
) -> list

A derived table as the module asserts it, for a pass reading it as an INPUT (RM136).

The compiler applies overrides.csv before any check reads a row, which is the whole point: a check must report on what the module asserts. The enricher did not — its passes re-read the raw derived file — so an author who corrected a resolution.csv cell through the overlay went on being told the same finding by the tier that writes it, on every run, forever, with nothing saying their correction had been recorded and honoured one tier over.

This is not a second implementation of the overlay, and the distinction is the entry's own. RM136 refuses "teaching every enricher pass to apply the overlay" on the grounds that a second apply_overrides would drift on the normalization seam. This calls the apply_overrides, the one the compiler calls, through the one loader — so there is nothing to drift from.

INPUT reads only, and merge baselines must never come through here. A pass that reads its own output file to merge against it writes that file back; feeding it post-overlay rows would bake the correction into the derived table, and the enricher would be writing through the overlay — the author's answer restated as the tier's, which is RM83's standing refusal. The rule is the one the sidecar rules already state from the other side: read the file you write, and write what you read.

Errors from the overlay are raised as the caller's own exception rather than swallowed: an overlay that does not apply is a broken module, and a pass that quietly used the raw rows instead would be the silent-success shape this workspace keeps closing.

Source code in enricher/src/just_dna_enricher/licensing.py
def overlaid_input_rows(spec_dir: Path, table: str, rows: list, *, error: type[Exception]) -> list:
    """A derived table as the module **asserts** it, for a pass reading it as an INPUT (RM136).

    The compiler applies `overrides.csv` before any check reads a row, which is the whole point: a
    check must report on what the module asserts. The enricher did not — its passes re-read the raw
    derived file — so an author who corrected a `resolution.csv` cell through the overlay went on being
    told the same finding by the tier that writes it, on every run, forever, with nothing saying their
    correction had been recorded and honoured one tier over.

    **This is not a second implementation of the overlay, and the distinction is the entry's own.**
    RM136 refuses "teaching every enricher pass to apply the overlay" on the grounds that a second
    `apply_overrides` would drift on the normalization seam. This calls *the* `apply_overrides`, the
    one the compiler calls, through the one loader — so there is nothing to drift from.

    **INPUT reads only, and merge baselines must never come through here.** A pass that reads its own
    output file to merge against it writes that file back; feeding it post-overlay rows would bake the
    correction into the derived table, and the enricher would be writing *through* the overlay — the
    author's answer restated as the tier's, which is RM83's standing refusal. The rule is the one the
    sidecar rules already state from the other side: read the file you write, and write what you read.

    Errors from the overlay are raised as the caller's own exception rather than swallowed: an overlay
    that does not apply is a broken module, and a pass that quietly used the raw rows instead would be
    the silent-success shape this workspace keeps closing.
    """
    overrides, overlay_errors, _ = load_overlay(Path(spec_dir))
    if overlay_errors:
        raise error(f"overrides.csv is invalid: {overlay_errors[0]}")
    if not overrides or table not in VALID_OVERRIDE_TABLES:
        return rows
    applied, apply_errors, _ = apply_overrides(table, rows, overrides)
    if apply_errors:
        raise error(f"overrides.csv could not be applied to {table}: {apply_errors[0]}")
    return applied

overlay_answers

overlay_answers(
    spec_dir: Path, table: str
) -> set[tuple[str, str]]

The (subject, field) pairs this module's overlay has already answered for table (RM136).

Per field, and that is the decision. A finding is answered when the overlay updates the very cell the finding is about — so correcting a coordinate silences the coordinate check and leaves an unrelated clin_sig finding standing. Per row was the cheaper rule and was refused: an author correcting one cell would silence findings they never looked at, which is the silent-suppress hole the overlay's own design calls its worst case.

Only update counts. An insert supplies a row the source had no answer for, so there was no finding to answer; a suppress removes the row, and RM131 already reports that removal in its own right. Returns an empty set when the module has no overlay, which is every module today — a check that consults this must therefore behave exactly as before on one.

Read-only. Nothing here writes, and a malformed overlay is the caller's problem to raise on through overlaid_input_rows; this answers set() rather than guessing.

Source code in enricher/src/just_dna_enricher/licensing.py
def overlay_answers(spec_dir: Path, table: str) -> set[tuple[str, str]]:
    """The `(subject, field)` pairs this module's overlay has already answered for `table` (RM136).

    **Per field, and that is the decision.** A finding is answered when the overlay `update`s the very
    cell the finding is about — so correcting a coordinate silences the coordinate check and leaves an
    unrelated `clin_sig` finding standing. Per *row* was the cheaper rule and was refused: an author
    correcting one cell would silence findings they never looked at, which is the silent-suppress hole
    the overlay's own design calls its worst case.

    Only `update` counts. An `insert` supplies a row the source had no answer for, so there was no
    finding to answer; a `suppress` removes the row, and RM131 already reports that removal in its own
    right. Returns an empty set when the module has no overlay, which is every module today — a check
    that consults this must therefore behave exactly as before on one.

    Read-only. Nothing here writes, and a malformed overlay is the caller's problem to raise on
    through `overlaid_input_rows`; this answers `set()` rather than guessing.
    """
    overrides, overlay_errors, _ = load_overlay(Path(spec_dir))
    if overlay_errors:
        return set()
    return {
        (row.subject, row.field or "")
        for row in overrides
        if row.table == table and row.operation == "update" and row.field
    }

overlay_answered_subjects

overlay_answered_subjects(
    spec_dir: Path, table: str
) -> list[tuple[str, str]]

The (subject, member) pairs this module's overlay answers for table (RM151).

Every operation and every field, and that is the rule rather than an omission. What this feeds is the staleness question — has the value a recorded judgement was written about moved since? — and the judgement is the reason, which the model makes mandatory on every overlay row whatever it does. An author who suppresses a contested subject has reasoned about the same values as one who updates its call, so a per-field rule would have to name a field the reason does not live in.

It is deliberately not overlay_answers, whose per-field rule is the right one for the opposite direction. That one decides whether a finding may be silenced, so it insists the overlay touched the very cell the finding is about — anything looser would silence findings the author never looked at, the overlay design's own worst case. This one raises a finding, and a finding raised too widely costs a reader one line rather than hiding one.

An empty member is group-scoped and is returned as ""; the caller decides what a group means for its table, because only it knows the members. Ordered by first appearance in the overlay, deduplicated, so a caller's message is deterministic.

Read-only, and empty for a module whose overlay does not parse — a malformed overlay is raised on by overlaid_input_rows, and answering [] here rather than guessing keeps one loader owning that diagnosis.

Source code in enricher/src/just_dna_enricher/licensing.py
def overlay_answered_subjects(spec_dir: Path, table: str) -> list[tuple[str, str]]:
    """The `(subject, member)` pairs this module's overlay answers for `table` (RM151).

    **Every operation and every field, and that is the rule rather than an omission.** What this
    feeds is the staleness question — has the value a recorded judgement was written *about* moved
    since? — and the judgement is the `reason`, which the model makes mandatory on every overlay row
    whatever it does. An author who suppresses a contested subject has reasoned about the same
    values as one who updates its call, so a per-field rule would have to name a field the reason
    does not live in.

    **It is deliberately not `overlay_answers`, whose per-field rule is the right one for the
    opposite direction.** That one decides whether a finding may be *silenced*, so it insists the
    overlay touched the very cell the finding is about — anything looser would silence findings the
    author never looked at, the overlay design's own worst case. This one *raises* a finding, and a
    finding raised too widely costs a reader one line rather than hiding one.

    An empty `member` is group-scoped and is returned as `""`; the caller decides what a group means
    for its table, because only it knows the members. Ordered by first appearance in the overlay,
    deduplicated, so a caller's message is deterministic.

    Read-only, and empty for a module whose overlay does not parse — a malformed overlay is raised on
    by `overlaid_input_rows`, and answering `[]` here rather than guessing keeps one loader owning
    that diagnosis.
    """
    overrides, overlay_errors, _ = load_overlay(Path(spec_dir))
    if overlay_errors:
        return []
    seen: list[tuple[str, str]] = []
    for row in overrides:
        if row.table != table:
            continue
        pair = (row.subject, row.member or "")
        if pair not in seen:
            seen.append(pair)
    return seen