just_dna_enricher.licensing¶
just_dna_enricher.licensing ¶
Data-source terms, and the declared-use gate (0.5).
The enricher is the only tier that fetches, so it is the only tier that knows where a fact came from and on what terms. This module owns both halves of that: the terms for each service it can reach, and the refusal that happens at the moment of acquisition.
Why the refusal lives here and not in the compiler. Under a data-usage policy, the terms are
accepted when the data is taken. Refusing here means nothing is fetched; refusing at compile would
only mean nothing is written, after the copy already exists on disk. The compiler still has a gate,
but it is a different one — it enforces that whatever a module carries is accompanied by a
declaration, computed purely from the injected sources.csv.
Why the constants sit beside the client that uses them, not in a shared registry. The pair
(endpoint, terms) is one fact about one service; separating them is how they drift. TERMS here is
the small residue that cannot be read from the payload: a service that ships no licence file with its
data has to be described somewhere. Where a source does ship its terms — ClinPGx bundles a
LICENSE.txt inside every archive — the pass reads them out of the same bytes it took the data from
and overrides the constant, which makes the recorded licence provably contemporaneous with the
recorded data rather than a lookup that was true once.
That distinction is not theoretical. Both halves of the static table went stale inside a single
release: api.pharmgkb.org was retired on 2026-07-20, and CPIC's licence page moved to the ClinPGx
policy when the two merged. A recorded license_sha256 turns the next such change into a finding.
LicenseRefusal ¶
Bases: RuntimeError
Raised when a declared use is incompatible with a source's terms.
Fatal in both modes, unlike most enricher findings. The mode ladder grades how confident we
are in a finding; this is not a finding about the data, it is a statement that the fetch is not
permitted. best_effort means "resolve what you can", never "take what you may not".
SourceTerms
dataclass
¶
SourceTerms(
source: str,
license: str | None = None,
license_url: str | None = None,
attribution: str | None = None,
notice: str | None = None,
share_alike: bool | None = None,
commercial_use: bool | None = None,
redistribution: bool | None = None,
)
The terms a service publishes, as far as they can be established without the payload.
row ¶
row(
layer: str,
*,
declared_use: str,
dataset: str | None = None,
license_text: str | None = None,
) -> SourceRow
A SourceRow for this source at layer.
license_text, when the pass could read the terms out of the payload, is hashed into
license_sha256 — pinning the terms to the same moment as the data.
Blank is absent, and the normalization lives here rather than at the four call sites. A
licence file that exists and says nothing is not terms, and hashing it produces
sha256:e3b0c442…b855 — a definite answer to a question nobody answered, indistinguishable
from a real pin once it is in sources.csv. Whitespace-only counts as blank: the readers
upstream guard on is_file(), so an empty file reached this far, and the tri-state rule says
an unknown withholds rather than takes a default that is itself an answer. One normalizer
because there are four sinks (two archive readers, the ClinPGx drafter, and a registry's
status field) and a rule restated per caller is a rule three callers will drift from.
Source code in enricher/src/just_dna_enricher/licensing.py
ArticleTerms
dataclass
¶
ArticleTerms(
share_alike: bool | None = None,
commercial_use: bool | None = None,
redistribution: bool | None = None,
)
The three rights a cited article carries, read off its own licence (RM46).
A separate shape from SourceTerms because it answers a different question about a different
thing. SourceTerms describes a service and produces a SourceRow; this describes one paper
and lands on the LiteratureRow for it. There is deliberately no pubmed entry in
TERMS_BY_SOURCE and there will not be one: a literature source's terms are per article, not
per source — PubMed's metadata is one thing and the publisher's article is another, and Europe
PMC's open subset spans CC-BY, CC-BY-NC and bronze. One pubmed row would be right for a module
citing only ids and a false all-clear for one carrying a provenance_quote lifted from a
CC-BY-NC article, since that quote is publisher text in the module's own annotation layer.
ScoreRights
dataclass
¶
ScoreRights(
share_alike: bool | None = None,
commercial_use: bool | None = None,
redistribution: bool | None = None,
)
The three rights one PGS Catalog score's own license string grants.
A separate shape from SourceTerms for ArticleTerms' reason: this describes one licensed work
inside a hosting service rather than the service. pgs_score_terms folds it back into a
SourceTerms so the row is written by the one constructor every other pass uses.
article_terms ¶
The rights a licence string grants, or all-unknown when it names nothing this tier knows.
Case- and whitespace-insensitive, and tolerant of the CC-BY-NC spelling as well as Europe PMC's
own cc by-nc, because the same value reaches this function from a hand-edited sidecar. Nothing
is inferred from a substring: a licence this tier has not read is unknown, and an unknown right is
withheld rather than guessed in either direction.
Source code in enricher/src/just_dna_enricher/licensing.py
pgs_license_class ¶
(class name, the licence name to record, the rights) for one score's license string.
All three are empty for a string this tier has not read — unknown on every axis, withheld rather than guessed in either direction, and logged so a new class becomes visible instead of being absorbed into the permissive-looking default.
Source code in enricher/src/just_dna_enricher/licensing.py
pgs_score_terms ¶
The terms for ONE score: PGS_TERMS as the floor, the score's own license on top.
The row's source namespaces the accession under the service (pgs_catalog:PGS000013) because
SourceRow is keyed (source, layer) and the terms genuinely differ per score — one row per
service would have to pick one of them, and picking the majority is picking the permissive answer
for the minority. Nothing joins this value: PgsRow carries no source column, and the
annotation layer is structurally exempt from the compiler's orphan check
(@orphan-check-exempt), so the namespaced name costs nothing and says which score it is about.
The published string goes in notice; a short name goes in license. license_sha256 is
computed over the verbatim text by the caller, and that is what pins the terms to the moment they
were read — putting the sentence itself into license would push a paragraph into
manifest.sources.licenses and into the compiler's declared-licence comparison, which is an
equality test the author would then have no way to satisfy.
Source code in enricher/src/just_dna_enricher/licensing.py
resolution_authority ¶
The licensed source a resolution link speaks for, or None when there is no external one.
record_source_terms ¶
record_source_terms(
source_names: Iterable[str],
layer: str,
spec_dir: Path,
*,
error: type[Exception],
declared_use: str = "unstated",
datasets: Mapping[str, str] | None = None,
license_texts: Mapping[str, str] | None = None,
) -> list[SourceRow]
Record the terms of every licensed source a pass consulted, at layer.
A pass that consults a source must write its SourceRow — the rule clingen.py and then
pgx_draft.py each shipped without. The compile gate and manifest.sources read sources.csv and
nothing else, so a source that is only used is a source the module cannot account for. The three
machine-fact passes (resolution, frequency, gene metrics) all skipped it, which is why
VALID_SOURCE_LAYERS has reserved members nothing ever wrote.
The fact layers cannot taint a module — taints_commercial_use requires the annotation
layer, because a coordinate or an AC/AN is a fact the source reports rather than expression it
owns. For those passes what this records is attribution — which gnomAD, Ensembl and ClinVar
all request and none of them enforces — and that is precisely the case the table exists to carry,
not only prohibitions. declared_use defaults to unstated for the same reason: no fact-layer
source here forbids sale, so those passes never have to ask the author for a declaration.
But annotation callers exist and this said they did not (RM222). civic_draft records at
annotation with an explicit declared_use, so the paragraph above — which read "None of these
layers can taint" about all of them — sent a reader to the wrong conclusion about whether a
drafting pass can taint. It can: that layer is exactly the one the gate reads.
datasets maps a source name to the release its rows came from, for a caller that knows one. A
drafting pass does; a fact pass usually does not, and an absent entry leaves dataset unset rather
than guessing. It matters because SourceRow.dataset is what --verify-datasets compares and what
withdraw_stale_dataset withdraws — a row without one puts the module outside the currency check
altogether, which is where every CIViC-drafted module was.
license_texts is the same shape one axis over, added for the same reason (RM228): a pass that
read the terms out of the payload can pin license_sha256 to the same moment as the data, and
SourceTerms.row has always accepted one — this function simply had no way to pass it, so the two
PGx drafters that do extract a licence file had to build their rows by hand and were therefore
outside every other guarantee this function gives.
A name with no terms constant is skipped rather than guessed at: TERMS_BY_SOURCE is what this tier
can state, and inventing a row for the rest would be worse than the compiler's honest warning that
the terms are unrecorded. Existing rows are never clobbered (merge_sources_file), so a human's
hand-written terms survive a re-run.
Source code in enricher/src/just_dna_enricher/licensing.py
check_declared_use ¶
Decide whether a fetch may proceed. Returns a skip reason, or raises, or returns None to go.
Three outcomes rather than two, and the middle one is the point:
- raise — the source forbids sale and the caller declared
commercial. A direct contradiction; refuse in both modes rather than take the data. - skip (a reason string) — either the caller declared nothing (
unstated) and the source forbids sale, or the source's terms are unknown. Conservative by default: the tool must not assert a purpose on the user's behalf, and "we could not establish the terms" is not permission. Mirrors--offlinemaking a pass a no-op with a warning rather than a failure. - None — proceed.
Source code in enricher/src/just_dna_enricher/licensing.py
effective_declared_use ¶
effective_declared_use(
spec_dir: Path,
terms: SourceTerms,
declared_use: str,
layer: str = "annotation",
) -> tuple[str, str | None]
The declaration a gate should judge, and where it came from (S105, RM252).
The caller's flag when it states one. Otherwise the module's own recorded declaration for this
source at this layer — the row an earlier run wrote into the licence table, which is the same
author, the same module, and the very file the compile gate keys on. pgx on a module drafted
with --use non-commercial was saying "no use was declared" about CPIC while the declaration sat
in the file it had just read, and asking the author to assert the same position a second time —
which is the fabrication risk the draft's own message warns about, arriving at check time.
Three things this is not. It is not a default: with nothing recorded the answer is still
unstated, and the tool asserts no purpose (@declared-use-third-axis). It is not per module: a
declaration for CPIC says nothing about PharmVar, so a leg with no row still asks. And it is not a
reading of the flag's meaning, which is unchanged — an explicit --use outranks the file, in
both directions, and commercial against a forbidding source still refuses.
Returns (declaration, origin): origin is None when the flag decided, else the licence file's
name, for the sentence that tells the author where the declaration was read from.
Source code in enricher/src/just_dna_enricher/licensing.py
write_sources_csv ¶
Write sources.csv in a fixed column order (normalized, like every reverse writer).
Source code in enricher/src/just_dna_enricher/licensing.py
merge_sources_csv ¶
merge_sources_csv(
rows: list[SourceRow],
path: Path,
existing: list[SourceRow],
) -> list[SourceRow]
Merge emitted rows into whatever is already recorded, never clobbering (as enrich() does).
Sorted by (source, layer) so the emitted order is deterministic (Principle 7).
Source code in enricher/src/just_dna_enricher/licensing.py
sidecar_path ¶
Where a machine-written sidecar lives for this module — the copy it has, else a fresh one.
The enricher's single entry point into just_dna_format.layout, so a pass never joins a filename
onto a spec directory itself. That matters twice over: the licence table has two accepted
spellings (RM51), and any of them may sit under derived/ (RM49). A pass keeping its own literal
would read the split copy and write the flat one, leaving the module with two — which is the
collision, arrived at by following the documented workflow rather than by misuse.
SidecarCollision is re-raised as the caller's own error, so a pass still fails as itself rather
than as a schema-tier ValueError nobody up the stack is catching.
Source code in enricher/src/just_dna_enricher/licensing.py
sources_path ¶
Where this module's licence table lives — the file it already has, else the current spelling.
Nine passes used to write spec_dir / "sources.csv" by hand; they all come through here now.
Source code in enricher/src/just_dna_enricher/licensing.py
require_sources_file ¶
The module's licence table as it stands, [] when it has none — refusing one that does not load.
The read half of merge_sources_file, published on its own so a pass can run it before the
fetch it is about to pay for (S98, RM231): a scaffold's placeholder row used to be found only at
the merge, after a 47-minute query and after the data table was already on disk. Same refusal,
same error type, moved to where it costs a second. The gentle counterpart for a reader is
read_sources_file below, which withholds instead of refusing.
Source code in enricher/src/just_dna_enricher/licensing.py
merge_sources_file ¶
merge_sources_file(
rows: list[SourceRow],
spec_dir: Path,
*,
error: type[Exception],
) -> list[SourceRow]
Read the module's licence table if it is there, merge rows in, and write it back.
The read-merge-write every terms-emitting pass performs, in one place: a pass that consulted a
source has to record it, and each of them was otherwise growing its own copy of these nine lines.
An unparseable existing file raises rather than being overwritten — merging into a table that did
not load would silently drop the rows already recorded. error is the caller's own exception
type, so a failure still surfaces as that pass's error rather than as a licensing one.
Takes the spec directory, not a path: resolving the filename here is what stops a caller naming a spelling the module does not use.
Source code in enricher/src/just_dna_enricher/licensing.py
withdraw_stale_dataset ¶
withdraw_stale_dataset(
spec_dir: Path,
source: str,
layer: str,
dataset: str | None,
*,
error: type[Exception],
) -> str | None
Blank a recorded dataset that this run's rows did not come from. Returns what it withdrew.
The one place anything overwrites a cell merge_sources_file would have kept, and it is narrow on
purpose: merge_sources_csv is never-clobber so a curator's hand-written terms survive a
re-run, which is right, and dataset inherited that protection at the moment RM4 made it
load-bearing. A module drafted from one release and then widened from a newer one kept the older
label — a licence row asserting a release half its rows did not come from, in the column a
published manifest.sources carries and the clinical cross-check keys on.
It only ever withdraws, never re-labels, because the honest value for a module carrying rows
from two releases is not the newer label either — one column cannot name two releases, so the
answer is unknown and unknown is withheld (the house rule). That is also the safe direction for
everything downstream: an empty dataset skips nothing, so the cross-check simply runs.
None when there was nothing to withdraw — no row, or a row already naming this run's release.
The caller decides whether its rows even changed the module's provenance; a re-draft that added
nothing must not reach this at all.
Source code in enricher/src/just_dna_enricher/licensing.py
read_sources_file ¶
The module's licence rows as recorded, or [] when there are none that can be read.
The gentle counterpart to the strict load inside merge_sources_file, for a reader whose only
power is to let a check be skipped — clinical.tautology_reason, which asks whether the licence
row says these annotation rows were drafted from the snapshot the check is about to read (RM4).
Gentle deliberately, and in the same direction the rest of this codebase withholds: a table that
could not be read has established nothing, so [] leaves every check running. A pass that
writes must still fail loudly on an unreadable table — merging into one that did not load would
drop rows already recorded — and merge_sources_file does.
Source code in enricher/src/just_dna_enricher/licensing.py
overlaid_input_rows ¶
A derived table as the module asserts it, for a pass reading it as an INPUT (RM136).
The compiler applies overrides.csv before any check reads a row, which is the whole point: a
check must report on what the module asserts. The enricher did not — its passes re-read the raw
derived file — so an author who corrected a resolution.csv cell through the overlay went on being
told the same finding by the tier that writes it, on every run, forever, with nothing saying their
correction had been recorded and honoured one tier over.
This is not a second implementation of the overlay, and the distinction is the entry's own.
RM136 refuses "teaching every enricher pass to apply the overlay" on the grounds that a second
apply_overrides would drift on the normalization seam. This calls the apply_overrides, the
one the compiler calls, through the one loader — so there is nothing to drift from.
INPUT reads only, and merge baselines must never come through here. A pass that reads its own output file to merge against it writes that file back; feeding it post-overlay rows would bake the correction into the derived table, and the enricher would be writing through the overlay — the author's answer restated as the tier's, which is RM83's standing refusal. The rule is the one the sidecar rules already state from the other side: read the file you write, and write what you read.
Errors from the overlay are raised as the caller's own exception rather than swallowed: an overlay that does not apply is a broken module, and a pass that quietly used the raw rows instead would be the silent-success shape this workspace keeps closing.
Source code in enricher/src/just_dna_enricher/licensing.py
overlay_answers ¶
The (subject, field) pairs this module's overlay has already answered for table (RM136).
Per field, and that is the decision. A finding is answered when the overlay updates the very
cell the finding is about — so correcting a coordinate silences the coordinate check and leaves an
unrelated clin_sig finding standing. Per row was the cheaper rule and was refused: an author
correcting one cell would silence findings they never looked at, which is the silent-suppress hole
the overlay's own design calls its worst case.
Only update counts. An insert supplies a row the source had no answer for, so there was no
finding to answer; a suppress removes the row, and RM131 already reports that removal in its own
right. Returns an empty set when the module has no overlay, which is every module today — a check
that consults this must therefore behave exactly as before on one.
Read-only. Nothing here writes, and a malformed overlay is the caller's problem to raise on
through overlaid_input_rows; this answers set() rather than guessing.
Source code in enricher/src/just_dna_enricher/licensing.py
overlay_answered_subjects ¶
The (subject, member) pairs this module's overlay answers for table (RM151).
Every operation and every field, and that is the rule rather than an omission. What this
feeds is the staleness question — has the value a recorded judgement was written about moved
since? — and the judgement is the reason, which the model makes mandatory on every overlay row
whatever it does. An author who suppresses a contested subject has reasoned about the same
values as one who updates its call, so a per-field rule would have to name a field the reason
does not live in.
It is deliberately not overlay_answers, whose per-field rule is the right one for the
opposite direction. That one decides whether a finding may be silenced, so it insists the
overlay touched the very cell the finding is about — anything looser would silence findings the
author never looked at, the overlay design's own worst case. This one raises a finding, and a
finding raised too widely costs a reader one line rather than hiding one.
An empty member is group-scoped and is returned as ""; the caller decides what a group means
for its table, because only it knows the members. Ordered by first appearance in the overlay,
deduplicated, so a caller's message is deterministic.
Read-only, and empty for a module whose overlay does not parse — a malformed overlay is raised on
by overlaid_input_rows, and answering [] here rather than guessing keeps one loader owning
that diagnosis.