Skip to content

just_dna_enricher.currency

just_dna_enricher.currency

Has the source a module was drafted from published since? (RM85)

SourceRow.dataset already records which release a module's rows came from — that is what the tautology skip reads, and what withdraw_stale_dataset blanks when a module ends up mixing two. What nothing did was act on it. A module drafted from ClinVar inherits ClinVar's weekly cadence and needs a source-refresh pass; one built from a paper inherits the literature's and needs an evidence pass; and neither an author who has forgotten nor a curator who inherited the module was ever told which.

So this is a comparison, not a column: the recorded label against the label the source publishes now. It needs the network, which is why it is an enricher check (the validation-ceiling rule) and not a compiler one, and it reads and never writes — the column-shaped repair (a second field saying what this module was made from and what would age it) was refused one table over in RM71, on the grounds that it restates dataset and then rots where dataset is maintained.

It is --rederive's cheap neighbour. Both ask has the world moved: --rederive re-asks every source about every subject and reports the rows that changed, which costs a full run; this asks about the release label alone and costs one request per source. So the label check is what tells an author whether the expensive one is worth running, and the two compose rather than overlap.

Tri-state, and --offline is where it bites. A source that could not be asked is unchecked — never up to date. That distinction is the whole value here: a check that reported a clean bill for a source nobody reached would be the S4 defect wearing the badge of the mechanism built to end it. So behind is True / False / None, the unaskable legs are named in the record's detail rather than counted into the denominator, and with no leg answering the pass records a skip instead of a zero.

Comparability is tri-state too. clinvar_dataset_label has two forms — clinvar_2026-08-25 from the VCF header, and clinvar_sha256:… for a snapshot built from a VCF whose header stated no date. The two name the same release space and cannot be tested for equality across forms, so a recorded digest against a live date is uncomparable, not behind. Withheld, and named.

Two probes ship. ClinVar is the source the whole item is about; the PGS Catalog joined it in RM163 because that source publishes its own release record — /rest/info states the date and the score count outright, so the label to compare against was already there and needed reading rather than building. Every source outside PROBE_SOURCES is honestly unsupported rather than quietly current, and the registry is what a reader consults. Adding a probe is adding a member; nothing else here changes.

ReleaseProbeError

Bases: RuntimeError

A release probe failed in a way its caller must be able to tell from a real answer.

ReleaseUnavailable

Bases: ReleaseProbeError

The source could not be asked at all — a failed request, never "it publishes nothing".

A subclass, not a second type (P3): an existing except ReleaseProbeError keeps firing, while a caller that needs to separate unreachable from unreadable can. @client-exception-contract, which also makes a handler's except order load-bearing — the narrow arm goes first.

ClinVarReleaseClient dataclass

ClinVarReleaseClient(
    url: str = DEFAULT_CLINVAR_URL,
    timeout: float = 30.0,
    client: Client | None = None,
    gate: PacingGate | None = None,
)

What release ClinVar publishes now, read from the live VCF's own header.

The label is built with CLINVAR_DATASET_PREFIX + the ##fileDate= value, through the same reader clinvar_build uses on a downloaded file — so a probe result and a snapshot's recorded dataset are the same string when they are the same release. A second spelling here would make the check quietly never match, which is the failure clinvar_dataset_label exists to prevent one function over.

It streams and abandons rather than asking for a byte range. A Range header is the obvious move and depends on the server honouring it; a server that ignores one answers 200 with the whole 200 MB body and the probe silently becomes a download. Reading the first HEADER_PROBE_BYTES off a normal stream and closing it needs no such promise.

@client-exception-contract: retry, then translate, both legs — a persistent 5xx and an exhausted transport failure both arrive as ReleaseUnavailable, never as an httpx type.

current_release

current_release() -> str | None

clinvar_<fileDate>, or None when what came back stated no release date.

None is the withhold and is not the same as raising: the source answered, and what it said carries no label to compare against. The caller records those two as different reasons.

Source code in enricher/src/just_dna_enricher/currency.py
def current_release(self) -> str | None:
    """`clinvar_<fileDate>`, or `None` when what came back stated no release date.

    `None` is the withhold and is *not* the same as raising: the source answered, and what it said
    carries no label to compare against. The caller records those two as different reasons.
    """
    try:
        raw = self._header_bytes()
    except (httpx.TransportError, httpx.HTTPStatusError) as exc:
        raise ReleaseUnavailable(f"ClinVar release probe failed: {exc}") from exc
    date = file_date_from_header(_gunzip_prefix(raw))
    return f"{CLINVAR_DATASET_PREFIX}{date}" if date else None

DatasetCurrency dataclass

DatasetCurrency(
    source: str,
    layer: str,
    recorded: str,
    current: str | None = None,
    unchecked: str | None = None,
)

One (source, layer) row's recorded release, against what its source publishes now.

behind property

behind: bool | None

Tri-state: True the source has published since, False still current, None unknown.

None is never False. A leg nobody could ask, and a pair of labels written in two different forms, both land here — and a caller that read this as "up to date" would publish exactly the reassurance this check exists to withhold.

CurrencyCheck dataclass

CurrencyCheck(
    compared: tuple[DatasetCurrency, ...] = (),
    unchecked: tuple[DatasetCurrency, ...] = (),
    not_checked: str | None = None,
)

What the pass compared, what it could not, and why — the denominator travelling with the finding.

compared is the honest subject set: the legs that were asked and answered comparably. A leg nobody could reach is in unchecked and counted nowhere, because counting it would claim a comparison that was never made — the coverage lie the reference-allele pass shipped once already.

subjects property

subjects: int

How many recorded releases were actually compared against a stated current one.

behind property

behind: list[DatasetCurrency]

The legs whose source has published since — the findings, in recorded order.

default_probes

default_probes(
    *,
    clinvar: ClinVarReleaseClient | None = None,
    pgs_catalog: PgsCatalogClient | None = None,
) -> dict[str, ReleaseProbe]

The probes this tier ships, keyed by the SourceRow.source they answer for.

A registry rather than a chain of if source == …: what a reader needs is the set of sources that can be asked, and PROBE_SOURCES below is derived from this function so the two cannot disagree.

Source code in enricher/src/just_dna_enricher/currency.py
def default_probes(
    *,
    clinvar: ClinVarReleaseClient | None = None,
    pgs_catalog: PgsCatalogClient | None = None,
) -> dict[str, ReleaseProbe]:
    """The probes this tier ships, keyed by the `SourceRow.source` they answer for.

    A registry rather than a chain of `if source == …`: what a reader needs is the set of sources that
    *can* be asked, and `PROBE_SOURCES` below is derived from this function so the two cannot disagree.
    """
    clinvar_client = clinvar if clinvar is not None else ClinVarReleaseClient()
    pgs_client = pgs_catalog if pgs_catalog is not None else PgsCatalogClient()
    return {
        "clinvar": clinvar_client.current_release,
        PGS_SOURCE: lambda: _pgs_release(pgs_client),
    }

check_dataset_currency

check_dataset_currency(
    rows: Sequence[SourceRow],
    *,
    probes: Mapping[str, ReleaseProbe] | None = None,
    offline: bool = False,
) -> CurrencyCheck

Compare every recorded dataset against the release its source publishes now.

Reads and writes nothing: it takes the rows a caller already loaded and returns what it found, so a run that reports a gap leaves the spec directory byte-for-byte as it was. Repairing a stale label is a re-draft, which is an author's decision and a different command.

One request per source, not per row: two layers of one source share a probe result, because "what does ClinVar publish now" has one answer whatever a module used it for.

probes is injected, the way resolver and gnomad_client are one module over, and it is what the tests drive: the shipped registry is built only once a leg could actually be asked, so an offline run opens no client at all.

Source code in enricher/src/just_dna_enricher/currency.py
def check_dataset_currency(
    rows: Sequence[SourceRow],
    *,
    probes: Mapping[str, ReleaseProbe] | None = None,
    offline: bool = False,
) -> CurrencyCheck:
    """Compare every recorded `dataset` against the release its source publishes now.

    Reads and writes nothing: it takes the rows a caller already loaded and returns what it found, so
    a run that reports a gap leaves the spec directory byte-for-byte as it was. Repairing a stale
    label is a re-draft, which is an author's decision and a different command.

    One request per *source*, not per row: two layers of one source share a probe result, because
    "what does ClinVar publish now" has one answer whatever a module used it for.

    `probes` is injected, the way `resolver` and `gnomad_client` are one module over, and it is what
    the tests drive: the shipped registry is built only once a leg could actually be asked, so an
    offline run opens no client at all.
    """
    subjects = [row for row in rows if (row.dataset or "").strip()]
    if not subjects:
        # Not a skip that a flag or egress would clear: the module records no release, so there is no
        # claim to have an opinion about. `nothing_to_check` is exactly that member.
        return CurrencyCheck(not_checked="nothing_to_check")

    if offline:
        # Returned before the registry is built, so an offline run opens no client at all — the
        # off-switch has to be provable by the probe never being invoked, not by the reason it wrote.
        # Every recorded release is `unchecked`: an offline run has not established that any source
        # stands still, and saying otherwise is the one thing this check may never do.
        return CurrencyCheck(
            (),
            tuple(
                DatasetCurrency(row.source, row.layer, (row.dataset or "").strip(), unchecked="offline")
                for row in subjects
            ),
            not_checked="offline",
        )

    # The **effective** registry decides who can be asked, never `PROBE_SOURCES`: an injected registry
    # is entitled to answer for a source this tier ships no probe for, and testing the shipped set
    # here would make the injection point unable to widen the check.
    registry = probes if probes is not None else default_probes()
    #: `source -> (label, reason)` — one probe result per source, reused across its layers. Cached
    #: rather than re-asked so a module recording ClinVar at two layers costs one request, and so the
    #: two layers can never be told different things about one source.
    answered: dict[str, tuple[str | None, str | None]] = {}
    compared: list[DatasetCurrency] = []
    unchecked: list[DatasetCurrency] = []

    for row in subjects:
        recorded = (row.dataset or "").strip()
        if row.source not in answered:
            answered[row.source] = _ask(registry, row.source)
        label, reason = answered[row.source]
        if reason is not None:
            unchecked.append(DatasetCurrency(row.source, row.layer, recorded, unchecked=reason))
        elif label is None or _label_kind(label) != _label_kind(recorded):
            # Read and unreadable are different absences. A source that stated no release, and a
            # source whose stated release is written in the other of the two label forms, both leave
            # nothing to compare — which is `no_reference`, and emphatically not "still current".
            unchecked.append(
                DatasetCurrency(row.source, row.layer, recorded, current=label, unchecked="no_reference")
            )
        else:
            compared.append(DatasetCurrency(row.source, row.layer, recorded, current=label))

    if compared:
        return CurrencyCheck(tuple(compared), tuple(unchecked))
    reasons = {leg.unchecked for leg in unchecked}
    reason = next((member for member in _SKIP_PRECEDENCE if member in reasons), "unreachable")
    return CurrencyCheck((), tuple(unchecked), not_checked=reason)

unchecked_sentences

unchecked_sentences(check: CurrencyCheck) -> list[str]

One sentence per reason for the legs that could not be settled — never one per row.

Shared between the record's detail and the CLI's report, so what an author reads on the terminal and what the attestation carries cannot drift into two accounts of one run.

Source code in enricher/src/just_dna_enricher/currency.py
def unchecked_sentences(check: CurrencyCheck) -> list[str]:
    """One sentence per *reason* for the legs that could not be settled — never one per row.

    Shared between the record's `detail` and the CLI's report, so what an author reads on the terminal
    and what the attestation carries cannot drift into two accounts of one run.
    """
    by_reason: dict[str, list[DatasetCurrency]] = {}
    for leg in check.unchecked:
        by_reason.setdefault(leg.unchecked or "unreachable", []).append(leg)
    return [
        f"{len(by_reason[reason])} recorded release(s) unchecked ({reason}): "
        + examples([f"{c.source} {c.recorded}" for c in by_reason[reason]])
        for reason in sorted(by_reason)
    ]

summarize_currency

summarize_currency(check: CurrencyCheck) -> list[str]

What a record's detail carries: the superseded releases, then the legs nobody could settle.

The shortfall travels with the finding rather than beside it — a coverage figure whose denominator is stated elsewhere is the defect _vrs_coverage exists for, one check over.

Source code in enricher/src/just_dna_enricher/currency.py
def summarize_currency(check: CurrencyCheck) -> list[str]:
    """What a record's `detail` carries: the superseded releases, then the legs nobody could settle.

    The shortfall travels with the finding rather than beside it — a coverage figure whose
    denominator is stated elsewhere is the defect `_vrs_coverage` exists for, one check over.
    """
    behind = check.behind
    superseded = (
        [
            f"{len(behind)} recorded release(s) have been superseded: "
            + examples([f"{c.source} {c.recorded} → {c.current}" for c in behind])
        ]
        if behind
        else []
    )
    return superseded + unchecked_sentences(check)