just_dna_enricher.pgs¶
just_dna_enricher.pgs ¶
The PGS Catalog — the registry behind pgs.csv's accession, and the terms behind each score (RM163).
PgsRow is keyed on pgs_id, and until now that was the one authored identifier in the format no
pass ever put to its registry. This module is the client and the source-shape knowledge behind the
fourth registry in identifiers.py; the statuses, the comparison and the two attestations live there,
beside dbSNP's, OLS4's and HGNC's.
Read the body, never the status. Measured against the live REST surface on 2026-09-01:
GET /rest/score/PGS000001 returns the record, GET /rest/score/PGS999999 — a never-assigned id —
returns HTTP 200 with {}, and GET /rest/score/PGSXXXX, which is not even a well-formed
accession, returns HTTP 200 with {} as well. So the status code carries no existence information
at all and every negative looks identical from the outside. @existence-not-identity: a lookup that
answers "does this exist" has to say what it found, which is why a recognised accession carries the
score's name, release date and traits rather than a bare True.
A 404 is therefore a failure here, not an answer, and that is the opposite of OntologyClient's
rule one module over. OLS4 and HGNC say "no such term" with a 404; this service says it with an empty
body on a 200, so a 404 from it means the request went somewhere unexpected — a moved path, a proxy,
a maintenance page — and reading that as "the Catalog does not hold this score" would turn an
infrastructure problem into a permanent negative about an author's row.
The id space is sparse, and that decides how the absence is worded. 6,982 scores span a range
reaching PGS019960 — roughly a third of it assigned — and a random sample of 40 well-formed in-range
ids returned 12 records and 28 empty bodies. So an unrecognised accession is overwhelmingly one that
was never issued rather than one that was withdrawn. @rsid-absent-two-readings earns its
equal-weight treatment because dbSNP's id space is densely assigned and merges are frequent; here the
base rate runs the other way, so the message names the typo reading first and withdrawal as the rarer
one. The Catalog publishes no supersession field at all, and that is stated as a limit of the source
rather than resolved by guessing.
The release record is read, not built. /rest/info publishes latest_release — date, score
count, publication and trait counts — beside the REST API's own version, so the currency question this
item wanted needed no builder and no snapshot. dataset_label spells that release the way
SourceRow.dataset records it, and currency.default_probes asks for it with the same function, so a
probe result and a recorded label are the same string when they are the same release.
PgsCatalogError ¶
Bases: RuntimeError
The PGS Catalog answered something this tier cannot use.
PgsCatalogUnavailable ¶
Bases: PgsCatalogError
The Catalog could not be asked at all — a failed request, never "it holds no such score".
A subclass rather than a second type, for ReleaseUnavailable's reason (P3): an existing
except PgsCatalogError keeps firing, and a caller that needs to separate unreachable from
unreadable can.
identifiers catches this rather than re-raising it, and that is worth knowing before writing
a handler for it. check-identifiers puts questions to four registries and writes one record per
check, so letting this out as an IdentifierUnavailable would make the command stamp unreachable
against OLS4's and HGNC's records too — for services that answered. The failure lands on
IdentifierReport.pgs_not_checked and only the two PGS records read it. currency._pgs_release is
the one caller that does translate, into ReleaseUnavailable, because a ReleaseProbe's contract
is an exception and its reason is read off the subclass.
CatalogRelease
dataclass
¶
CatalogRelease(
date: str | None = None,
scores: int | None = None,
rest_api_version: str | None = None,
terms_of_use: str | None = None,
)
What /rest/info says the Catalog published most recently.
PgsCatalogClient ¶
PgsCatalogClient(
*,
base: str = PGS_REST_BASE,
client: Client | None = None,
gate: PacingGate | None = None,
timeout: float = 30.0,
)
One score record, or the Catalog's own release, paced and with a per-run cache.
A class rather than a function because the pacing gate and the cache are per-run state, the same
reason ClingenAlleleClient is one. @client-exception-contract: retry, then translate, both
legs — a persistent 5xx and an exhausted transport failure both arrive as
PgsCatalogUnavailable, never as an httpx type.
Source code in enricher/src/just_dna_enricher/pgs.py
score ¶
One score record, or {} when the Catalog holds none under that accession.
The empty dict is the Catalog's own answer and is cached like any other: a module naming the same accession on two rows costs one request, and the two rows can never be told different things.
Source code in enricher/src/just_dna_enricher/pgs.py
info ¶
release ¶
The release /rest/info names. Raises PgsCatalogUnavailable if it cannot be asked.
parse_release ¶
/rest/info → the release it names. Pure, so it is testable against a recorded payload.
Source code in enricher/src/just_dna_enricher/pgs.py
dataset_label ¶
pgs_catalog_2026-08-26, or None when the Catalog named no release.
One function for the SourceRow.dataset stamp and for currency's probe, which is
clinvar_dataset_label's rule: a second spelling would make the currency check quietly never
match, and "never matches" renders as unreadable rather than as a bug.
Source code in enricher/src/just_dna_enricher/pgs.py
score_ancestries ¶
(the format codes this score's dev/eval ancestries name, the categories that block a negative).
Both halves are returned because the second is what makes the comparison honest. A category blocks
a negative for either of two reasons, and they are different facts kept in one answer because the
caller does the same thing with both: NR, ASN, GME and OTH name an ancestry this format has
no member for, and MAE/MAO name a bag of superpopulations the Catalog did not break down. In
both cases an authored code the first half does not contain may be exactly the one behind the
category, so the caller withholds rather than reporting a difference it cannot stand behind.
A MAE/MAO category still contributes multi to the first half: it is a perfectly good positive
answer, and only the negative is unsafe.
Every stratum with a positive share counts, however small. A threshold would be a second judgement about what "the score was developed and evaluated in" means, on top of the stage choice above.
Source code in enricher/src/just_dna_enricher/pgs.py
ancestry_agrees ¶
Whether one authored training_ancestry member is covered by what the Catalog publishes.
multi is the only member that is not a population, so it is the only one with a rule of its own:
it is covered by either of the Catalog's two multi-ancestry categories, and also by a published
set naming two or more superpopulations, which is what multi-ancestry means. Anything else is a
plain membership test.
Source code in enricher/src/just_dna_enricher/pgs.py
score_cohort_names ¶
Everything samples_training says about who the score was trained on, first-occurrence order.
Cohort short names, full names and alternate names, plus the per-sample ancestry prose the
Catalog carries beside them (ancestry_broad, ancestry_free, ancestry_country,
ancestry_additional) — because training_cohort's own examples are FIN and Ashkenazi, which
are ancestries rather than cohorts, and comparing them against cohort names alone would accuse
every correct row that used one.
Empty when the record lists no training samples, which the caller reads as nothing to compare rather than as a disagreement.
Source code in enricher/src/just_dna_enricher/pgs.py
cohort_agrees ¶
Whether an authored training_cohort shares any real word with what the record says.
Three-valued: True agrees, False shares nothing, None means the question could not be
put — the authored cell carries no word this can test with. Short tokens are dropped because two
letters hit by accident inside unrelated prose, and the structural words in
_GENERIC_COHORT_WORDS are dropped because they hit on purpose: a match by accident is a clean
bill nobody earned.
Deliberately a weak test, and the weakness is the point. training_cohort is free-form by
design — 'FIN', 'Ashkenazi', 'UK Biobank NW-EUR' — and no reliable equality exists between
prose a curator wrote and prose the Catalog wrote. A comparison that cannot be made reliable is
one that reports unknown, so False is reserved for the case where no reading is available at
all: not one word of the authored cell appears anywhere in the record's own account of its
training samples. Substring rather than whole-word matching, so FIN is vouched for by
Finnish — insisting on token equality would accuse a correct row.