just_dna_format.pgs¶
just_dna_format.pgs ¶
Polygenic score declaration (0.4 — see docs/CHANGELOG.md).
pgs.csv is a manifest of PGS Catalog IDs, not authored weights — just-prs resolves a PGSxxxxxx
id to a harmonized scoring file itself and scores each id independently, so per-PGS weights would be
dead data. It is therefore a declared interface (like GenePanelSpec), not a measure→phenotype
binning table: a PRS yields a Z/percentile within a matched reference distribution, which the
format does not bin.
The one-way-door fields (consumer round-2 Q8) are pinned here from day one so a
consumer can refuse or caveat an out-of-ancestry application instead of silently miscalibrating:
- training_ancestry — the superpopulation(s) the score was developed and evaluated in (the PGS
Catalog's dev and eval samples; required floor), plus an optional free-form training_cohort
for the sub-superpop precision superpop codes can't express (a Northwest-EUR-trained score applied
to a Finnish/Ashkenazi sample).
- match_rate_floor — the author-set floor (a > ~20% variant mismatch invalidates the score). Only
the floor lives here: the observed per-sample match rate is a measurement, so by the
data-agnostic north star (CLAUDE.md) it is consumer/runtime-side and must NOT live in the module.
- research_tier — pins as data that a PRS is a within-reference Z/percentile, never an
ancestry-calibrated absolute risk; |Z| >= 2.5 in a healthy proband is a population-stratification
signal, not a disease prediction.
Data-agnostic (design north star — see CLAUDE.md): this declares which scores a module curates and how to caveat them; no sample, genotype, or computed score lives here.
PgsRow ¶
Bases: AuthoredModel
One curated PGS Catalog entry. Inherits AuthoredModel (reserved-namespace guard, which keeps
the namespace closed, + the shared trait_efo_id validator).