Skip to content

just_dna_format.pgs

just_dna_format.pgs

Polygenic score declaration (0.4 — see docs/CHANGELOG.md).

pgs.csv is a manifest of PGS Catalog IDs, not authored weights — just-prs resolves a PGSxxxxxx id to a harmonized scoring file itself and scores each id independently, so per-PGS weights would be dead data. It is therefore a declared interface (like GenePanelSpec), not a measure→phenotype binning table: a PRS yields a Z/percentile within a matched reference distribution, which the format does not bin.

The one-way-door fields (consumer round-2 Q8) are pinned here from day one so a consumer can refuse or caveat an out-of-ancestry application instead of silently miscalibrating: - training_ancestry — the superpopulation(s) the score was developed and evaluated in (the PGS Catalog's dev and eval samples; required floor), plus an optional free-form training_cohort for the sub-superpop precision superpop codes can't express (a Northwest-EUR-trained score applied to a Finnish/Ashkenazi sample). - match_rate_floor — the author-set floor (a > ~20% variant mismatch invalidates the score). Only the floor lives here: the observed per-sample match rate is a measurement, so by the data-agnostic north star (CLAUDE.md) it is consumer/runtime-side and must NOT live in the module. - research_tier — pins as data that a PRS is a within-reference Z/percentile, never an ancestry-calibrated absolute risk; |Z| >= 2.5 in a healthy proband is a population-stratification signal, not a disease prediction.

Data-agnostic (design north star — see CLAUDE.md): this declares which scores a module curates and how to caveat them; no sample, genotype, or computed score lives here.

PgsRow

Bases: AuthoredModel

One curated PGS Catalog entry. Inherits AuthoredModel (reserved-namespace guard, which keeps the namespace closed, + the shared trait_efo_id validator).