Skip to content

pgs.csv

One row is one curated PGS Catalog entry — the accession and the trait, not the weights. The scoring file itself is deliberately not authored here: a polygenic score's variant weights are a data file, and carrying them as an authored table is tracked separately rather than assumed.

This table reaches no gate, and research_tier is not a licence axis. Its two members are a calibration judgement — research_only pins as data that a score yields a within-reference Z/percentile, never an ancestry-calibrated absolute risk, and calibrated says the opposite — so research_only is about what the number means, not about who may use it. Nothing in pgs.csv changes what compiles. The paragraph here said the reverse until 2026-09-13, and an author who had read it believed the accession carried its restriction to the gate by itself, so an empty licence ledger read as nothing restrictive here rather than as nobody declared anything (S101).

What the table does owe is a licence row. The Catalog publishes license per score record and it varies, so a score whose record restricts use needs its own row in licensing.csv — and that row is what the gate reads, keyed on that file and nothing else.

Identity

Row model PgsRow (just_dna_format.pgs)
Becomes pgs.parquet — lead parquet: pgs.parquet
Authored or derived authored — a person writes it (a drafter may append rows)
Draftable yes — draft can append rows
Drafted by no drafting provider targets this table — draft writes no row of it
Natural key pgs_id, trait_efo_id
Fact signature no
In the attestation binding yes — manifest.inputs[]

Columns

Column Type Required Values Meaning
pgs_id str required PGS Catalog id, e.g. PGS000135
trait_efo_id str | None optional EFO/MONDO/OBA/HP trait ontology id(s) — joins with variant modules
note str | None optional Free-text note
group str | None optional Grouping label within the module
training_ancestry list[str] | None optional one of: AFR, AMR, EAS, EUR, SAS, multi Superpopulation(s) the score was developed and evaluated in: the PGS Catalog's development and evaluation samples, not the discovery GWAS behind its weights (1000G superpop codes; multi-valued)
training_cohort str | None optional Optional free-form sub-superpop cohort, e.g. 'FIN', 'Ashkenazi', 'UK Biobank NW-EUR'
match_rate_floor float | None optional Author-set variant-match floor in [0,1]; a score computed below it is invalid. Only the floor lives in-module — the observed per-sample match rate is a measurement (consumer-side).
research_tier str | None optional one of: calibrated, research_only Calibration frame, not a licence term (VALID_RESEARCH_TIERS): research_only is a within-reference Z/percentile, calibrated is ancestry-calibrated absolute risk. Reaches no compile gate — a score whose Catalog record restricts use needs a licensing.csv row, which is the only thing the gate reads.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.