pgs.csv¶
One row is one curated PGS Catalog entry — the accession and the trait, not the weights. The scoring file itself is deliberately not authored here: a polygenic score's variant weights are a data file, and carrying them as an authored table is tracked separately rather than assumed.
This table reaches no gate, and research_tier is not a licence axis. Its two members are a
calibration judgement — research_only pins as data that a score yields a within-reference
Z/percentile, never an ancestry-calibrated absolute risk, and calibrated says the opposite — so
research_only is about what the number means, not about who may use it. Nothing in pgs.csv
changes what compiles. The paragraph here said the reverse until 2026-09-13, and an author who had
read it believed the accession carried its restriction to the gate by itself, so an empty licence
ledger read as nothing restrictive here rather than as nobody declared anything (S101).
What the table does owe is a licence row. The Catalog publishes license per score record and it
varies, so a score whose record restricts use needs its own row in
licensing.csv — and that row is what the gate reads, keyed on that file and
nothing else.
Identity¶
| Row model | PgsRow (just_dna_format.pgs) |
| Becomes | pgs.parquet — lead parquet: pgs.parquet |
| Authored or derived | authored — a person writes it (a drafter may append rows) |
| Draftable | yes — draft can append rows |
| Drafted by | no drafting provider targets this table — draft writes no row of it |
| Natural key | pgs_id, trait_efo_id |
| Fact signature | no |
| In the attestation binding | yes — manifest.inputs[] |
Columns¶
| Column | Type | Required | Values | Meaning |
|---|---|---|---|---|
pgs_id |
str |
required | PGS Catalog id, e.g. PGS000135 | |
trait_efo_id |
str | None |
optional | EFO/MONDO/OBA/HP trait ontology id(s) — joins with variant modules | |
note |
str | None |
optional | Free-text note | |
group |
str | None |
optional | Grouping label within the module | |
training_ancestry |
list[str] | None |
optional | one of: AFR, AMR, EAS, EUR, SAS, multi |
Superpopulation(s) the score was developed and evaluated in: the PGS Catalog's development and evaluation samples, not the discovery GWAS behind its weights (1000G superpop codes; multi-valued) |
training_cohort |
str | None |
optional | Optional free-form sub-superpop cohort, e.g. 'FIN', 'Ashkenazi', 'UK Biobank NW-EUR' | |
match_rate_floor |
float | None |
optional | Author-set variant-match floor in [0,1]; a score computed below it is invalid. Only the floor lives in-module — the observed per-sample match rate is a measurement (consumer-side). | |
research_tier |
str | None |
optional | one of: calibrated, research_only |
Calibration frame, not a licence term (VALID_RESEARCH_TIERS): research_only is a within-reference Z/percentile, calibrated is ancestry-calibrated absolute risk. Reaches no compile gate — a score whose Catalog record restricts use needs a licensing.csv row, which is the only thing the gate reads. |
Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.