association_id |
str |
required |
|
The Catalog's own association id. Identity, the same source-accession pattern as PgsRow.pgs_id and PharmVariantRow.annotation_id — one variant carries dozens of associations and only the archive's id tells them apart. |
variant_key |
str |
required |
|
Coordinate-derived identity of the locus (matches the post-expansion weights key) |
rsid |
str | None |
optional |
|
dbSNP identifier, as the Catalog itself names it. Position-level, so it names the locus rather than the allele — which is effect_allele's job. |
effect_allele |
str | None |
optional |
|
The allele effect_size is stated relative to, parsed from the Catalog's riskAlleleName (rs4149056-C → C). Null when the source wrote -?, which it does often — 2 of the first 3 rs4149056 associations — meaning the study did not establish which allele carries the effect. An effect size relative to an unknown allele cannot be used as a weight, so the row is kept and counted rather than dropped: that is a fact a consumer needs, not a defect to hide. |
effect_size |
float | None |
optional |
|
The reported magnitude. Its meaning is effect_measure; its scale is effect_unit. |
effect_measure |
str | None |
optional |
suggested: HR, NR, OR, RR, beta, log(HR), log(OR) |
What kind of magnitude it is — OR or beta in practice. The Catalog keeps these in two mutually exclusive fields (orPerCopyNum, betaNum) rather than the single ambiguous 'OR or BETA' column its download TSV uses, so this maps cleanly. Open, matching StudyRow.effect_measure. |
effect_unit |
str | None |
optional |
|
The unit a beta is in, verbatim from the source — 'umol/l', 'unit', 'cm'. Kept verbatim including the uninformative ones: 'unit' really is all many studies report, and recording that honestly is the difference between a scale a consumer can reason about and one it cannot. Null for an OR, which is dimensionless. |
effect_direction |
str | None |
optional |
one of: decrease, increase |
Whether the effect allele increases or decreases the measured trait. Not clinical direction — VariantRow.direction is protective|risk|neutral|unknown and answers a different question, since increasing a trait may be good, bad or neither. Folding the two together would overload one field with two axes (Principle 5); the source keeps them apart (betaDirection) and so does this. |
standard_error |
float | None |
optional |
|
Standard error of the effect size, when the study reported one. |
confidence_interval |
str | None |
optional |
|
The reported 95% CI, verbatim — '[0.14-0.17]', and also '[NR]' where the study reported none. A string because that is what the source publishes and the bracket forms vary; parsing it into two floats here would discard the cases that do not fit. |
risk_allele_frequency |
float | None |
optional |
|
Frequency of the effect allele in the study's sample. The source writes 'NR' for not reported, which is an ABSENCE and arrives here as null — distinct from a reported 0. |
p_value |
str | None |
optional |
|
Raw p-value string, verbatim, mirroring StudyRow.p_value. |
p_value_num |
float | None |
optional |
|
The queryable p-value, mirroring StudyRow.p_value_num. The Catalog's own mantissa/exponent pair is deliberately not carried, and the recorded 0.5 rejection has two halves of which only one still applies: 'a cost every author pays' does not transfer to a machine-written table, but 'a catalogue-of-millions problem rather than a curated-module one' does, and a module cites tens of associations, not millions. |
trait |
str | None |
optional |
|
The reported trait in the Catalog's words. Descriptive — it is re-worded between releases for an unchanged ontology term, so trait_efo_id is the fact and this is what makes the row readable. |
trait_efo_id |
str | None |
optional |
|
EFO trait id(s) — the join back to VariantRow.trait_efo_id. |
pmid |
str | None |
optional |
|
PubMed id of the publication, free-form under the same grammar as StudyRow.pmid. |
study_accession |
str | None |
optional |
|
The Catalog's study accession, e.g. 'GCST001234'. |
ancestry |
str | None |
optional |
|
The study population, free text as the Catalog records it ('European', 'East Asian', 'Hispanic or Latin American'). Free rather than a vocabulary because it is prose there and closing it would make a future release unloadable; and NOT merged with PgsRow.training_ancestry, which is a 1000G superpopulation code list — vocab.py forbids collapsing two ancestry vocabularies into one. |
dataset |
str |
required |
|
Which Catalog release this row came from, e.g. 'gwas_catalog_2026-08-01'. A FACT: the Catalog re-curates, and an association whose effect size moved is a different fact. |
source |
str | None |
optional |
|
The licensed data source: gwas_catalog|manual|reversed (open). Joins sources.csv.source. Names the SOURCE, never the route. |
status |
str | None |
optional |
one of: ambiguous, not_found, resolved |
Outcome: resolved|not_found|ambiguous. not_found is a FACT — the Catalog was consulted and reports no association for this variant — and differs from a variant never queried, which has no row at all. |
fetched_at |
str | None |
optional |
|
ISO-8601 UTC timestamp, second resolution. Canonicalized on load; records when this row was written by a pass, not when the Catalog published anything — that is dataset. |