Skip to content

gwas_effects.csv

Identity

Row model GwasEffectRow (just_dna_format.gwas)
Becomes gwas_effects.parquet
Authored or derived derived — an enricher pass writes it; an author corrects it via overrides.csv
Draftable no
Written and checked by just-dna-enricher gwas (gwas) from gwas_catalog
Natural key association_id
Fact signature yes — its sidecar carries one
In the attestation binding no

Columns

Column Type Required Values Meaning
association_id str required The Catalog's own association id. Identity, the same source-accession pattern as PgsRow.pgs_id and PharmVariantRow.annotation_id — one variant carries dozens of associations and only the archive's id tells them apart.
variant_key str required Coordinate-derived identity of the locus (matches the post-expansion weights key)
rsid str | None optional dbSNP identifier, as the Catalog itself names it. Position-level, so it names the locus rather than the allele — which is effect_allele's job.
effect_allele str | None optional The allele effect_size is stated relative to, parsed from the Catalog's riskAlleleName (rs4149056-C → C). Null when the source wrote -?, which it does often — 2 of the first 3 rs4149056 associations — meaning the study did not establish which allele carries the effect. An effect size relative to an unknown allele cannot be used as a weight, so the row is kept and counted rather than dropped: that is a fact a consumer needs, not a defect to hide.
effect_size float | None optional The reported magnitude. Its meaning is effect_measure; its scale is effect_unit.
effect_measure str | None optional suggested: HR, NR, OR, RR, beta, log(HR), log(OR) What kind of magnitude it is — OR or beta in practice. The Catalog keeps these in two mutually exclusive fields (orPerCopyNum, betaNum) rather than the single ambiguous 'OR or BETA' column its download TSV uses, so this maps cleanly. Open, matching StudyRow.effect_measure.
effect_unit str | None optional The unit a beta is in, verbatim from the source — 'umol/l', 'unit', 'cm'. Kept verbatim including the uninformative ones: 'unit' really is all many studies report, and recording that honestly is the difference between a scale a consumer can reason about and one it cannot. Null for an OR, which is dimensionless.
effect_direction str | None optional one of: decrease, increase Whether the effect allele increases or decreases the measured trait. Not clinical direction — VariantRow.direction is protective|risk|neutral|unknown and answers a different question, since increasing a trait may be good, bad or neither. Folding the two together would overload one field with two axes (Principle 5); the source keeps them apart (betaDirection) and so does this.
standard_error float | None optional Standard error of the effect size, when the study reported one.
confidence_interval str | None optional The reported 95% CI, verbatim — '[0.14-0.17]', and also '[NR]' where the study reported none. A string because that is what the source publishes and the bracket forms vary; parsing it into two floats here would discard the cases that do not fit.
risk_allele_frequency float | None optional Frequency of the effect allele in the study's sample. The source writes 'NR' for not reported, which is an ABSENCE and arrives here as null — distinct from a reported 0.
p_value str | None optional Raw p-value string, verbatim, mirroring StudyRow.p_value.
p_value_num float | None optional The queryable p-value, mirroring StudyRow.p_value_num. The Catalog's own mantissa/exponent pair is deliberately not carried, and the recorded 0.5 rejection has two halves of which only one still applies: 'a cost every author pays' does not transfer to a machine-written table, but 'a catalogue-of-millions problem rather than a curated-module one' does, and a module cites tens of associations, not millions.
trait str | None optional The reported trait in the Catalog's words. Descriptive — it is re-worded between releases for an unchanged ontology term, so trait_efo_id is the fact and this is what makes the row readable.
trait_efo_id str | None optional EFO trait id(s) — the join back to VariantRow.trait_efo_id.
pmid str | None optional PubMed id of the publication, free-form under the same grammar as StudyRow.pmid.
study_accession str | None optional The Catalog's study accession, e.g. 'GCST001234'.
ancestry str | None optional The study population, free text as the Catalog records it ('European', 'East Asian', 'Hispanic or Latin American'). Free rather than a vocabulary because it is prose there and closing it would make a future release unloadable; and NOT merged with PgsRow.training_ancestry, which is a 1000G superpopulation code list — vocab.py forbids collapsing two ancestry vocabularies into one.
dataset str required Which Catalog release this row came from, e.g. 'gwas_catalog_2026-08-01'. A FACT: the Catalog re-curates, and an association whose effect size moved is a different fact.
source str | None optional The licensed data source: gwas_catalog|manual|reversed (open). Joins sources.csv.source. Names the SOURCE, never the route.
status str | None optional one of: ambiguous, not_found, resolved Outcome: resolved|not_found|ambiguous. not_found is a FACT — the Catalog was consulted and reports no association for this variant — and differs from a variant never queried, which has no row at all.
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution. Canonicalized on load; records when this row was written by a pass, not when the Catalog published anything — that is dataset.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.