Skip to content

studies.csv

One row is one (rsid, pmid) evidence link. Not a description of a study — a link between a variant and a paper, so one study supporting five variants is five rows, and a variant with three citations is three.

Grounding evidence is mandatory, and that is the point of the table. A binning table that states thresholds with no study rows behind them warns in both modes, strict and best-effort: a bin boundary is a claim about the world and an unsourced one is the kind this format refuses to publish quietly.

A row cites when its claim is finer-grained than this table's key. studies.csv is keyed on (rsid, pmid), so when a claim is per-bin or per-genotype rather than per-variant, the citation lives on the claiming row and this table describes the study instead. The two are not alternatives; asking which one a citation belongs in is asking how specific the claim is.

The symptom of confusing it with literature.csv is a dropped row. literature.csv is machine-produced and keeps every article it found; the compiler discards an uncited literature row when nothing references it. studies.csv is the authored side and is never dropped. If a citation disappeared from a compile, it was in the wrong one of the two.

Identity

Row model StudyRow (just_dna_format.spec)
Becomes studies.parquet
Authored or derived authored — a person writes it (a drafter may append rows)
Draftable yes — draft can append rows
Drafted by no drafting provider targets this table — draft writes no row of it
Natural key variant_key, pmid
Fact signature no
In the attestation binding yes — manifest.inputs[]

Columns

Column Type Required Values Meaning
rsid str | None optional dbSNP identifier or variant key
chrom str | None optional Chromosome (for position-only variants)
start int | None optional 1-based position, VCF POS convention (position-only variants) — do not subtract one
ref str | None optional Reference allele (position-only variants)
pmid str required PubMed ID or reference — free-form, must be non-empty
population str | None optional Study population
p_value str | None optional Raw p-value string (free-form)
conclusion str | None optional Study-specific conclusion
study_design str | None optional e.g. meta-analysis, GWAS
stat_significance str | None optional one of: not_significant, significant, suggestive, unknown Per-study statistical significance: significant|suggestive|not_significant|unknown.
effect_size float | None optional Per-study effect magnitude (unit given by effect_measure).
effect_measure str | None optional suggested: HR, NR, OR, RR, beta, log(HR), log(OR) Unit of effect_size, e.g. OR|HR|beta|RR (recommended, open).
effect_allele str | None optional The allele this study's effect_size is stated relative to — bases, or a symbolic/structural allele carrying its length (e.g. ). Absent means the study did not state one, which is not the same as the reference allele.
trait_efo_id str | None optional EFO/MONDO/OBA/HP trait ontology id(s) for this study.
statistical_test str | None optional Which analysis produced this row's p_value/effect_size — the test or model, and what it was adjusted for, e.g. Fisher's exact (allelic) or logistic regression adjusted for age and sex. Free text. study_design describes the study; this describes the analysis, and one study may report several. Absent means the paper's analysis was not recorded, never that it had only one.
confidence str | None optional How far the citing source stands behind this evidence link, in ITS OWN units and unconverted (CIViC's submitted/accepted, a review-star count). Meaningless without confidence_unit beside it, and refused without one.
confidence_unit str | None optional Which instrument confidence is measured on, e.g. civic_evidence_status. Required whenever confidence is set: a magnitude with no unit beside it is a value nothing can read, and this format has paid for that once already on weight.
doi str | None optional Digital Object Identifier — wider than pmid (covers preprints/books/datasets with no PubMed id). Free-form, kept verbatim; a validator may cross-fill doi↔pmid.
provenance_quote str | None optional Optional keyword phrase / literal passage locating this study's claim in the cited article's fulltext. Human-legible; a validator confirms fulltext-contains, yes/no.
provenance_regex str | None optional Optional regex locating the claim in fulltext — a declarative pattern grammar (Principle 1: data, not code), matched consumer-side by a linear-time/ReDoS-safe engine.
curator str | None optional Who located this row's provenance quote/regex — a name, handle, or model id, resolvable against the manifest's authorship. Row-level because real work is mixed at row granularity: a human may read a review while an agent traverses its citations, in one module, in one pass. Records the distribution of labour so a reviewer can route scrutiny; it does not move responsibility, which the human author holds regardless.
p_value_num float | None optional The p-value as a number, so it can be sorted and thresholded — p_value above is a free-form string and cannot be. In (0, 1]: a p-value of exactly 0 is a source's own underflow, not a probability, so it is rejected rather than stored as a confident zero.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.