studies.csv¶
One row is one (rsid, pmid) evidence link. Not a description of a study — a link between a variant and a paper, so one study supporting five variants is five rows, and a variant with three citations is three.
Grounding evidence is mandatory, and that is the point of the table. A binning table that states thresholds with no study rows behind them warns in both modes, strict and best-effort: a bin boundary is a claim about the world and an unsourced one is the kind this format refuses to publish quietly.
A row cites when its claim is finer-grained than this table's key. studies.csv is keyed on
(rsid, pmid), so when a claim is per-bin or per-genotype rather than per-variant, the citation lives
on the claiming row and this table describes the study instead. The two are not alternatives; asking
which one a citation belongs in is asking how specific the claim is.
The symptom of confusing it with literature.csv is a dropped row. literature.csv is
machine-produced and keeps every article it found; the compiler discards an uncited literature row
when nothing references it. studies.csv is the authored side and is never dropped. If a citation
disappeared from a compile, it was in the wrong one of the two.
Identity¶
| Row model | StudyRow (just_dna_format.spec) |
| Becomes | studies.parquet |
| Authored or derived | authored — a person writes it (a drafter may append rows) |
| Draftable | yes — draft can append rows |
| Drafted by | no drafting provider targets this table — draft writes no row of it |
| Natural key | variant_key, pmid |
| Fact signature | no |
| In the attestation binding | yes — manifest.inputs[] |
Columns¶
| Column | Type | Required | Values | Meaning |
|---|---|---|---|---|
rsid |
str | None |
optional | dbSNP identifier or variant key | |
chrom |
str | None |
optional | Chromosome (for position-only variants) | |
start |
int | None |
optional | 1-based position, VCF POS convention (position-only variants) — do not subtract one | |
ref |
str | None |
optional | Reference allele (position-only variants) | |
pmid |
str |
required | PubMed ID or reference — free-form, must be non-empty | |
population |
str | None |
optional | Study population | |
p_value |
str | None |
optional | Raw p-value string (free-form) | |
conclusion |
str | None |
optional | Study-specific conclusion | |
study_design |
str | None |
optional | e.g. meta-analysis, GWAS | |
stat_significance |
str | None |
optional | one of: not_significant, significant, suggestive, unknown |
Per-study statistical significance: significant|suggestive|not_significant|unknown. |
effect_size |
float | None |
optional | Per-study effect magnitude (unit given by effect_measure). |
|
effect_measure |
str | None |
optional | suggested: HR, NR, OR, RR, beta, log(HR), log(OR) |
Unit of effect_size, e.g. OR|HR|beta|RR (recommended, open). |
effect_allele |
str | None |
optional | The allele this study's effect_size is stated relative to — bases, or a symbolic/structural allele carrying its length (e.g. |
|
trait_efo_id |
str | None |
optional | EFO/MONDO/OBA/HP trait ontology id(s) for this study. | |
statistical_test |
str | None |
optional | Which analysis produced this row's p_value/effect_size — the test or model, and what it was adjusted for, e.g. Fisher's exact (allelic) or logistic regression adjusted for age and sex. Free text. study_design describes the study; this describes the analysis, and one study may report several. Absent means the paper's analysis was not recorded, never that it had only one. |
|
confidence |
str | None |
optional | How far the citing source stands behind this evidence link, in ITS OWN units and unconverted (CIViC's submitted/accepted, a review-star count). Meaningless without confidence_unit beside it, and refused without one. |
|
confidence_unit |
str | None |
optional | Which instrument confidence is measured on, e.g. civic_evidence_status. Required whenever confidence is set: a magnitude with no unit beside it is a value nothing can read, and this format has paid for that once already on weight. |
|
doi |
str | None |
optional | Digital Object Identifier — wider than pmid (covers preprints/books/datasets with no PubMed id). Free-form, kept verbatim; a validator may cross-fill doi↔pmid. |
|
provenance_quote |
str | None |
optional | Optional keyword phrase / literal passage locating this study's claim in the cited article's fulltext. Human-legible; a validator confirms fulltext-contains, yes/no. | |
provenance_regex |
str | None |
optional | Optional regex locating the claim in fulltext — a declarative pattern grammar (Principle 1: data, not code), matched consumer-side by a linear-time/ReDoS-safe engine. | |
curator |
str | None |
optional | Who located this row's provenance quote/regex — a name, handle, or model id, resolvable against the manifest's authorship. Row-level because real work is mixed at row granularity: a human may read a review while an agent traverses its citations, in one module, in one pass. Records the distribution of labour so a reviewer can route scrutiny; it does not move responsibility, which the human author holds regardless. |
|
p_value_num |
float | None |
optional | The p-value as a number, so it can be sorted and thresholded — p_value above is a free-form string and cannot be. In (0, 1]: a p-value of exactly 0 is a source's own underflow, not a probability, so it is rejected rather than stored as a confident zero. |
Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.