variant_key |
str |
required |
|
Coordinate-derived identity of the allele (matches the post-expansion weights key) |
rsid |
str | None |
optional |
|
dbSNP identifier for this allele's locus, as the resolution table already established it. Descriptive: an rsID is position/multi-allelic-level, so it names the locus rather than this record's allele, and variant_key is what identifies the row. |
chrom |
str | None |
optional |
|
Chromosome without 'chr' prefix |
start |
int | None |
optional |
|
1-based genomic position (VCF POS convention) |
ref |
str | None |
optional |
|
Reference allele |
alt |
str | None |
optional |
|
The ONE alt allele this record is about — singular like FrequencyRow.alt and unlike ResolutionRow.alts, because a clinical call is per-allele: one rsID at one locus can carry a pathogenic, a benign and an uncertain allele (rs33922842 in HBB does). |
genome_build |
str |
defaulted |
|
Assembly the coordinate is in. Load-bearing rather than decorative: the ClinVar snapshot is GRCh38 and its lookup key carries no assembly, so a coordinate from another build is a well-formed query that returns a different variant's record. |
clin_sig |
str | None |
optional |
one of: affects, association, benign, conflicting, drug_response, likely_benign, likely_pathogenic, not_provided, other, pathogenic, protective, risk_factor, uncertain_significance |
The archive's clinical significance, normalized to the module vocabulary (vocab.VALID_CLIN_SIG) at the enricher boundary. A FACT about the archive, never an adjudication: this table records the call, it does not decide who is right. |
clin_sig_raw |
str | None |
optional |
|
The archive's verbatim token (Pathogenic/Likely_pathogenic), kept so the normalization above stays auditable and a value the mapping does not model is still visible. |
review_status |
str | None |
optional |
|
The archive's own review wording, verbatim — e.g. 'criteria_provided,_multiple_submitters,_no_conflicts'. Open, not a vocabulary: it is ClinVar's phrasing and it changes on ClinVar's schedule, so closing it here would make a future release unloadable for no gain. |
review_stars |
int | None |
optional |
|
The 0-to-4 rating that wording maps to. Stored beside the prose rather than derived from it, because the mapping is a ClinVar convention and this tier holds no source conventions (Principle 2) — the enricher owns it. Null means the archive stated no review status; 0 is a rating ('no assertion criteria provided') and is not the same thing. |
condition |
str | None |
optional |
|
The condition the call is scoped to, as the archive names it. Descriptive: the archive holds several records per allele under different conditions, and this is what separates them for a reader. |
variation_id |
str | None |
optional |
|
The archive's stable record id (ClinVar's VariationID). The thing a consumer cites, and the join key back into the archive — which is why it is in the fact set. |
dataset |
str |
required |
|
Which release this record is from, e.g. 'clinvar_2026-06-27'. A FACT: a re-reviewed record is a new fact, and telling the two apart is the point of this table. |
source |
str | None |
optional |
|
The licensed data source: clinvar|manual|reversed (open). Joins sources.csv.source. It names the SOURCE, never the route — which release answered is dataset's job. |
status |
str | None |
optional |
one of: ambiguous, not_found, resolved |
Outcome: resolved|not_found|ambiguous (the ResolutionRow vocabulary). not_found is a FACT — the archive was consulted and has no record for this allele — and is different from an allele that was never queried, which has no row at all. |
fetched_at |
str | None |
optional |
|
ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-13T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the archive published anything — that is dataset. |