Skip to content

clinical_assertions.csv

Identity

Row model ClinicalAssertionRow (just_dna_format.assertions)
Becomes clinical_assertions.parquet
Authored or derived derived — an enricher pass writes it; an author corrects it via overrides.csv
Draftable no
Written and checked by just-dna-enricher assertions (assertions) from clinvar
Natural key variant_key, variation_id
Fact signature yes — its sidecar carries one
In the attestation binding no

Columns

Column Type Required Values Meaning
variant_key str required Coordinate-derived identity of the allele (matches the post-expansion weights key)
rsid str | None optional dbSNP identifier for this allele's locus, as the resolution table already established it. Descriptive: an rsID is position/multi-allelic-level, so it names the locus rather than this record's allele, and variant_key is what identifies the row.
chrom str | None optional Chromosome without 'chr' prefix
start int | None optional 1-based genomic position (VCF POS convention)
ref str | None optional Reference allele
alt str | None optional The ONE alt allele this record is about — singular like FrequencyRow.alt and unlike ResolutionRow.alts, because a clinical call is per-allele: one rsID at one locus can carry a pathogenic, a benign and an uncertain allele (rs33922842 in HBB does).
genome_build str defaulted Assembly the coordinate is in. Load-bearing rather than decorative: the ClinVar snapshot is GRCh38 and its lookup key carries no assembly, so a coordinate from another build is a well-formed query that returns a different variant's record.
clin_sig str | None optional one of: affects, association, benign, conflicting, drug_response, likely_benign, likely_pathogenic, not_provided, other, pathogenic, protective, risk_factor, uncertain_significance The archive's clinical significance, normalized to the module vocabulary (vocab.VALID_CLIN_SIG) at the enricher boundary. A FACT about the archive, never an adjudication: this table records the call, it does not decide who is right.
clin_sig_raw str | None optional The archive's verbatim token (Pathogenic/Likely_pathogenic), kept so the normalization above stays auditable and a value the mapping does not model is still visible.
review_status str | None optional The archive's own review wording, verbatim — e.g. 'criteria_provided,_multiple_submitters,_no_conflicts'. Open, not a vocabulary: it is ClinVar's phrasing and it changes on ClinVar's schedule, so closing it here would make a future release unloadable for no gain.
review_stars int | None optional The 0-to-4 rating that wording maps to. Stored beside the prose rather than derived from it, because the mapping is a ClinVar convention and this tier holds no source conventions (Principle 2) — the enricher owns it. Null means the archive stated no review status; 0 is a rating ('no assertion criteria provided') and is not the same thing.
condition str | None optional The condition the call is scoped to, as the archive names it. Descriptive: the archive holds several records per allele under different conditions, and this is what separates them for a reader.
variation_id str | None optional The archive's stable record id (ClinVar's VariationID). The thing a consumer cites, and the join key back into the archive — which is why it is in the fact set.
dataset str required Which release this record is from, e.g. 'clinvar_2026-06-27'. A FACT: a re-reviewed record is a new fact, and telling the two apart is the point of this table.
source str | None optional The licensed data source: clinvar|manual|reversed (open). Joins sources.csv.source. It names the SOURCE, never the route — which release answered is dataset's job.
status str | None optional one of: ambiguous, not_found, resolved Outcome: resolved|not_found|ambiguous (the ResolutionRow vocabulary). not_found is a FACT — the archive was consulted and has no record for this allele — and is different from an allele that was never queried, which has no row at all.
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-13T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the archive published anything — that is dataset.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.