Skip to content

gene_validity.csv

Identity

Row model GeneValidityRow (just_dna_format.gene_validity)
Becomes gene_validity.parquet
Authored or derived derived — an enricher pass writes it; an author corrects it via overrides.csv
Draftable no
Written and checked by just-dna-enricher gene-validity (gene_validity) from clingen, gencc
Natural key assertion_id
Fact signature yes — its sidecar carries one
In the attestation binding no

Columns

Column Type Required Values Meaning
gene str required HGNC-style symbol, matching the gene column authored in variants.csv
gene_id str | None optional HGNC id (HGNC:20) — the stable identity behind the mutable symbol, which both sources publish. Carried for the reason GeneMetricsRow.gene_id carries the ENSG: symbols are renamed and the id is not, so a module authored against an old symbol still matches.
disease_id str | None optional CURIE for the disease term, as the source states it — MONDO:0013212 from ClinGen and GenCC, OMIM:…/ORPHA:… from a source that publishes those. Stored verbatim: a CURIE is an identity, and rewriting one across ontologies is a claim this tier cannot make.
disease_label str | None optional The term's human-readable name as the source publishes it, so the table is legible without resolving the CURIE. Descriptive, never the join key, and deliberately OUTSIDE the fact hash: one real export carries MONDO:0017146 under two labels at once, so a label says when it was read rather than what was asserted.
moi str | None optional one of: autosomal_dominant, autosomal_recessive, mitochondrial, semidominant, undetermined, x_linked, x_linked_dominant, x_linked_recessive, y_linked Mode of inheritance the assertion is scoped to: autosomal_dominant|autosomal_recessive|x_linked|x_linked_dominant|x_linked_recessive|y_linked|mitochondrial|semidominant|undetermined. Part of the KEY, not decoration — 59 ClinGen (gene, disease) pairs carry two rows differing only here. undetermined is a stated finding; empty means the source has no such concept.
classification str | None optional one of: animal_model_only, definitive, disputed, limited, moderate, no_known_disease_relationship, refuted, strong, supportive Strength of the gene–disease assertion: definitive|strong|moderate|limited|supportive|disputed|refuted|no_known_disease_relationship|animal_model_only. Normalized from the submitter's own wording at the enricher boundary. Empty where the source asserts an association without grading it — which is NOT no_known_disease_relationship, a graded verdict against. A FACT, never this workspace's opinion.
classification_raw str | None optional The submitter's verbatim wording (Disputed Evidence, Definitive), kept so the mapping above stays auditable and a term this release does not model is still visible. Same role clin_sig_raw plays beside clin_sig.
classification_date str | None optional ISO-8601 UTC timestamp of the curation itself — when the panel ruled, not when this row was written (that is fetched_at). Inside the fact set: a re-curation of the same pair is a new fact.
submitter str | None optional Who made the assertion — 'Charcot-Marie-Tooth Disease Gene Curation Expert Panel' from ClinGen, 'Ambry Genetics' from GenCC. Part of the key on an aggregate: one gene-disease pair routinely carries several submitters at different strengths, and they are all data.
assertion_id str | None optional The source's own stable id for this assertion (ClinGen's CGGV:assertion_…, GenCC's uuid). The identity half of report_url, which is why that one is outside the fact set and this one is inside.
report_url str | None optional Where a human can read the curation. Outside the fact set — a location, not a fact.
dataset str required Which release this assertion is from, e.g. 'clingen_gene_validity_2026-08-13'. A FACT, for the reason it is one on every sibling table: two releases are two facts.
source str | None optional The licensed data source this assertion came from: clingen|gencc|manual|reversed (open). Joins sources.csv.source. It names the SOURCE, never the route — the release and the route are dataset's job, which is why that column is inside the fact set and this is not.
status str | None optional one of: ambiguous, not_found, resolved Outcome: resolved|not_found|ambiguous (the ResolutionRow vocabulary)
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-13T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the curation was made — that is classification_date.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.