Skip to content

gene_metrics.csv

Identity

Row model GeneMetricsRow (just_dna_format.gene_metrics)
Becomes gene_metrics.parquet
Authored or derived derived — an enricher pass writes it; an author corrects it via overrides.csv
Draftable no
Written and checked by just-dna-enricher gene-metrics (gene_metrics) from gnomad — constraint and dosage sensitivity are two different measurements of one gene from two authorities, and they share a table because a consumer asks one question of it
just-dna-enricher dosage (clingen) from clingen — ClinGen's haploinsufficiency and triplosensitivity calls are gene metrics like gnomAD's constraint, and are CC0, which is what keeps a module carrying them sellable
Natural key gene, dataset
Fact signature yes — its sidecar carries one
In the attestation binding no

Columns

Column Type Required Values Meaning
gene str required HGNC-style symbol, matching the gene column authored in variants.csv
gene_id str | None optional Ensembl gene id (ENSG…) — the stable identity behind the mutable symbol. Carried because symbols are aliases that get renamed while the ENSG does not, so a module authored against an old symbol can still be matched.
transcript str | None optional Ensembl transcript (ENST…) the metrics were computed on
mane_select bool | None optional Whether transcript is the MANE Select transcript. Load-bearing for reproducibility, not decoration: the source is per-transcript and the row pick must be deterministic.
pli float | None optional Probability of being loss-of-function intolerant
loeuf float | None optional LoF observed/expected upper bound fraction — the source's oe_lof_upper, stored under the name clinical readers actually ask for it by. oe_lof/oe_lof_lower sit beside it so the point estimate and the full interval are never lost.
oe_lof float | None optional LoF observed/expected ratio
oe_lof_lower float | None optional Lower bound of the LoF o/e 90% CI
lof_z float | None optional LoF constraint Z score
obs_lof int | None optional Observed LoF variant count
exp_lof float | None optional Expected LoF variant count
oe_mis float | None optional Missense observed/expected ratio
mis_z float | None optional Missense constraint Z score
syn_z float | None optional Synonymous constraint Z score — near zero for a well-behaved gene, so it doubles as a sanity check
constraint_flags str | None optional The source's own caveat list, pipe-joined and sorted (e.g. 'no_exp_lof', 'outlier_mis|outlier_syn'), or absent when the source flagged nothing — never an empty string and never an empty container, so if row.constraint_flags: is the right test. The flag TOKENS are the source's, verbatim; the container is not, because gnomAD spells the same list two ways depending on which route answered (a JSON array from the live API, its array literal in the bulk TSV cell) and this column is inside GENE_METRICS_FACT_FIELDS — one gene fetched two ways would otherwise carry two signatures. A flagged gene's scores are not to be read at face value, and folding that warning away would be the format editorializing over its source; normalizing how the list is written down is not folding it away.
haploinsufficiency str | None optional one of: autosomal_recessive, dosage_sensitivity_unlikely, little_evidence, no_evidence, some_evidence, sufficient_evidence ClinGen haploinsufficiency rating: no_evidence|little_evidence|some_evidence|sufficient_evidence|autosomal_recessive|dosage_sensitivity_unlikely. NOT an ordinal — see VALID_DOSAGE_SENSITIVITY. A FACT.
triplosensitivity str | None optional one of: autosomal_recessive, dosage_sensitivity_unlikely, little_evidence, no_evidence, some_evidence, sufficient_evidence ClinGen triplosensitivity rating, same vocabulary. Empty where ClinGen says 'Not yet evaluated' — an absence, not a rating. A FACT.
dataset str required Which release these metrics are from, e.g. 'gnomad_v4.1_constraint'. A FACT.
source str | None optional The licensed data source these metrics came from: gnomad|clingen|manual|reversed (open). Joins sources.csv.source. It names the SOURCE, not the route — which release and which route answered is dataset's job, and a v2.1.1 API figure and a v4.1 bulk figure are different facts precisely because dataset is inside the fact set and this column is not.
status str | None optional one of: ambiguous, not_found, resolved Outcome: resolved|not_found (the ResolutionRow vocabulary)
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-03T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the source published anything

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.