Skip to content

frequencies.csv

Identity

Row model FrequencyRow (just_dna_format.frequency)
Becomes frequencies.parquet
Authored or derived derived — an enricher pass writes it; an author corrects it via overrides.csv
Draftable no
Written and checked by just-dna-enricher frequencies (frequencies) from gnomad
Natural key variant_key, population
Fact signature yes — its sidecar carries one
In the attestation binding no

Columns

Column Type Required Values Meaning
variant_key str required Coordinate-derived identity of the allele (matches the post-expansion weights key)
rsid str | None optional dbSNP identifier, when known
chrom str | None optional Chromosome without 'chr' prefix
start int | None optional 1-based genomic position (VCF POS convention)
ref str | None optional Reference allele
alt str | None optional The ONE alt allele this frequency is for — deliberately singular, unlike ResolutionRow.alts: an allele frequency is per-allele, so a multi-allelic site is several rows, never one comma-joined cell.
genome_build str defaulted Assembly the coordinate is in (the RM15 forward hook)
population str required Ancestry group the counts are for, or 'global' for the whole dataset. Open, seeded vocabulary (vocab.RECOMMENDED_ANCESTRY_GROUPS) — a label is interpretable only together with dataset.
allele_count int | None optional AC — observed copies of alt in this group
allele_number int | None optional AN — total called alleles in this group (the denominator)
homozygote_count int | None optional Individuals homozygous for alt in this group
hemizygote_count int | None optional Hemizygous calls (X/Y outside the PAR, and MT)
faf95 float | None optional Filtering allele frequency, 95% CI lower bound — the number an ACMG BA1/BS1 filter actually uses. The source reports ONE faf95 with a named owning group, so it is set only on that group's row and null everywhere else (no extra column, no overloaded field).
dataset str required Which release these counts are from, e.g. 'gnomad_v4.1_joint'. A FACT, not provenance: the same allele has different (equally correct) counts in different releases.
vrs_id str | None optional GA4GH VRS allele id (ga4gh:VA.…)
caid str | None optional ClinGen Allele Registry canonical allele id (CA<digits>)
source str | None optional Which link filled this: gnomad|manual|reversed (open)
status str | None optional one of: not_covered, not_found, resolved Outcome: resolved|not_found|not_covered (VALID_FREQUENCY_STATUS)
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-03T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the source published anything

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.