variant_key |
str |
required |
|
The variant half of this row's identity. Joins weights.parquet.variant_key where the module happens to author the variant — and most rows will not be joined at all, because an interval query answers for every scored variant in the window and the distal ones are what the query was for. |
rsid |
str | None |
optional |
|
The rsID, when one is known for this position. Null is the common case and not a defect: the Atlas answers by coordinate and names no rsID at all, so this is filled only where the module or its resolution.csv already knew one. |
chrom |
str | None |
optional |
|
Contig, as the module spells it — 22, not chr22. The Atlas ships chr22 and the conversion happens in the pass, which is where RM193 found a silent-wrong-answer bug: a join that matched nothing looked exactly like an uncovered region. |
start |
int | None |
optional |
|
The 1-based VCF position (@start-1based). Bounded ge=0 rather than ge=1 because VCF permits POS 0; the Atlas speaks 0-based intervals and 1-based variants, and atlas_client converts at its own boundary so nothing downstream has to know. |
ref |
str | None |
optional |
|
Reference allele, as the source scored it. |
alt |
str | None |
optional |
|
The single alternate allele this score is for. One ALT per row, not a list: the score is per substitution, and three ALTs at one position are three different claims. |
gene |
str |
required |
|
HGNC symbol, as AlphaGenome attributed it — GeneScorerMetadata.name from the response, never a symbol looked up from a span the caller drew (@gene-map-is-another-sources-attribution). The gene half of this row's identity. |
gene_id |
str | None |
optional |
|
Ensembl gene accession, verbatim including the version — ENSG00000040608.14, not ENSG00000040608 (@verbatim-except-order). Upstream's own proto comment says this field is 'without version number' and the live service returns it versioned, which is why the version is recorded rather than trusted away: a consumer joining on bare accessions truncates on purpose, and one who believed the comment would have written a join that silently matches nothing. |
effect_size |
float | None |
optional |
|
The magnitude, aggregated over the scorer's tissue tracks for this gene. Its meaning is effect_measure; it has no unit, and effect_unit says so rather than guessing one. |
effect_measure |
str | None |
optional |
|
What kind of magnitude it is — the scorer that produced it, RNA_SEQ in practice. Open prose rather than a vocabulary: the Atlas publishes twenty-two scorers and RM200 adopted one, so closing this would make adopting a second a schema change rather than a pass change. |
effect_unit |
str | None |
optional |
|
The unit the magnitude is in, verbatim from the source when it states one. Null for every AlphaGenome row, and that is the honest record rather than a hole: the service publishes no unit for a scorer's output, so these scores are comparable within the scorer and across nothing else. Present so a future source that does state a unit has somewhere to put it (@weight-has-no-unit). |
effect_direction |
str | None |
optional |
one of: decrease, increase |
Whether the variant increases or decreases this gene's predicted expression — the majority sign across the scorer's tracks. Not clinical direction: VariantRow.direction is protective|risk|neutral|unknown and answers a different question, since raising a gene may be good, bad or neither. Shares VALID_EFFECT_DIRECTIONS with GwasEffectRow because it is the same axis — the sign of a magnitude, not a judgement. |
tracks_agreeing |
int | None |
optional |
|
How many of tracks_total carried effect_direction's sign. A record, never a confidence: RM200 measured consensus against effect size across six variants and found it flat — CAGE is 97% unanimous at a PHRED of 0.007 — so a high count says the tracks agree and says nothing about whether the variant matters. |
tracks_total |
int | None |
optional |
|
How many tissue tracks the scorer reported for this gene — 371 for RNA_SEQ today. Stored beside tracks_agreeing rather than divided into it, because a fraction cannot say that its denominator moved. |
distance_to_gene |
int | None |
optional |
|
Base pairs from the variant to the gene's span, 0 inside it. The reason this column exists is that a threshold without it is wrong: AlphaGenome attributes a variant to genes across the model's whole 1 Mb input window, reaching +/-512 kb and stopping dead beyond, and distal scores run ~10x lower than scores at the gene — so a flat --min-score silently keeps only the proximal rows, which is exactly the failure RM194 exists to prevent. Null is unknown, not zero: the span comes from the MANE lane, and a deployment without that lane provisioned records the absence rather than a distance it could not compute. |
dataset |
str |
required |
|
Which query produced this row, e.g. 'alphagenome_atlas_2026-09-11'. A FACT, and dated rather than versioned because the service publishes no release label: AlphaGenome's Output Terms pin the applicable terms version to the date the Output was generated, so the date is a real property of the row rather than a stand-in for a missing one. |
source |
str | None |
optional |
|
The licensed data source: alphagenome_atlas|manual (open). Joins sources.csv.source. Never alphagenome_avi — one name cannot carry two licence classes, and the AVI artifact is Permissive Use while this scorer's output is non-commercial. |
status |
str | None |
optional |
one of: ambiguous, not_found, resolved |
Outcome: resolved|not_found|ambiguous. not_found is a FACT — the interval was queried and the service attributed this variant to no gene — and differs from a variant never queried, which has no row at all (@unreachable-not-absent). |
fetched_at |
str | None |
optional |
|
ISO-8601 UTC timestamp, second resolution. Canonicalized on load; records when this row was written by a pass, not when AlphaGenome generated anything — that is dataset. |