just_dna_format.expression¶
just_dna_format.expression ¶
The per-gene expression-effect table (0.7, RM194/RM200).
expression_effects.csv is the tenth derived-fact sidecar. It records, per (variant, gene) pair,
which way a variant moves that gene's predicted expression, how large the move is, how many of the
source's tissue tracks agree on the direction, and how far the variant sits from the gene. Filled by
just-dna-enricher's alphagenome expression pass, consumed and hashed by the compiler, never
fetched by it.
Why it exists. The AlphaGenome AVI artifact RM191 adopted is one number per variant with the
sign discarded — its eighteen features are all MAX_ABS_*, so it can say a variant matters and
cannot say which way. The Atlas API has the direction that artifact threw away, and RM200 measured
all twenty-two scorers to find which of it is recordable. Exactly one survived: RNA_SEQ is the only
scorer with a gene axis, the only one whose disagreement across tracks is meaningful rather than
flat, and the only one that attributes its own claim to a gene — which is what
@gene-map-is-another-sources-attribution requires, since a gene claim must come from a source's own
per-record attribution and never from a span the caller drew. Every other scorer either restates
AVI_SCORE or ranks noise: the top-5 of CAGE's 546 tracks carry 2–7% of the effect, and the
concentration runs backwards to effect size.
Why a derived sidecar rather than an authored column. Atlas Output is non-commercial (only the
AVI SNV scores are Permissive Use), and RM193's position is that non-commercial Output enters as a
finding and never as a stored value. A sidecar keeps that rule literally — nothing here is
authored, nothing enters content_signature, and the compile gate still reads sources.csv and
nothing else — while carrying an axis a finding cannot: per-gene direction survives, and a finding
would have to collapse it to prose. Half cost under Principle 9 rather than full.
One row is one (variant, gene) pair. Not one per variant: the whole reason to keep the gene
axis is that a variant genuinely raises one gene and lowers another, which is also why RNA_SEQ
never reaches consensus across its tracks (52–61% throughout, measured over six variants spanning
four decades of effect size). That disagreement is the signal, so a summary that collapses genes
would destroy the thing that makes this scorer worth having. And not one row per track: 371 tissue
tracks per gene is lossless and unreadable, and a table a human can never open is not a table this
format ships.
This row carries coordinates, and gwas_effects.csv deliberately does not — same reason, opposite
outcome. There, a coordinate would have to be copied from the module's own resolution.csv,
making it the module's fact rather than the source's. Here the coordinate is the source's fact:
the Atlas answers an interval query with the variant it scored, and the rows are locus-wide by
design — most of them name variants the module does not author, because finding the distal ones is
the entire point of the item. A row with no coordinate would be unjoinable to anything.
What is deliberately not here.
tracks_agreeingis a count, never a confidence. RM200 tested consensus-as-confidence and killed it:CAGEandPROCAPare near-unanimous at every variant, 97% agreement at aPHREDof 0.007, so agreement does not separate a consequential variant from an inconsequential one. The count is recorded because a reader may want it; nothing in this tier gates on it, and no threshold withholds a direction.- No
gene_strandcolumn. Strand is a property of the gene, recoverable from any gene annotation, and it is not this source's fact about this variant. Recording it beside a signed score would invite a consumer to "correct" the sign for a minus-strand gene, which would invert the claim. It is an additive optional column if a caller ever needs it (Principle 3). - No unit on
effect_size, stated rather than invented. AlphaGenome publishes no unit for a scorer's output, soeffect_unitisNoneandeffect_measurenames the scorer instead. That is@weight-has-no-unitobeyed rather than dodged: the rule is that a magnitude needs its unit beside it, and where no unit exists the honest record is the absence plus the name of what produced the number. Scores are comparable within the scorer and across nothing else.
ExpressionEffectRow ¶
Bases: BaseModel
One (variant, gene) pair: which way the variant moves that gene, and how far away it is.
Standalone (not an AuthoredModel) for the same reason every fact row is — machine-produced
reference fact, not an authored annotation — with extra="forbid" so a typo'd column is caught
rather than silently dropped.