Skip to content

just_dna_format.expression

just_dna_format.expression

The per-gene expression-effect table (0.7, RM194/RM200).

expression_effects.csv is the tenth derived-fact sidecar. It records, per (variant, gene) pair, which way a variant moves that gene's predicted expression, how large the move is, how many of the source's tissue tracks agree on the direction, and how far the variant sits from the gene. Filled by just-dna-enricher's alphagenome expression pass, consumed and hashed by the compiler, never fetched by it.

Why it exists. The AlphaGenome AVI artifact RM191 adopted is one number per variant with the sign discarded — its eighteen features are all MAX_ABS_*, so it can say a variant matters and cannot say which way. The Atlas API has the direction that artifact threw away, and RM200 measured all twenty-two scorers to find which of it is recordable. Exactly one survived: RNA_SEQ is the only scorer with a gene axis, the only one whose disagreement across tracks is meaningful rather than flat, and the only one that attributes its own claim to a gene — which is what @gene-map-is-another-sources-attribution requires, since a gene claim must come from a source's own per-record attribution and never from a span the caller drew. Every other scorer either restates AVI_SCORE or ranks noise: the top-5 of CAGE's 546 tracks carry 2–7% of the effect, and the concentration runs backwards to effect size.

Why a derived sidecar rather than an authored column. Atlas Output is non-commercial (only the AVI SNV scores are Permissive Use), and RM193's position is that non-commercial Output enters as a finding and never as a stored value. A sidecar keeps that rule literally — nothing here is authored, nothing enters content_signature, and the compile gate still reads sources.csv and nothing else — while carrying an axis a finding cannot: per-gene direction survives, and a finding would have to collapse it to prose. Half cost under Principle 9 rather than full.

One row is one (variant, gene) pair. Not one per variant: the whole reason to keep the gene axis is that a variant genuinely raises one gene and lowers another, which is also why RNA_SEQ never reaches consensus across its tracks (52–61% throughout, measured over six variants spanning four decades of effect size). That disagreement is the signal, so a summary that collapses genes would destroy the thing that makes this scorer worth having. And not one row per track: 371 tissue tracks per gene is lossless and unreadable, and a table a human can never open is not a table this format ships.

This row carries coordinates, and gwas_effects.csv deliberately does not — same reason, opposite outcome. There, a coordinate would have to be copied from the module's own resolution.csv, making it the module's fact rather than the source's. Here the coordinate is the source's fact: the Atlas answers an interval query with the variant it scored, and the rows are locus-wide by design — most of them name variants the module does not author, because finding the distal ones is the entire point of the item. A row with no coordinate would be unjoinable to anything.

What is deliberately not here.

  • tracks_agreeing is a count, never a confidence. RM200 tested consensus-as-confidence and killed it: CAGE and PROCAP are near-unanimous at every variant, 97% agreement at a PHRED of 0.007, so agreement does not separate a consequential variant from an inconsequential one. The count is recorded because a reader may want it; nothing in this tier gates on it, and no threshold withholds a direction.
  • No gene_strand column. Strand is a property of the gene, recoverable from any gene annotation, and it is not this source's fact about this variant. Recording it beside a signed score would invite a consumer to "correct" the sign for a minus-strand gene, which would invert the claim. It is an additive optional column if a caller ever needs it (Principle 3).
  • No unit on effect_size, stated rather than invented. AlphaGenome publishes no unit for a scorer's output, so effect_unit is None and effect_measure names the scorer instead. That is @weight-has-no-unit obeyed rather than dodged: the rule is that a magnitude needs its unit beside it, and where no unit exists the honest record is the absence plus the name of what produced the number. Scores are comparable within the scorer and across nothing else.

ExpressionEffectRow

Bases: BaseModel

One (variant, gene) pair: which way the variant moves that gene, and how far away it is.

Standalone (not an AuthoredModel) for the same reason every fact row is — machine-produced reference fact, not an authored annotation — with extra="forbid" so a typo'd column is caught rather than silently dropped.