just_dna_format.binning¶
just_dna_format.binning ¶
The measure → phenotype binning primitive (0.4 — see docs/CHANGELOG.md).
One declarative shape shared by every quantity-carrying locus: a per-locus table that maps a measured quantity (activity score, copy number, repeat count, heteroplasmy fraction, PRS percentile) to a phenotype by range. The tables differ only in which quantity is measured and in their explicit key columns (multicolumn keying — never a packed tuple; the keying stance is a coding standard, see CLAUDE.md). Aligning the column vocabulary gives a consumer one "bin-a-measure" code path.
Data-agnostic (design north star — see CLAUDE.md). These rows are pure annotation: a lookup table declaring range→phenotype. The module contains no measurement — the measured quantity is supplied by the consumer at query time; the table never sees a sample. The bins themselves are a generalization over a practical subset of real loci/ranges, not an all-encompassing model, so a data item that doesn't fit is a schema gap to widen additively.
Ranges are inclusive [measure_min, measure_max]: min == max is a sharp value (e.g. exactly
0 copies), min < max is a range (HTT 36–39 CAG), and measure_max = None is open-ended (≥40 CAG,
3+ copies). There is no copy_number column — a sharp copy number is measure_min == measure_max.
On a continuous measure, two adjacent bins may share an endpoint, and the higher bin owns it.
The lookup rule, which a consumer implements once: select the row with the greatest measure_min ≤ x
(within the group). Written out for allele_fraction bins 0.0–0.1, 0.1–0.3, 0.3–1.0, a
heteroplasmy of exactly 0.1 selects the MIDD bin and 1.0 selects the top one.
This is a rule about tiling, not a second meaning for measure_max, and it exists because the
alternative was unsatisfiable (RM35). Inclusive-at-both-ends, overlap-is-an-error and
any-positive-hole-is-a-warning cannot all hold on a dense domain: two adjacent continuous bins either
share an endpoint or leave a gap, so every allele_fraction/prs_percentile table carried a finding
forever — a check that could not be satisfied rather than one that was failing. No epsilon escapes it
([0, 0.0999999] + [0.1, 1.0] still warns). Half-open [min, max) for continuous kinds was the other
candidate and lost on authorship: it makes one column mean two things depending on measure_kind, the
number written in the cell is then not in the bin, and a bounded domain's top value (AF = 1.0 is
homoplasmy, and real) becomes unreachable unless the last bin is authored open. Here measure_max means
the same thing on every kind and the top bin stays closed.
Discrete kinds tile the other way, and that is still their default. repeat_count/copy_number
tile exactly under inclusive bounds — HTT [6,35], [36,39], [40,∞) is gapless if the domain is
integral — so for them a shared endpoint is a real overlap and an error unless the table says
otherwise.
…except that the domain is not integral, and the spec says so (RM55). VCF 4.4 §7.2 "Redefined INFO
and FORMAT CN to support non-integer copy numbers" and its worked examples are fractional throughout
(CN=3,0.9666, CN=1.25); §5.6 leaves the granularity of a copy number deliberately undefined and
allows a segment mean "at a highly granular megabase level of resolution". §3 types RUC, the repeat
count VCF 4.4 standardises, as a Float. So the premise the paragraph above rests on was withdrawn
for both kinds, and the consequence was worse than the allele_fraction case RM35 fixed: a tiling
[0,0] [1,1] [2,2] [3,∞) was accepted with no warning at all, a measured 2.4 matched no bin, and the
module compiled green under --strict. A hole of exactly one is invisible to the gap check by
construction — and, the half nobody had written down, the schema also refused the tiling that would
fix it, since a shared endpoint on these kinds was an overlap error.
How 0.6 closes it: tiling becomes its own axis. measure_tiling ({quantised, continuous},
optional, VALID_MEASURE_TILINGS) says how the axis is divided, which is a different question from
measure_kind's what is measured — folding the two would be the overloaded-field anti-pattern (P5),
and a product rather than a sum. The shared-endpoint and gap rules read the group's effective
tiling (resolve_tiling): the declared value, else continuous where the kind would default to
quantised and the group carries a value no grid of whole numbers can hold, else the kind's default
(DEFAULT_MEASURE_TILING). Absence means the kind's default, never a value, which is what makes the
column additive — every module published before 0.6 keeps its exact meaning, and an author meets the
column only when departing from it. The inference announces itself, runs one way only —
fractional-ness contradicts a stated grid, integer-ness contradicts nothing, since [0,1] [2,3] is what
a continuous measure looks like when its author has only seen whole-number data — and fires only
against a quantised default, because that is the only reading a fraction falsifies. See
resolve_tiling for why activity_score, which is fractional by nature, must not be moved by it.
Three repairs were refused. Moving the kinds into _DENSE_KINDS is one line and silently re-reads every
published table ([2,2] beside [3,3] is a legal quantised tiling, and both rules change meaning under
dense semantics) with no notice and no way to say otherwise. A sixth measure_kind is the wrong axis,
above. And deriving the tiling from the rows with no column at all fails on the half that has to be
right: absence of a fractional value implies nothing, so it would read the ambiguous table one way,
silently, and leave the curator no way to correct it.
The one genuine int here is CopyNumberRow.modifier_cn, a modifier's dosage rather than a bound,
and it gets the parallel-float treatment the roadmap entry had proposed for the bounds:
modifier_copy_number beside it, read through effective_modifier_copy_number, both-set refused,
modifier_cn deprecated in 0.6 and removed at 1.0. So 1.0 inherits a removal, not a retype.
A measurement can also span several bins, and there is no state for that (RM56). RUC travels with
CIRUC and CN with CICN, and the spec is explicit that the upper bound may be missing and then means
unbounded (§3: "a reasonable limit of total length of the repeat could not be determined"). §5.7's
canonical CAG form is RUS=CAG;RUC=65;CIRUC=-15,., and says why: "Many of these techniques result in
imprecise variant calls." Imprecision is the normal case. The consumer contract below has exactly three
states — a bin matched, no bin matched, or the measurement absent — and none of them is the measurement
spans bins. reference_examples/htt_repeat_expansion has thresholds at 26/27, 35/36 and 39/40, so a
real RUC=38, CIRUC=-5,5 spans [33,43] and crosses all three: benign, uncertain and fully
penetrant, with no honest answer among them.
The policy vocabulary that settles it — withhold / take the worst bin / take the point estimate — lands
with the rest of the repeat work, when a real caller VCF is in hand, and its grain (per table or per row)
is deliberately undecided. Until then the stated placeholder behaviour is the house default: a
consumer that reads an interval spanning two or more bins withholds. It does not pick among them, and
it does not fall back to the unresolved sentinel either — that row means no measurement was
available, and here one was. Note what is not on the table: widening the measurement itself into an
interval puts a measurement in the module, which the data-agnostic north star forbids outright. What
belongs here is the rule for an interval that spans bins, which is annotation.
unresolved (T1) is mandatory. A table can state the outcome for measurement absent / not
callable, and the consumer contract is that a missing measurement selects the unresolved row,
never the lowest/reference bin (no activity score ⇒ not "Normal Metabolizer"; no CN ⇒ not "2
copies"; no heteroplasmy read ⇒ not "homoplasmic reference"). An unresolved row carries no bounds.
A measurement that is present but matches no bin is a distinct third state ("no matching bin", not
unresolved); validate_bins below rejects overlaps and flags coverage gaps so a table stays
coherent (consumer round-2 C1).
pmid grounds the BOUNDARY, and the line is: the bin row cites, the citation table describes
(RM47, 0.6). A threshold is the most interpretive claim this format carries — where 36 rather than 35
CAG becomes "reduced penetrance" is a clinical judgement drawn from a specific paper — and until 0.6
nothing could point at one: studies.csv identifies its subject by rsid or chrom, while a
repeat_alleles.csv row is keyed (gene, repeat_unit). One optional column on this base reaches all
four kinds and answers the question actually asked, which is why 36 rather than why this table.
It carries a pointer only. Everything about the paper — population, p_value_num, effect_size,
provenance_quote — stays in studies.csv, whose subject requirement was relaxed in the same release
so a citation row may exist without naming a variant. Copying that column set here one column at a
time would restate the bin inside its own evidence, which is exactly what killed the alternative
designs (a bin_evidence.csv join table keys on the thresholds, and they are floats).
source_field (round-2 3a) is a declarative pointer, not code. It optionally names the VCF
field the consumer extracts the measure from (FORMAT/REPCN, FORMAT/AF, INFO/CN|FORMAT/DS) —
pure indirection/addressing, deliberately constrained to a field-name key (optionally namespace-
qualified, optionally |-alternated) so it can never become an expression. That keeps it inside
Principle 1 (declarative, non-Turing): a name that says where the measurement lives, never a
transform that computes one. The module still holds no measurement.
A VCF field is identified by namespace and described by cardinality, and source_field used to
carry neither (RM53/RM54, 0.6). INFO and FORMAT are two reserved-key tables that collide on DP,
AD, ADF, ADR, MQ, AF and — since 4.4 — CN, so AF alone names the cohort frequency of an
ALT or this sample's fraction of it, and both are floats in [0, 1] that bin without complaint.
Number then decides how many values come back: a pointer at a Number=R field returns one value
per allele, reference first, of which none is the answer. So the namespace goes in the pointer
(FORMAT/AF, bare still legal and still meaning unqualified) and the element goes in source_element
— a closed set of named rules rather than an index, because AD[1] is the first line of an
expression grammar and Principle 1 refuses it.
"Element" is one of the values the field carries for a record, which is wider than a Number slot on
purpose: ExpansionHunter reports both repeat alleles in a single REPCN cell as 17/42, and a rule
that only spoke about Number would have nothing to say about the case it was built for
(reference_examples/htt_repeat_expansion, where the clinical rule is the larger of the two). How a
caller encodes multiplicity is the caller's business and this tier holds no opinion on it (Principle
2); which value the annotation means is the module's, and that is all this column states.
MeasureBinRow ¶
Bases: AuthoredModel
Base row of a binning table: a measured quantity range → the same orthogonal axes a
VariantRow carries. Subclasses add the explicit key columns for their quantity.
Inherits AuthoredModel (and passes it to its subclasses): extra="forbid" + the
reserved-namespace guard, and the shared direction/clin_sig/trait_efo_id validators.
ActivityPhenotypeRow ¶
Bases: MeasureBinRow
PGx metabolizer phenotype by activity score, per gene (CYP2D6 PM/IM/NM/UM). The score is a consumer call (Σ activity×copies over the diplotype); this table only bins it.
CopyNumberRow ¶
Bases: MeasureBinRow
Whole-gene dosage phenotype by copy number (SMN1 SMA). Sharp dosages are
measure_min == measure_max (0 copies = [0, 0]); 3+ is measure_min=3, measure_max=None.
Optional modifier_gene plus a modifier dosage express a second locus read in context (SMN1
phenotype depends on SMN2 copy number) — explicit named columns (multicolumn keying), never a
tuple. The gene and the dosage are set together or both left null.
The dosage has two spellings and one meaning (RM55, 0.6). modifier_copy_number is the
float column; modifier_cn is the original int one, deprecated in 0.6 and removed at 1.0.
VCF 4.4 §7.2 redefined CN to support non-integer copy numbers, so an int cannot hold a real
modifier dosage — and retyping the column is major-only, which is what the companion exists to
avoid. Everything reads effective_modifier_copy_number, including _KEY_FIELDS: a coalesced
value is one spelling by the time grouping and dedup see it, so the key never holds two
spellings of one number. Setting both is an error rather than a precedence rule.
effective_modifier_copy_number
property
¶
The modifier dosage, from whichever column carries it — the only reader of the pair.
is not None rather than or, because 0 is a legal dosage: SMN2 = 0 copies is a real
row, and a truthiness fallback would silently read it as "unset" and then fall through to
the deprecated column. (The boolean effective_pathogenic/effective_benign aliases are
the precedent here, not the string effective_direction, which may use or because an
empty string is not a legal value there.)
A fixed point by construction — coalescing an already-coalesced value returns it unchanged (P7 idempotency).
RepeatAlleleRow ¶
Bases: MeasureBinRow
VNTR/STR phenotype by repeat count, keyed on (gene, repeat_unit) — the motif is part of
the identity (T3): a count is only comparable within its motif definition. The count is a
consumer call (ExpansionHunter / adVNTR / a span genotyper) that MUST state the motif it
counted.
HeteroplasmyRow ¶
Bases: MeasureBinRow
mtDNA phenotype by heteroplasmy allele fraction (0–1), keyed on
(gene, reference_sequence, tissue). The reference sequence is part of the key (A3): rCRS/
NC_012920 vs legacy NC_001807 disagree and genome_build does not disambiguate. Bounds are
constrained to [0, 1].
tissue/assay_context are optional but load-bearing (round-2 Q6): heteroplasmy bins are
tissue-conditional — a blood-derived fraction systematically under-represents the
affected-tissue burden, and the penetrance threshold itself shifts by tissue, so the same
fraction bins to different phenotypes across tissues. A heteroplasmy table with no tissue context
is quietly unsafe; state the tissue the bins assume.
The variant identity is part of the key too (0.5.1), and it was missing. A mitochondrial gene
carries several pathogenic variants with genuinely different thresholds — MT-TL1 has m.3243A>G
and m.3271T>C, both causing MELAS; MT-ATP6 has m.8993T>G and m.9176T>C. Keyed on the gene alone,
their bins landed in one group and validate_bins rejected the module outright with "overlapping
bins", which is an error, not a warning, so the module could not compile at all. There was no
honest way out: trait_efo_id is in the group key and would have separated them, but only by
giving one disease two ontology ids. The documented example never showed two variants in a gene,
so the limitation was invisible rather than decided.
The columns mirror PharmVariantRow exactly — rsid, else chrom+start(+ref/alts) — and are
optional, so an existing single-variant table groups as it always did (P3/P8). They enter the
key through the derived variant_key property rather than one-by-one, so all the identity shapes
collapse to the format's own notion of which variant a row is about.
TilingResolution ¶
Bases: NamedTuple
How one bin group's tiling was decided, and from what — so a caller can say which it was.
value is the effective tiling the rules run under: "continuous", "quantised", or None
for the kind (activity_score) that is neither. The other four members are the evidence:
declared is the authored measure_tiling if any row carried one, default is what the kind
would have been read under, fractional is the first value that no quantised reading can hold
(with the column it sits in), and disagreement names two rows of one group that declared
different tilings.
format_group_key ¶
A bin group's key as it appears in a message, with an integral float rendered as an integer.
A warning's text is an API — compile_module copies its warnings into
manifest.compilation.warnings, and a catalog reindexing from a published manifest has nothing
else — so recompiling an unchanged module must not move the string. CopyNumberRow._KEY_FIELDS
keys on effective_modifier_copy_number since 0.6, which coalesces the deprecated int column
to a float, and a bare interpolation would have turned every published
…for key ('SMN1', 'SMN2', 2, None) into 2.0 with no other change in the module. Normalizing
here keeps the coalesce invisible to a reader who never writes the new column, and leaves new
text appearing only where genuinely fractional data exists.
Grouping itself is unaffected either way — 2 == 2.0 and they hash equal, so the two spellings
were always one dict key. This is a rendering rule and nothing else. Same normalization
just_dna_compiler.compiler._scalar_cell applies to a parquet value on the way back to a CSV,
for the same reason: a whole number should read as one.
Source code in schema/src/just_dna_format/binning.py
resolve_tiling ¶
The effective tiling for one bin group: declared, else inferred, else the kind's default.
One resolution with two callers (validate_bins for the shared-endpoint and gap rules,
measurement_shape_warnings for the RM55 finding), because a second copy would answer the same
question against a different reading of the same rows.
Order matters. Agreement is checked first — with two rows declaring different tilings there
is no declared value to read — and it is returned rather than raised so the warning path cannot
be crashed by a table the error path is about to refuse anyway. None beside a declared value is
absence, not disagreement: the house algebra is three-valued and None is never a value.
The inference runs one way only. Fractional-ness implies continuous — a value between two
grid points is incompatible with quantised semantics — while integer-ness implies nothing, since
[0,1] [2,3] is exactly what a continuous measure looks like when its author has only ever seen
whole-number data. That asymmetry is why one direction can be automatic and the other has to be
declared. An explicit quantised beside a fractional value stands: the author said something
the data contradicts, and picking a winner silently is what a three-valued algebra exists to
avoid, so the caller warns instead.
And it fires only against a quantised default, because that is the only reading a fractional
value contradicts. quantised asserts a step, which is what lets the gap rule tolerate a hole
of one, and a fractional bound falsifies the assertion. The other two defaults assert nothing a
fraction can falsify: continuous is already what the evidence would say, and activity_score's
None makes no claim about the grid at all — it reports no interior hole precisely because
the score is consumer-summed onto a grid this schema does not know the step of. Reading a
fractional activity score as continuous would therefore invent findings rather than reveal them:
reference_examples/cyp2d6_structural states bins at 0.25/0.5/1.25/2.25, and under continuous
rules those produce three "coverage gap" warnings for intervals no activity score can land in.
The rule is the data contradicts the reading, not the data is fractional.
Source code in schema/src/just_dna_format/binning.py
measurement_shape_warnings ¶
What a table of quantised bins cannot express about its own source measurement (RM55/RM56).
Both findings are stated against the kind, once per table, never per row or per group: the finding is about how the axis is divided, so a per-row line would be the same sentence repeated as many times as the author wrote bins. See the module docstring for the spec quotations.
- RM55 — VCF 4.4 makes both
CNandRUCnon-integral, so a fractional measurement between two adjacent quantised bins matches neither, and the coverage-gap check does not see the hole because under quantised tiling it only reports one wider than a step. Note the exposure is at the boundaries and on a sharp tiling, not everywhere: a fraction inside a wide range bin is answered fine, which is why this is stated against the kind rather than derived per bound. Fires only where it is still true — a kind VCF 4.4 types as fractional and at least one group of that kind whose effective tiling isquantised. A table that declaresmeasure_tiling: continuous, or that carries a fractional bound and is read as continuous because of it, answers its own boundaries and is silent here. A group declaringquantisedbeside a fractional value keeps the declaration, so it keeps this finding too — which is the right reading: the author has said the axis is a grid and written a value that is not on it. - RM56 — the same two fields travel with a confidence interval whose upper bound may be
unbounded, so a measurement is an interval and can cross a threshold. Fires only where there is a
threshold to cross: two or more resolved bins in one group. With a single bin there is nothing to
span, and saying so anyway would be a finding about a table that does not have the problem. The
count is of bins, and the message says only that — an earlier draft called them adjacent,
which is a property this function never computes and which
[0,0]beside[50,60]falsifies. Adjacency is also not what matters here: an interval crosses two bins whether or not there is a hole between them, and where there is one the hole isvalidate_bins' finding, not this one. Unaffected by tiling: an interval spans bins however the axis is divided.
Warnings in both modes, and deliberately never a strict error — the not_covered /
VRS-coverage class. RM56 needs a policy vocabulary that has not been designed, so no authored edit
clears it at all. RM55 now does have an authored answer where the measure really is continuous,
and still does not escalate, for a different reason: a genuinely quantised catalog count is
correct, the default is quantised precisely so no published table is silently re-read, and
refusing would refuse a correct module over a property of its source's type. strict means
reproducible artifact, an unrelated axis (P5), and both tables reproduce exactly.
Source code in schema/src/just_dna_format/binning.py
901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 | |
deprecation_warnings ¶
Columns this release still reads and behaves exactly as before on, and an author should stop writing (Principle 3's two-step retirement: deprecate in a minor, remove at the next major).
Once per table, never once per row. A copy-number table states one dosage per SMN2 copy number and the sentence is the same for every one of them; the repeated-warning rule and the embedded-count rule both apply, so the line carries no count and is emitted at most once for the rows handed over.
The warning is actionable, which is the 0.6 cadence amendment's own condition for deprecating
inside a minor: modifier_copy_number exists in this same release, holds everything the integer
column held, and an author can move the value the day they read this.
Source code in schema/src/just_dna_format/binning.py
validate_bins ¶
Table-level coherence check for a set of binning rows of one kind (consumer round-2 C1).
Rows are grouped by their explicit key columns (_KEY_FIELDS) plus trait_efo_id. Within a
group of resolved rows — a consumer measurement selects at most one — inclusive ranges
[measure_min, measure_max] (a null bound = -inf/+inf) must not overlap; an overlap would
select two phenotypes for one measurement and raises ValueError. Overlap across different
trait_efo_id is allowed (pleiotropy — the same measurement legitimately binning to two traits).
unresolved sentinel rows carry no range and are ignored.
Both rules read the group's effective tiling (resolve_tiling), which since 0.6 is the
declared measure_tiling, else the reading a fractional value forces, else the kind's default
(RM55). Under continuous two bins are touching, which is the only way to tile a dense domain
at all, so the error is lo < prev_hi and the shared value belongs to the higher bin, and any
positive hole is reported. Under quantised adjacent bins sharing an endpoint really do both
claim that grid point, so lo <= prev_hi is the error, and only a hole wider than one step is a
hole. Under neither (activity_score, which is consumer-summed onto a coarse grid) a shared
endpoint is an overlap and interior holes are not reported at all. Two bins sharing a lower
bound refuse in every case: the boundary rule picks the greatest measure_min at or below the
measurement, and there is nothing to pick between equals.
Three findings come out of the tiling itself. Two rows of one group declaring different
tilings raise — there is no group tiling to run the rules under. A tiling inferred from a
fractional value emits an informational line naming the group, the value and the rules it
applied, because an inference a reader cannot see is the thing this repo distrusts about
inference. And an explicit quantised beside a fractional value warns that the data contradicts
the declaration, while the declaration stands — neither side silently overrides the other.
Returns a list of warnings (the two tiling notices above, plus interior coverage gaps: a value between two authored bins that matches no row). Edge coverage below the lowest bin (the "author the reference bin" contract, C1) is a consumer-contract matter, not auto-detected here — it would false-positive without a known domain floor. Callers decide what to do with the warnings (log, fail, ignore).
Source code in schema/src/just_dna_format/binning.py
1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 | |