Skip to content

just_dna_format.binning

just_dna_format.binning

The measure → phenotype binning primitive (0.4 — see docs/CHANGELOG.md).

One declarative shape shared by every quantity-carrying locus: a per-locus table that maps a measured quantity (activity score, copy number, repeat count, heteroplasmy fraction, PRS percentile) to a phenotype by range. The tables differ only in which quantity is measured and in their explicit key columns (multicolumn keying — never a packed tuple; the keying stance is a coding standard, see CLAUDE.md). Aligning the column vocabulary gives a consumer one "bin-a-measure" code path.

Data-agnostic (design north star — see CLAUDE.md). These rows are pure annotation: a lookup table declaring range→phenotype. The module contains no measurement — the measured quantity is supplied by the consumer at query time; the table never sees a sample. The bins themselves are a generalization over a practical subset of real loci/ranges, not an all-encompassing model, so a data item that doesn't fit is a schema gap to widen additively.

Ranges are inclusive [measure_min, measure_max]: min == max is a sharp value (e.g. exactly 0 copies), min < max is a range (HTT 36–39 CAG), and measure_max = None is open-ended (≥40 CAG, 3+ copies). There is no copy_number column — a sharp copy number is measure_min == measure_max.

On a continuous measure, two adjacent bins may share an endpoint, and the higher bin owns it. The lookup rule, which a consumer implements once: select the row with the greatest measure_min ≤ x (within the group). Written out for allele_fraction bins 0.0–0.1, 0.1–0.3, 0.3–1.0, a heteroplasmy of exactly 0.1 selects the MIDD bin and 1.0 selects the top one.

This is a rule about tiling, not a second meaning for measure_max, and it exists because the alternative was unsatisfiable (RM35). Inclusive-at-both-ends, overlap-is-an-error and any-positive-hole-is-a-warning cannot all hold on a dense domain: two adjacent continuous bins either share an endpoint or leave a gap, so every allele_fraction/prs_percentile table carried a finding forever — a check that could not be satisfied rather than one that was failing. No epsilon escapes it ([0, 0.0999999] + [0.1, 1.0] still warns). Half-open [min, max) for continuous kinds was the other candidate and lost on authorship: it makes one column mean two things depending on measure_kind, the number written in the cell is then not in the bin, and a bounded domain's top value (AF = 1.0 is homoplasmy, and real) becomes unreachable unless the last bin is authored open. Here measure_max means the same thing on every kind and the top bin stays closed.

Discrete kinds tile the other way, and that is still their default. repeat_count/copy_number tile exactly under inclusive bounds — HTT [6,35], [36,39], [40,∞) is gapless if the domain is integral — so for them a shared endpoint is a real overlap and an error unless the table says otherwise.

…except that the domain is not integral, and the spec says so (RM55). VCF 4.4 §7.2 "Redefined INFO and FORMAT CN to support non-integer copy numbers" and its worked examples are fractional throughout (CN=3,0.9666, CN=1.25); §5.6 leaves the granularity of a copy number deliberately undefined and allows a segment mean "at a highly granular megabase level of resolution". §3 types RUC, the repeat count VCF 4.4 standardises, as a Float. So the premise the paragraph above rests on was withdrawn for both kinds, and the consequence was worse than the allele_fraction case RM35 fixed: a tiling [0,0] [1,1] [2,2] [3,∞) was accepted with no warning at all, a measured 2.4 matched no bin, and the module compiled green under --strict. A hole of exactly one is invisible to the gap check by construction — and, the half nobody had written down, the schema also refused the tiling that would fix it, since a shared endpoint on these kinds was an overlap error.

How 0.6 closes it: tiling becomes its own axis. measure_tiling ({quantised, continuous}, optional, VALID_MEASURE_TILINGS) says how the axis is divided, which is a different question from measure_kind's what is measured — folding the two would be the overloaded-field anti-pattern (P5), and a product rather than a sum. The shared-endpoint and gap rules read the group's effective tiling (resolve_tiling): the declared value, else continuous where the kind would default to quantised and the group carries a value no grid of whole numbers can hold, else the kind's default (DEFAULT_MEASURE_TILING). Absence means the kind's default, never a value, which is what makes the column additive — every module published before 0.6 keeps its exact meaning, and an author meets the column only when departing from it. The inference announces itself, runs one way only — fractional-ness contradicts a stated grid, integer-ness contradicts nothing, since [0,1] [2,3] is what a continuous measure looks like when its author has only seen whole-number data — and fires only against a quantised default, because that is the only reading a fraction falsifies. See resolve_tiling for why activity_score, which is fractional by nature, must not be moved by it.

Three repairs were refused. Moving the kinds into _DENSE_KINDS is one line and silently re-reads every published table ([2,2] beside [3,3] is a legal quantised tiling, and both rules change meaning under dense semantics) with no notice and no way to say otherwise. A sixth measure_kind is the wrong axis, above. And deriving the tiling from the rows with no column at all fails on the half that has to be right: absence of a fractional value implies nothing, so it would read the ambiguous table one way, silently, and leave the curator no way to correct it.

The one genuine int here is CopyNumberRow.modifier_cn, a modifier's dosage rather than a bound, and it gets the parallel-float treatment the roadmap entry had proposed for the bounds: modifier_copy_number beside it, read through effective_modifier_copy_number, both-set refused, modifier_cn deprecated in 0.6 and removed at 1.0. So 1.0 inherits a removal, not a retype.

A measurement can also span several bins, and there is no state for that (RM56). RUC travels with CIRUC and CN with CICN, and the spec is explicit that the upper bound may be missing and then means unbounded (§3: "a reasonable limit of total length of the repeat could not be determined"). §5.7's canonical CAG form is RUS=CAG;RUC=65;CIRUC=-15,., and says why: "Many of these techniques result in imprecise variant calls." Imprecision is the normal case. The consumer contract below has exactly three states — a bin matched, no bin matched, or the measurement absent — and none of them is the measurement spans bins. reference_examples/htt_repeat_expansion has thresholds at 26/27, 35/36 and 39/40, so a real RUC=38, CIRUC=-5,5 spans [33,43] and crosses all three: benign, uncertain and fully penetrant, with no honest answer among them.

The policy vocabulary that settles it — withhold / take the worst bin / take the point estimate — lands with the rest of the repeat work, when a real caller VCF is in hand, and its grain (per table or per row) is deliberately undecided. Until then the stated placeholder behaviour is the house default: a consumer that reads an interval spanning two or more bins withholds. It does not pick among them, and it does not fall back to the unresolved sentinel either — that row means no measurement was available, and here one was. Note what is not on the table: widening the measurement itself into an interval puts a measurement in the module, which the data-agnostic north star forbids outright. What belongs here is the rule for an interval that spans bins, which is annotation.

unresolved (T1) is mandatory. A table can state the outcome for measurement absent / not callable, and the consumer contract is that a missing measurement selects the unresolved row, never the lowest/reference bin (no activity score ⇒ not "Normal Metabolizer"; no CN ⇒ not "2 copies"; no heteroplasmy read ⇒ not "homoplasmic reference"). An unresolved row carries no bounds. A measurement that is present but matches no bin is a distinct third state ("no matching bin", not unresolved); validate_bins below rejects overlaps and flags coverage gaps so a table stays coherent (consumer round-2 C1).

pmid grounds the BOUNDARY, and the line is: the bin row cites, the citation table describes (RM47, 0.6). A threshold is the most interpretive claim this format carries — where 36 rather than 35 CAG becomes "reduced penetrance" is a clinical judgement drawn from a specific paper — and until 0.6 nothing could point at one: studies.csv identifies its subject by rsid or chrom, while a repeat_alleles.csv row is keyed (gene, repeat_unit). One optional column on this base reaches all four kinds and answers the question actually asked, which is why 36 rather than why this table. It carries a pointer only. Everything about the paper — population, p_value_num, effect_size, provenance_quote — stays in studies.csv, whose subject requirement was relaxed in the same release so a citation row may exist without naming a variant. Copying that column set here one column at a time would restate the bin inside its own evidence, which is exactly what killed the alternative designs (a bin_evidence.csv join table keys on the thresholds, and they are floats).

source_field (round-2 3a) is a declarative pointer, not code. It optionally names the VCF field the consumer extracts the measure from (FORMAT/REPCN, FORMAT/AF, INFO/CN|FORMAT/DS) — pure indirection/addressing, deliberately constrained to a field-name key (optionally namespace- qualified, optionally |-alternated) so it can never become an expression. That keeps it inside Principle 1 (declarative, non-Turing): a name that says where the measurement lives, never a transform that computes one. The module still holds no measurement.

A VCF field is identified by namespace and described by cardinality, and source_field used to carry neither (RM53/RM54, 0.6). INFO and FORMAT are two reserved-key tables that collide on DP, AD, ADF, ADR, MQ, AF and — since 4.4 — CN, so AF alone names the cohort frequency of an ALT or this sample's fraction of it, and both are floats in [0, 1] that bin without complaint. Number then decides how many values come back: a pointer at a Number=R field returns one value per allele, reference first, of which none is the answer. So the namespace goes in the pointer (FORMAT/AF, bare still legal and still meaning unqualified) and the element goes in source_element — a closed set of named rules rather than an index, because AD[1] is the first line of an expression grammar and Principle 1 refuses it.

"Element" is one of the values the field carries for a record, which is wider than a Number slot on purpose: ExpansionHunter reports both repeat alleles in a single REPCN cell as 17/42, and a rule that only spoke about Number would have nothing to say about the case it was built for (reference_examples/htt_repeat_expansion, where the clinical rule is the larger of the two). How a caller encodes multiplicity is the caller's business and this tier holds no opinion on it (Principle 2); which value the annotation means is the module's, and that is all this column states.

MeasureBinRow

Bases: AuthoredModel

Base row of a binning table: a measured quantity range → the same orthogonal axes a VariantRow carries. Subclasses add the explicit key columns for their quantity.

Inherits AuthoredModel (and passes it to its subclasses): extra="forbid" + the reserved-namespace guard, and the shared direction/clin_sig/trait_efo_id validators.

ActivityPhenotypeRow

Bases: MeasureBinRow

PGx metabolizer phenotype by activity score, per gene (CYP2D6 PM/IM/NM/UM). The score is a consumer call (Σ activity×copies over the diplotype); this table only bins it.

CopyNumberRow

Bases: MeasureBinRow

Whole-gene dosage phenotype by copy number (SMN1 SMA). Sharp dosages are measure_min == measure_max (0 copies = [0, 0]); 3+ is measure_min=3, measure_max=None.

Optional modifier_gene plus a modifier dosage express a second locus read in context (SMN1 phenotype depends on SMN2 copy number) — explicit named columns (multicolumn keying), never a tuple. The gene and the dosage are set together or both left null.

The dosage has two spellings and one meaning (RM55, 0.6). modifier_copy_number is the float column; modifier_cn is the original int one, deprecated in 0.6 and removed at 1.0. VCF 4.4 §7.2 redefined CN to support non-integer copy numbers, so an int cannot hold a real modifier dosage — and retyping the column is major-only, which is what the companion exists to avoid. Everything reads effective_modifier_copy_number, including _KEY_FIELDS: a coalesced value is one spelling by the time grouping and dedup see it, so the key never holds two spellings of one number. Setting both is an error rather than a precedence rule.

effective_modifier_copy_number property

effective_modifier_copy_number: float | None

The modifier dosage, from whichever column carries it — the only reader of the pair.

is not None rather than or, because 0 is a legal dosage: SMN2 = 0 copies is a real row, and a truthiness fallback would silently read it as "unset" and then fall through to the deprecated column. (The boolean effective_pathogenic/effective_benign aliases are the precedent here, not the string effective_direction, which may use or because an empty string is not a legal value there.)

A fixed point by construction — coalescing an already-coalesced value returns it unchanged (P7 idempotency).

RepeatAlleleRow

Bases: MeasureBinRow

VNTR/STR phenotype by repeat count, keyed on (gene, repeat_unit) — the motif is part of the identity (T3): a count is only comparable within its motif definition. The count is a consumer call (ExpansionHunter / adVNTR / a span genotyper) that MUST state the motif it counted.

HeteroplasmyRow

Bases: MeasureBinRow

mtDNA phenotype by heteroplasmy allele fraction (0–1), keyed on (gene, reference_sequence, tissue). The reference sequence is part of the key (A3): rCRS/ NC_012920 vs legacy NC_001807 disagree and genome_build does not disambiguate. Bounds are constrained to [0, 1].

tissue/assay_context are optional but load-bearing (round-2 Q6): heteroplasmy bins are tissue-conditional — a blood-derived fraction systematically under-represents the affected-tissue burden, and the penetrance threshold itself shifts by tissue, so the same fraction bins to different phenotypes across tissues. A heteroplasmy table with no tissue context is quietly unsafe; state the tissue the bins assume.

The variant identity is part of the key too (0.5.1), and it was missing. A mitochondrial gene carries several pathogenic variants with genuinely different thresholds — MT-TL1 has m.3243A>G and m.3271T>C, both causing MELAS; MT-ATP6 has m.8993T>G and m.9176T>C. Keyed on the gene alone, their bins landed in one group and validate_bins rejected the module outright with "overlapping bins", which is an error, not a warning, so the module could not compile at all. There was no honest way out: trait_efo_id is in the group key and would have separated them, but only by giving one disease two ontology ids. The documented example never showed two variants in a gene, so the limitation was invisible rather than decided.

The columns mirror PharmVariantRow exactly — rsid, else chrom+start(+ref/alts) — and are optional, so an existing single-variant table groups as it always did (P3/P8). They enter the key through the derived variant_key property rather than one-by-one, so all the identity shapes collapse to the format's own notion of which variant a row is about.

TilingResolution

Bases: NamedTuple

How one bin group's tiling was decided, and from what — so a caller can say which it was.

value is the effective tiling the rules run under: "continuous", "quantised", or None for the kind (activity_score) that is neither. The other four members are the evidence: declared is the authored measure_tiling if any row carried one, default is what the kind would have been read under, fractional is the first value that no quantised reading can hold (with the column it sits in), and disagreement names two rows of one group that declared different tilings.

inferred property

inferred: bool

A fractional value moved the group off its kind's default — announce it.

contradicted property

contradicted: bool

An explicit quantised sits beside a value quantised semantics cannot hold.

format_group_key

format_group_key(group_key: tuple) -> str

A bin group's key as it appears in a message, with an integral float rendered as an integer.

A warning's text is an API — compile_module copies its warnings into manifest.compilation.warnings, and a catalog reindexing from a published manifest has nothing else — so recompiling an unchanged module must not move the string. CopyNumberRow._KEY_FIELDS keys on effective_modifier_copy_number since 0.6, which coalesces the deprecated int column to a float, and a bare interpolation would have turned every published …for key ('SMN1', 'SMN2', 2, None) into 2.0 with no other change in the module. Normalizing here keeps the coalesce invisible to a reader who never writes the new column, and leaves new text appearing only where genuinely fractional data exists.

Grouping itself is unaffected either way — 2 == 2.0 and they hash equal, so the two spellings were always one dict key. This is a rendering rule and nothing else. Same normalization just_dna_compiler.compiler._scalar_cell applies to a parquet value on the way back to a CSV, for the same reason: a whole number should read as one.

Source code in schema/src/just_dna_format/binning.py
def format_group_key(group_key: tuple) -> str:
    """A bin group's key as it appears in a message, with an integral float rendered as an integer.

    **A warning's text is an API** — `compile_module` copies its warnings into
    `manifest.compilation.warnings`, and a catalog reindexing from a published manifest has nothing
    else — so recompiling an *unchanged* module must not move the string. `CopyNumberRow._KEY_FIELDS`
    keys on `effective_modifier_copy_number` since 0.6, which coalesces the deprecated `int` column
    to a `float`, and a bare interpolation would have turned every published
    `…for key ('SMN1', 'SMN2', 2, None)` into `2.0` with no other change in the module. Normalizing
    here keeps the coalesce invisible to a reader who never writes the new column, and leaves new
    text appearing only where genuinely fractional data exists.

    Grouping itself is unaffected either way — `2 == 2.0` and they hash equal, so the two spellings
    were always one dict key. This is a rendering rule and nothing else. Same normalization
    `just_dna_compiler.compiler._scalar_cell` applies to a parquet value on the way back to a CSV,
    for the same reason: a whole number should read as one.
    """
    members = tuple(
        int(v) if isinstance(v, float) and math.isfinite(v) and v.is_integer() else v for v in group_key
    )
    return repr(members)

resolve_tiling

resolve_tiling(
    grp: Sequence[MeasureBinRow],
) -> TilingResolution

The effective tiling for one bin group: declared, else inferred, else the kind's default.

One resolution with two callers (validate_bins for the shared-endpoint and gap rules, measurement_shape_warnings for the RM55 finding), because a second copy would answer the same question against a different reading of the same rows.

Order matters. Agreement is checked first — with two rows declaring different tilings there is no declared value to read — and it is returned rather than raised so the warning path cannot be crashed by a table the error path is about to refuse anyway. None beside a declared value is absence, not disagreement: the house algebra is three-valued and None is never a value.

The inference runs one way only. Fractional-ness implies continuous — a value between two grid points is incompatible with quantised semantics — while integer-ness implies nothing, since [0,1] [2,3] is exactly what a continuous measure looks like when its author has only ever seen whole-number data. That asymmetry is why one direction can be automatic and the other has to be declared. An explicit quantised beside a fractional value stands: the author said something the data contradicts, and picking a winner silently is what a three-valued algebra exists to avoid, so the caller warns instead.

And it fires only against a quantised default, because that is the only reading a fractional value contradicts. quantised asserts a step, which is what lets the gap rule tolerate a hole of one, and a fractional bound falsifies the assertion. The other two defaults assert nothing a fraction can falsify: continuous is already what the evidence would say, and activity_score's None makes no claim about the grid at all — it reports no interior hole precisely because the score is consumer-summed onto a grid this schema does not know the step of. Reading a fractional activity score as continuous would therefore invent findings rather than reveal them: reference_examples/cyp2d6_structural states bins at 0.25/0.5/1.25/2.25, and under continuous rules those produce three "coverage gap" warnings for intervals no activity score can land in. The rule is the data contradicts the reading, not the data is fractional.

Source code in schema/src/just_dna_format/binning.py
def resolve_tiling(grp: Sequence[MeasureBinRow]) -> TilingResolution:
    """The effective tiling for one bin group: declared, else inferred, else the kind's default.

    One resolution with two callers (`validate_bins` for the shared-endpoint and gap rules,
    `measurement_shape_warnings` for the RM55 finding), because a second copy would answer the same
    question against a different reading of the same rows.

    Order matters. **Agreement is checked first** — with two rows declaring different tilings there
    is no declared value to read — and it is *returned* rather than raised so the warning path cannot
    be crashed by a table the error path is about to refuse anyway. `None` beside a declared value is
    absence, not disagreement: the house algebra is three-valued and `None` is never a value.

    **The inference runs one way only.** Fractional-ness *implies* continuous — a value between two
    grid points is incompatible with quantised semantics — while integer-ness implies nothing, since
    `[0,1] [2,3]` is exactly what a continuous measure looks like when its author has only ever seen
    whole-number data. That asymmetry is why one direction can be automatic and the other has to be
    declared. An explicit `quantised` beside a fractional value **stands**: the author said something
    the data contradicts, and picking a winner silently is what a three-valued algebra exists to
    avoid, so the caller warns instead.

    **And it fires only against a `quantised` default, because that is the only reading a fractional
    value contradicts.** `quantised` asserts a step, which is what lets the gap rule tolerate a hole
    of one, and a fractional bound falsifies the assertion. The other two defaults assert nothing a
    fraction can falsify: `continuous` is already what the evidence would say, and `activity_score`'s
    `None` makes **no** claim about the grid at all — it reports no interior hole precisely because
    the score is consumer-summed onto a grid this schema does not know the step of. Reading a
    fractional activity score as continuous would therefore invent findings rather than reveal them:
    `reference_examples/cyp2d6_structural` states bins at 0.25/0.5/1.25/2.25, and under continuous
    rules those produce three "coverage gap" warnings for intervals no activity score can land in.
    The rule is *the data contradicts the reading*, not *the data is fractional*.
    """
    declared_rows = [r.measure_tiling for r in grp if r.measure_tiling is not None]
    disagreement: tuple[str, str] | None = None
    declared: str | None = None
    distinct = sorted(set(declared_rows))
    if len(distinct) > 1:
        disagreement = (distinct[0], distinct[1])
    elif distinct:
        declared = distinct[0]

    fractional: tuple[str, float] | None = None
    for row in grp:
        found = _fractional_values(row)
        if found:
            fractional = found[0]
            break

    default = DEFAULT_MEASURE_TILING.get(grp[0].measure_kind)
    if declared is not None:
        value = declared
    elif fractional is not None and default == "quantised":
        value = "continuous"
    else:
        value = default
    return TilingResolution(value, declared, default, fractional, disagreement)

measurement_shape_warnings

measurement_shape_warnings(
    rows: Sequence[MeasureBinRow],
) -> list[str]

What a table of quantised bins cannot express about its own source measurement (RM55/RM56).

Both findings are stated against the kind, once per table, never per row or per group: the finding is about how the axis is divided, so a per-row line would be the same sentence repeated as many times as the author wrote bins. See the module docstring for the spec quotations.

  • RM55 — VCF 4.4 makes both CN and RUC non-integral, so a fractional measurement between two adjacent quantised bins matches neither, and the coverage-gap check does not see the hole because under quantised tiling it only reports one wider than a step. Note the exposure is at the boundaries and on a sharp tiling, not everywhere: a fraction inside a wide range bin is answered fine, which is why this is stated against the kind rather than derived per bound. Fires only where it is still true — a kind VCF 4.4 types as fractional and at least one group of that kind whose effective tiling is quantised. A table that declares measure_tiling: continuous, or that carries a fractional bound and is read as continuous because of it, answers its own boundaries and is silent here. A group declaring quantised beside a fractional value keeps the declaration, so it keeps this finding too — which is the right reading: the author has said the axis is a grid and written a value that is not on it.
  • RM56 — the same two fields travel with a confidence interval whose upper bound may be unbounded, so a measurement is an interval and can cross a threshold. Fires only where there is a threshold to cross: two or more resolved bins in one group. With a single bin there is nothing to span, and saying so anyway would be a finding about a table that does not have the problem. The count is of bins, and the message says only that — an earlier draft called them adjacent, which is a property this function never computes and which [0,0] beside [50,60] falsifies. Adjacency is also not what matters here: an interval crosses two bins whether or not there is a hole between them, and where there is one the hole is validate_bins' finding, not this one. Unaffected by tiling: an interval spans bins however the axis is divided.

Warnings in both modes, and deliberately never a strict error — the not_covered / VRS-coverage class. RM56 needs a policy vocabulary that has not been designed, so no authored edit clears it at all. RM55 now does have an authored answer where the measure really is continuous, and still does not escalate, for a different reason: a genuinely quantised catalog count is correct, the default is quantised precisely so no published table is silently re-read, and refusing would refuse a correct module over a property of its source's type. strict means reproducible artifact, an unrelated axis (P5), and both tables reproduce exactly.

Source code in schema/src/just_dna_format/binning.py
def measurement_shape_warnings(rows: Sequence[MeasureBinRow]) -> list[str]:
    """What a table of quantised bins cannot express about its own source measurement (RM55/RM56).

    Both findings are stated against the *kind*, once per table, never per row or per group: the finding
    is about how the axis is divided, so a per-row line would be the same sentence repeated as many
    times as the author wrote bins. See the module docstring for the spec quotations.

    * **RM55** — VCF 4.4 makes both `CN` and `RUC` non-integral, so a fractional measurement *between
      two adjacent quantised bins* matches neither, and the coverage-gap check does not see the hole
      because under quantised tiling it only reports one wider than a step. Note the exposure is at
      the boundaries and on a sharp tiling, not everywhere: a fraction inside a wide range bin is
      answered fine, which is why this is stated against the kind rather than derived per bound.
      **Fires only where it is still true** — a kind VCF 4.4 types as fractional *and* at least one
      group of that kind whose effective tiling is `quantised`. A table that declares
      `measure_tiling: continuous`, or that carries a fractional bound and is read as continuous
      because of it, answers its own boundaries and is silent here. A group declaring `quantised`
      beside a fractional value keeps the declaration, so it keeps this finding too — which is the
      right reading: the author has said the axis is a grid and written a value that is not on it.
    * **RM56** — the same two fields travel with a confidence interval whose upper bound may be
      unbounded, so a measurement is an interval and can cross a threshold. Fires only where there is a
      threshold to cross: two or more resolved bins in one group. With a single bin there is nothing to
      span, and saying so anyway would be a finding about a table that does not have the problem. The
      count is of **bins**, and the message says only that — an earlier draft called them *adjacent*,
      which is a property this function never computes and which `[0,0]` beside `[50,60]` falsifies.
      Adjacency is also not what matters here: an interval crosses two bins whether or not there is a
      hole between them, and where there is one the hole is `validate_bins`' finding, not this one.
      Unaffected by tiling: an interval spans bins however the axis is divided.

    **Warnings in both modes, and deliberately never a `strict` error** — the `not_covered` /
    VRS-coverage class. RM56 needs a policy vocabulary that has not been designed, so no authored edit
    clears it at all. RM55 now *does* have an authored answer where the measure really is continuous,
    and still does not escalate, for a different reason: a genuinely quantised catalog count is
    **correct**, the default is quantised precisely so no published table is silently re-read, and
    refusing would refuse a correct module over a property of its source's type. `strict` means
    *reproducible artifact*, an unrelated axis (P5), and both tables reproduce exactly.
    """
    warnings: list[str] = []
    groups = _bin_groups(rows)
    kinds = {r.measure_kind for grp in groups.values() for r in grp} & set(_VCF_MEASURE_FIELDS)
    for kind in sorted(kinds):
        value_field, ci_field, spec_note = _VCF_MEASURE_FIELDS[kind]
        of_kind = {key: grp for key, grp in groups.items() if any(r.measure_kind == kind for r in grp)}
        if any(resolve_tiling(grp).value == "quantised" for grp in of_kind.values()):
            warnings.append(
                CodedWarning(
                    "measure_field_fractional",
                    f"{kind} bins here are tiled as whole numbers, but the field a consumer reads the "
                    f"measurement from ({value_field}) {FRACTIONAL_MEASURE_PHRASE}: {spec_note}. A "
                    f"fractional measurement falling between two adjacent bins matches neither — "
                    f"`[0,0] [1,1] [2,2] [3,∞)` is a legal quantised tiling and answers nothing at all "
                    f"for a 2.4 — and the coverage-gap check does not report the hole, because under "
                    f"quantised tiling it only flags one wider than a step. If this measure really is "
                    f"continuous, say so: `measure_tiling: continuous` on these rows makes adjacent bins "
                    f"able to share an endpoint (the higher one owns it) and makes any positive hole "
                    f"reportable, and the bounds are already floats so nothing else has to change. A "
                    f"group carrying a fractional bound is read as continuous without being asked. Left "
                    f"as a grid, expect an answer only from a caller that rounds, and none from a "
                    f"segment mean.",
                )
            )
        widest = max(len(grp) for grp in of_kind.values())
        if widest >= 2:
            warnings.append(
                CodedWarning(
                    "measurement_spans_bins",
                    f"{kind} bins: {SPANNING_MEASUREMENT_PHRASE}, and nothing in this format says what to "
                    f"do with one (RM56). A {value_field} call travels with {ci_field}, whose missing upper "
                    f"bound means *unbounded*, so the measurement is an interval — and the widest group "
                    f"here states {widest} bins for it to cross. The consumer contract has three "
                    f"states (a bin matched, no bin matched, the measurement absent) and none of them is "
                    f"this one. Not implemented, and stated rather than left silent: until the policy "
                    f"vocabulary lands, a conforming consumer **withholds** — it does not pick among the "
                    f"bins the interval touches, and it does not fall back to the `unresolved` row, which "
                    f"means no measurement was available and is a different claim.",
                )
            )
    return warnings

deprecation_warnings

deprecation_warnings(
    rows: Sequence[MeasureBinRow],
) -> list[str]

Columns this release still reads and behaves exactly as before on, and an author should stop writing (Principle 3's two-step retirement: deprecate in a minor, remove at the next major).

Once per table, never once per row. A copy-number table states one dosage per SMN2 copy number and the sentence is the same for every one of them; the repeated-warning rule and the embedded-count rule both apply, so the line carries no count and is emitted at most once for the rows handed over.

The warning is actionable, which is the 0.6 cadence amendment's own condition for deprecating inside a minor: modifier_copy_number exists in this same release, holds everything the integer column held, and an author can move the value the day they read this.

Source code in schema/src/just_dna_format/binning.py
def deprecation_warnings(rows: Sequence[MeasureBinRow]) -> list[str]:
    """Columns this release still reads and behaves exactly as before on, and an author should stop
    writing (Principle 3's two-step retirement: deprecate in a minor, remove at the next major).

    **Once per table, never once per row.** A copy-number table states one dosage per SMN2 copy
    number and the sentence is the same for every one of them; the repeated-warning rule and the
    embedded-count rule both apply, so the line carries no count and is emitted at most once for the
    rows handed over.

    The warning is **actionable**, which is the 0.6 cadence amendment's own condition for deprecating
    inside a minor: `modifier_copy_number` exists in this same release, holds everything the integer
    column held, and an author can move the value the day they read this.
    """
    if any(isinstance(r, CopyNumberRow) and r.modifier_cn is not None for r in rows):
        return [
            CodedWarning(
                "deprecated_bin_modifier",
                f"{DEPRECATED_MODIFIER_PHRASE} and is removed at 1.0: write the dosage in "
                f"`modifier_copy_number` instead, which is a float and can hold the non-integer copy "
                f"numbers VCF 4.4 §7.2 allows. It still reads and behaves exactly as before until then, "
                f"a whole number stays a whole number, and setting both columns is an error.",
            )
        ]
    return []

validate_bins

validate_bins(rows: Sequence[MeasureBinRow]) -> list[str]

Table-level coherence check for a set of binning rows of one kind (consumer round-2 C1).

Rows are grouped by their explicit key columns (_KEY_FIELDS) plus trait_efo_id. Within a group of resolved rows — a consumer measurement selects at most one — inclusive ranges [measure_min, measure_max] (a null bound = -inf/+inf) must not overlap; an overlap would select two phenotypes for one measurement and raises ValueError. Overlap across different trait_efo_id is allowed (pleiotropy — the same measurement legitimately binning to two traits). unresolved sentinel rows carry no range and are ignored.

Both rules read the group's effective tiling (resolve_tiling), which since 0.6 is the declared measure_tiling, else the reading a fractional value forces, else the kind's default (RM55). Under continuous two bins are touching, which is the only way to tile a dense domain at all, so the error is lo < prev_hi and the shared value belongs to the higher bin, and any positive hole is reported. Under quantised adjacent bins sharing an endpoint really do both claim that grid point, so lo <= prev_hi is the error, and only a hole wider than one step is a hole. Under neither (activity_score, which is consumer-summed onto a coarse grid) a shared endpoint is an overlap and interior holes are not reported at all. Two bins sharing a lower bound refuse in every case: the boundary rule picks the greatest measure_min at or below the measurement, and there is nothing to pick between equals.

Three findings come out of the tiling itself. Two rows of one group declaring different tilings raise — there is no group tiling to run the rules under. A tiling inferred from a fractional value emits an informational line naming the group, the value and the rules it applied, because an inference a reader cannot see is the thing this repo distrusts about inference. And an explicit quantised beside a fractional value warns that the data contradicts the declaration, while the declaration stands — neither side silently overrides the other.

Returns a list of warnings (the two tiling notices above, plus interior coverage gaps: a value between two authored bins that matches no row). Edge coverage below the lowest bin (the "author the reference bin" contract, C1) is a consumer-contract matter, not auto-detected here — it would false-positive without a known domain floor. Callers decide what to do with the warnings (log, fail, ignore).

Source code in schema/src/just_dna_format/binning.py
def validate_bins(rows: Sequence[MeasureBinRow]) -> list[str]:
    """Table-level coherence check for a set of binning rows of one kind (consumer round-2 C1).

    Rows are grouped by their explicit key columns (`_KEY_FIELDS`) plus `trait_efo_id`. Within a
    group of *resolved* rows — a consumer measurement selects at most one — inclusive ranges
    `[measure_min, measure_max]` (a null bound = -inf/+inf) **must not overlap**; an overlap would
    select two phenotypes for one measurement and raises ``ValueError``. Overlap *across* different
    `trait_efo_id` is allowed (pleiotropy — the same measurement legitimately binning to two traits).
    `unresolved` sentinel rows carry no range and are ignored.

    **Both rules read the group's effective tiling** (`resolve_tiling`), which since 0.6 is the
    declared `measure_tiling`, else the reading a fractional value forces, else the kind's default
    (RM55). Under `continuous` two bins are *touching*, which is the only way to tile a dense domain
    at all, so the error is `lo < prev_hi` and the shared value belongs to the higher bin, and any
    positive hole is reported. Under `quantised` adjacent bins sharing an endpoint really do both
    claim that grid point, so `lo <= prev_hi` is the error, and only a hole wider than one step is a
    hole. Under neither (`activity_score`, which is consumer-summed onto a coarse grid) a shared
    endpoint is an overlap and interior holes are not reported at all. Two bins sharing a *lower*
    bound refuse in every case: the boundary rule picks the greatest `measure_min` at or below the
    measurement, and there is nothing to pick between equals.

    Three findings come out of the tiling itself. Two rows of one group declaring **different**
    tilings raise — there is no group tiling to run the rules under. A tiling **inferred** from a
    fractional value emits an informational line naming the group, the value and the rules it
    applied, because an inference a reader cannot see is the thing this repo distrusts about
    inference. And an explicit `quantised` beside a fractional value warns that the data contradicts
    the declaration, while the declaration **stands** — neither side silently overrides the other.

    Returns a list of **warnings** (the two tiling notices above, plus interior coverage gaps: a
    value between two authored bins that matches no row). Edge coverage *below* the lowest bin (the
    "author the reference bin" contract, C1) is a consumer-contract matter, not auto-detected here —
    it would false-positive without a known domain floor. Callers decide what to do with the
    warnings (log, fail, ignore).
    """
    warnings: list[str] = []
    groups = _bin_groups(rows)

    for group_key, grp in groups.items():
        spans = sorted(
            (
                (
                    -math.inf if r.measure_min is None else r.measure_min,
                    math.inf if r.measure_max is None else r.measure_max,
                )
                for r in grp
            ),
            key=lambda t: (t[0], t[1]),
        )
        # Rendered once per group: every message below names the key, and an integral effective
        # modifier dosage must not start printing as `2.0` on a module nobody edited.
        shown_key = format_group_key(group_key)
        tiling = resolve_tiling(grp)
        if tiling.disagreement is not None:
            first, second = tiling.disagreement
            raise ValueError(
                f"conflicting measure_tiling for key {shown_key}: the rows of one bin group are "
                f"read under one tiling and these declare two, got {first!r} and {second!r} "
                f"(leave the column empty on the rows that do not state it — empty means the "
                f"kind's default, not a third answer)"
            )
        dense = tiling.value == "continuous"
        if tiling.inferred:
            column, value = tiling.fractional
            warnings.append(
                CodedWarning(
                    "bin_tiling_inferred",
                    f"tiling inferred for key {shown_key}: {column} is {value}, which no quantised "
                    f"reading can hold, so this group was read as continuous — adjacent bins may share "
                    f"an endpoint (the higher one owns it) and any positive hole is reported. Declare "
                    f"`measure_tiling` on these rows to state it rather than have it read off the data.",
                )
            )
        elif tiling.contradicted:
            column, value = tiling.fractional
            warnings.append(
                CodedWarning(
                    "bin_tiling_contradicted",
                    f"measure_tiling for key {shown_key} is declared 'quantised' and the data "
                    f"contradicts it: {column} is {value}, which is not a grid point. The declaration "
                    f"stands — nothing here overrides it either way — so these bins are still read "
                    f"under the quantised rules and that value sits between two of them.",
                )
            )
        for i in range(1, len(spans)):
            prev_lo, prev_hi = spans[i - 1]
            lo, hi = spans[i]
            if lo < prev_hi or (lo == prev_hi and not dense):
                raise ValueError(
                    f"overlapping bins for key {shown_key}: [{prev_lo}, {prev_hi}] and "
                    f"[{lo}, {hi}] both select a phenotype for a measurement in the overlap"
                )
            if lo == prev_lo:
                # Equal *lower* bounds, which the boundary rule cannot resolve: it selects the greatest
                # `measure_min` at or below the measurement, and these two are the same. Only reachable
                # under continuous tiling and only when the earlier bin is a single point (anything wider would
                # have tripped `lo < prev_hi` above) — e.g. `[0.1, 0.1]` beside `[0.1, 0.3]`, where a
                # measurement of exactly 0.1 has two answers and no rule to pick between them. That is
                # an ambiguous selection, so it refuses like any other overlap rather than warning.
                raise ValueError(
                    f"bins with the same lower bound for key {shown_key}: [{prev_lo}, {prev_hi}] and "
                    f"[{lo}, {hi}] both start at {lo}, so a measurement of {lo} selects two phenotypes "
                    f"and the shared-endpoint rule (the higher bin owns it) cannot separate them"
                )
            hole = lo - prev_hi
            # One step wide is not a hole on a grid, and any hole at all is one on a dense axis. The
            # third state reports neither: `activity_score` is summed onto a coarse grid the schema
            # does not know the step of, so a hole there is not a claim this tier can make.
            #
            # **`quantised`'s step is hardcoded to 1, and the schema has no way to state another.**
            # Right for `copy_number`/`repeat_count`, which is where the branch came from and the
            # only place its default applies; a limit everywhere else. Declaring `quantised` on a
            # bounded domain — `allele_fraction` in `[0, 1]` — therefore switches interior gap
            # reporting off entirely rather than tightening it, since no hole can exceed 1. Loud in
            # the realistic case (a fractional bound raises the `contradicted` warning) and silent
            # when the bounds happen to be integral. Closing it means a `measure_step` column, which
            # is a full-cost authored column nobody has asked for, so it waits for the demand that
            # would fix its shape (P5's one-way door). Documented on `measure_tiling`'s description,
            # which is what an author reads, rather than left as a surprise here.
            if tiling.value == "continuous":
                is_gap = hole > 1e-9
            elif tiling.value == "quantised":
                is_gap = hole > 1 + 1e-9
            else:
                is_gap = False
            if is_gap:
                warnings.append(
                    CodedWarning(
                        "bin_coverage_gap",
                        f"coverage gap for key {shown_key}: no bin covers ({prev_hi}, {lo})",
                    )
                )
    return warnings