just_dna_format.vrs¶
just_dna_format.vrs ¶
GA4GH VRS allele identity — derive_vrs_allele_id, the sibling of base.derive_variant_key.
A VRS Allele id (ga4gh:VA.<digest>) is a content-addressed name for one allele at one place
on one reference sequence. It is the identity RM15 was waiting for: unlike a bare chrom:start:ref,
it names its build, because the sequence is addressed by its refget accession — the digest of the
reference sequence itself — so GRCh38 and GRCh37 mint distinct, correctly non-colliding ids instead
of silently baking one build into the key.
Stdlib only. The identification algorithm is sha512t24u over a compact canonical JSON, which is
~20 lines of hashlib + base64 + json. The format tier therefore gains no dependency (it stays
pydantic + cryptography, CONSTITUTION Goal 2) and a verify-only consumer can recompute the identity
offline. This was probed, not assumed: the ids this module mints are byte-identical to the ones the
live gnomAD API returns (see schema/tests/test_vrs.py, which asserts against recorded ground truth).
Substitutions only, on purpose. derive_vrs_allele_id mints for a single-base substitution and
returns None for everything else (indel, MNV, multi-allelic cell, or a row with no coordinate). A
VRS allele id is defined over the fully justified (normalized) allele; for an SNV, normalization is
a provable no-op — there is no shared prefix or suffix to trim and nothing to shift — so the id is
computable offline and interoperable. For an indel, justification needs the reference sequence,
which this tier has no access to and will never fetch (Principle 2). Minting an unjustified indel id
would emit a ga4gh:VA.… string that looks interoperable and is not, which is worse than minting
nothing. Indel ids are minted upstream in just-dna-enricher's [dev] extras (which do have a
sequence proxy) and passed through as data, exactly as gnomAD's variantId is.
UnsupportedBuildError ¶
Bases: ValueError
Raised when a caller asks for a build this module has no refget table for (RM15).
sha512t24u ¶
The GA4GH sha512t24u digest: unpadded base64url of the first 24 bytes of SHA-512.
24 bytes is chosen so the base64 encoding is exactly 32 characters with no padding — which is why
the digest never carries a = and why VRS_ID_PATTERN can pin the length.
Source code in schema/src/just_dna_format/vrs.py
normalize_chrom ¶
Fold a chromosome label to this module's key form (chr7 → 7, M/chrM → MT).
Source code in schema/src/just_dna_format/vrs.py
in_pseudoautosomal_region ¶
in_pseudoautosomal_region(
chrom: str | None,
start: int | None,
*,
build: str = "GRCh38",
) -> bool | None
Is this locus in a PAR? Three-valued — True, False, or None for "cannot say".
None when there is no coordinate, when the contig is not X or Y, or when the build is one this
table does not describe. A caller must not read None as False: "this row has no position" and
"this position is outside the PAR" are different facts, and the second licenses a claim about
ploidy that the first does not.
Unlike refget_accession this does not raise on another build. That function answers a
question whose wrong answer would corrupt an identity, so silence is unacceptable; this one feeds
a warning, where the honest degradation is to withhold it.
Source code in schema/src/just_dna_format/vrs.py
par_partner ¶
par_partner(
chrom: str | None,
start: int | None,
*,
build: str = "GRCh38",
) -> tuple[str, int] | None
Where this pseudoautosomal locus is spelled on the other sex contig.
Returns (contig, position) — ("Y", 640851) for ("X", 640851) — or None, which is the
withhold answer this module uses everywhere: no coordinate, a contig that is not X or Y, a locus
outside both PARs, or a build with no PAR table. A caller must not read None as "there is no
partner", only as "this function did not name one".
The arithmetic is an index-matched offset between the two contigs' intervals, which is well-defined because GRCh38's Y PAR is a copy of the X PAR: the paired intervals have equal length, so PAR1 maps at offset 0 (X:640851 ↔ Y:640851) and PAR2 at a constant 98,813,480 (X:155770036 ↔ Y:56956556). A test pins the equal-length property over the table itself, so adding a build whose intervals do not pair cannot silently corrupt the mapping.
Position agreement is a necessary condition, never a sufficient one. Two loci at partner
coordinates are the same place, but a caller deciding they are the same variant must also compare
alleles — this function knows nothing about ref/alts, and fusing on geometry alone would merge
two different variants that happen to sit opposite each other.
Source code in schema/src/just_dna_format/vrs.py
contig_length ¶
How long this contig is on build, or None for a question this table cannot answer.
None covers three quite different unknowns and deliberately reads as one: no contig given, a
build with no table, and a name that is not a primary contig of that build (a scaffold, a patch,
an alt locus). All three mean the same thing to a caller — nothing may be concluded about a
position on it — and the three-valued rule says withhold rather than guess a bound.
Unlike refget_accession this does not raise for an untabled build, for the same reason
in_pseudoautosomal_region does not: that function answers a question whose wrong answer would
corrupt an identity, while this one feeds a diagnosis, where the honest degradation is silence.
Source code in schema/src/just_dna_format/vrs.py
builds_containing_position ¶
Every tabled build whose chrom reaches start, in sorted order.
The wrong-build half of the diagnosis: a position past the end of GRCh38's chromosome 1 that is comfortably inside GRCh37's is not merely impossible, it is impossible in a way that names its cause. Empty means no tabled build has a contig of that name long enough — which is a different finding, and a caller must render the two separately.
Only the upper bound is consulted. VCF's telomere convention writes POS 0 for a variant before the first base of a contig, so a low position is not evidence of anything and this function says nothing about one.
Source code in schema/src/just_dna_format/vrs.py
sole_build_naming_contig ¶
The build that alone names this contig, or None when that cannot be established.
None is the answer for every name the tables do not settle: the 25 primary contigs (identical
across both builds), a scaffold both builds carry (GL000194.1), a patch or alt locus (in no
top-level listing at all), and an unversioned accession (GL000205 — the suffix is what
separates GRCh37's .1 from GRCh38's .2). Withholding on all of those is the point: a name this
table cannot place is not evidence of a wrong build, and treating it as such would put a false
accusation on a correct row.
Source code in schema/src/just_dna_format/vrs.py
refget_accession ¶
The refget accession for a contig on build, or None for a contig outside the table.
Raises UnsupportedBuildError for a build with no table — a caller asking for GRCh37 today gets a
clear "not built yet" rather than a silently GRCh38-flavoured answer (RM15). An explicitly passed
None or "" is such a build: the GRCh38 default lives in this signature, so omitting the
argument is the default and handing over an empty one is a caller who has not established the
build. refget_supports_build reads the same predicate and says so first.
Source code in schema/src/just_dna_format/vrs.py
refget_supports_build ¶
Is there a refget table for this assembly at all — the question refget_accession raises on.
Separate from refget_accession because the two answer different things and only one of them is
a per-contig question. refget_accession(chrom, build) has three outcomes — an accession, None
for a contig outside the table, and UnsupportedBuildError for an assembly with no table — and a
caller that wants the third has to raise and catch an exception to ask a yes/no question. Worse,
the two negatives then arrive at the same except/is None and get treated alike, which is how
sequences.verify_reference_alleles came to swallow a whole GRCh37 module row by row and report
the pass as having run: an unbuilt assembly is a statement about the module, an unmapped contig is
a statement about one row.
Kept beside the table rather than in the enricher, for the reason PAR_GRCh38 is here: which
assemblies have a refget table is a property of this tier's constants, and a copy elsewhere is a
second list to keep in step when RM15 adds the second table.
None and "" answer False, and they answered True until R2-10. The docstring above
claims to answer "the question refget_accession raises on", and for those two inputs it was
answering the opposite of what that function does — a printed claim the code did not honour, in
the tier the other two build on, and precisely on the guard a caller reaches for in order to
avoid the exception. Latent, because the one caller filters if row.genome_build first.
The old reasoning was "an unset build is the format's default, not an unbuilt assembly", and it
imports a fact about the spec layer into the identity layer. ModuleSpecConfig.genome_build
does default to GRCh38, and so does each of these functions' own signature — but an explicit
None handed to a build parameter is not an omitted argument, it is a caller who did not thread
the row's build through, which is the bug class test_build_call_sites.py walks the AST to
prevent. Every other build gate in this module already reads it that way (in_pseudoautosomal_region,
par_partner, contig_length, and the two minting functions all withhold on anything that is not
literally GRCh38), so answering True here made this one function the outlier. Both sides now
read the same predicate, so they cannot drift apart again.
Source code in schema/src/just_dna_format/vrs.py
sequence_location_digest ¶
sequence_location_digest(
chrom: str | None,
start: int,
end: int,
*,
build: str = "GRCh38",
) -> str | None
The bare sha512t24u of a VRS SequenceLocation over interbase [start, end).
start/end are interbase (0-based, half-open) — VRS's coordinate convention, not VCF's.
Returns None when the contig has no accession. The full CURIE for this digest would be
ga4gh:SL.<digest>; the bare digest is what an enclosing Allele embeds.
Source code in schema/src/just_dna_format/vrs.py
is_substitution ¶
Whether (ref, alt) is a single-base substitution — the normalization-invariant case.
This is exactly the class derive_vrs_allele_id will mint: one ACGT base to a different one ACGT
base. For it, VRS's full justification is a provable no-op (nothing to trim, nothing to shift), so
the id computed here equals the id a normalizing implementation with sequence access would compute.
Source code in schema/src/just_dna_format/vrs.py
derive_vrs_allele_id ¶
derive_vrs_allele_id(
chrom: str | None,
start: int | None,
ref: str | None,
alt: str | None,
*,
build: str = "GRCh38",
) -> str | None
The ga4gh:VA.… allele id for a resolved substitution, or None when it cannot be minted.
start is the 1-based VCF position (the convention ResolutionRow.start and the Ensembl /
ClinVar snapshots already use); the interbase conversion happens here, once, so no caller has to
remember it.
Returns None — never guesses — when the row is not a mintable substitution:
- no coordinate (an rsid-only row, pre-resolution): there is nothing to address;
- an indel or MNV: justification needs the reference sequence (see the module docstring);
- a multi-allelic cell (
altcarrying a comma): a VA names one allele, so the caller must split first — silently picking one would be a data error wearing an id. "Split first" is the whole instruction: refusing to pick is not a reason to mint nothing, which is the mistakesplit_vrs_ids/join_vrs_idsexist to stop a caller making (see their docstrings); - a contig outside the primary assembly, or a position past the end of it.
It raises UnsupportedBuildError for exactly one input: a build with no refget table (today,
anything but GRCh38). That is not an oversight and must not be softened to None — None here
means "this row is not mintable", a per-row fact, while an unknown build means the caller's whole
frame of reference is unavailable, and answering None would let a GRCh37 module quietly compile
with no identities rather than being told minting is GRCh38-only (RM15). Every call site therefore
catches it and turns it into its own kind of report: compiler._recompute_vrs_id into an
"unverifiable" reason, enricher.vrs.VrsMinter.mint into an unmintable row.
The digest is taken over the VRS Allele serialization, in which the location appears as its own
digest (not inlined and not as a ga4gh:SL. CURIE) — that exact shape is what reproduces the ids
gnomAD serves, and the test suite pins it against recorded ground truth rather than trusting the
reading of any spec text.
Source code in schema/src/just_dna_format/vrs.py
split_vrs_ids ¶
A comma-joined vrs_id cell → one entry per ALT, None where no id was minted.
A VRS allele id names one allele, and alts is a column that may name several. Two ways to
reconcile that, and only one of them is honest. The first — mint nothing for a multi-allelic row —
is what this codebase did, borrowing derive_vrs_allele_id's refusal to pick an allele and
applying it to a column where nothing is being picked. It cost the id on 909 of 1,613 rows in one
real module while every input needed to compute all 2,110 of them sat in the same row.
The second is this: vrs_id is a parallel array of alts, so member i names alt i. A hole
(an empty member) is a real value — "this allele's id could not be minted here" — because a
substitution and an indel can share one site and only one of them mints offline. Losing the whole
row's ids to the indel beside them would be the same abstention one level down.
A single-alt row degenerates to a bare id, byte-identical to what every existing file carries, so
this widening reaches no module that does not need it. It also moves no signature: vrs_id is
outside RESOLUTION_FACT_FIELDS and reverse_module does not re-emit it.
Source code in schema/src/just_dna_format/vrs.py
join_vrs_ids ¶
Per-allele ids → the comma-joined cell, or None when not one of them was minted.
Holes are kept, so position is preserved; an all-holes result is None rather than ",,", since
a row that minted nothing is exactly the row that used to carry no id at all.
Source code in schema/src/just_dna_format/vrs.py
validate_vrs_id_list ¶
Validate a comma-joined vrs_id cell member by member, returning the canonical spelling.
Each member is either a well-formed VRS allele id or empty. Alignment with alts is not
checked here — a field validator cannot see a sibling field — so ResolutionRow checks the count
itself.
Source code in schema/src/just_dna_format/vrs.py
validate_vrs_id ¶
Validate an optional GA4GH VRS identifier of any identifiable type.
Well-formedness only: ga4gh:<TYPE>.<32-char digest>. Deliberately lenient (see
VRS_ID_PATTERN) so a validator rejects a malformed value rather than an
unfamiliar-but-well-formed one. Whether a given column may hold a non-allele type is a
separate question, and for the two vrs_id columns the answer is no — they use
validate_vrs_allele_id.
Source code in schema/src/just_dna_format/vrs.py
validate_vrs_allele_id ¶
Well-formed and an allele id — the rule ResolutionRow.vrs_id/FrequencyRow.vrs_id obey.
Settling this is what RM78's severity question was gated on. Both columns are described as
"GA4GH VRS allele id (ga4gh:VA.…) — one per ALT", and only the lenient format check ran — so a
ga4gh:SL.…, a sequence-location id naming a place rather than an allele, loaded cleanly. It
then reached _verify_vrs_ids, which recomputes with derive_vrs_allele_id (always a VA), found
a difference, and reported a mismatch: "recomputed and different, so corruption". That
verdict is already an error in both modes, so this tightening makes no passing module fail — it
replaces a confident wrong diagnosis with the true one, at load, naming the type it got.
Legal because it has no instantiation, which is the test this repo applies to a tightening
(the IUPAC probe is the precedent). Nothing mints a non-VA into either column — both minting paths
call derive_vrs_allele_id, and gnomAD's own id is an allele id — and a probe across all sixteen
reference examples found 844 ids, every one ga4gh:VA., zero of the other four types.
validate_vrs_id keeps its documented lenience and its place. What was wrong was using a format
check where a column rule was meant: a column that holds one kind of thing should say which.
Source code in schema/src/just_dna_format/vrs.py
validate_caid ¶
Validate an optional ClinGen Allele Registry canonical allele id (CA<digits>).