just_dna_enricher.vrs¶
just_dna_enricher.vrs ¶
VRS allele-id minting for the resolution table — the enricher half of the identity story.
The format tier mints a ga4gh:VA.… for a substitution with pure stdlib and no sequence access
(just_dna_format.vrs.derive_vrs_allele_id). That covers most rows and costs nothing, online or off.
What it cannot do is an indel: a VRS allele id is defined over the fully justified allele, and
justifying an indel means reading the reference sequence around it to see how far the event can slide.
Sequence access is network access, and network access belongs here (Principle 2) — so the indel half
of minting lives in this module and nowhere else.
Complete allele identity is the default, not an extra. The plan budgeted for ga4gh.vrs[extras]
— seqrepo + pysam + hgvs, a compiled extension and a multi-gigabyte local sequence store — on
the belief that normalization needs a local seqrepo. Probing showed it does not: core ga4gh.vrs
plus the seqrepo REST data proxy normalizes indels over HTTP, at a cost of 14 pure-Python packages
with no compiled dependencies. Reading a remote sequence is exactly what a network tier is for, so at
that price there is no reason to make complete identity opt-in. ga4gh.vrs is therefore a core
dependency, indels are minted by default, and --offline is the only thing that turns it off.
(AlleleTranslator is deliberately unused despite being the obvious entry point: it imports hgvs at
module scope, which drags the heavy tree back in for a convenience wrapper we do not need.
normalize() + ga4gh_identify() are the two functions that matter.)
Three ways a row can get an id, in precedence order:
- Minted locally, stdlib — a substitution. Zero egress, and verified against the library itself
to be byte-identical (see
schema/tests/test_vrs.py), so the offline answer and the online answer for a substitution are the same answer. - Minted, normalized — an indel/MNV, justified against the reference over the REST proxy. Needs
the network, so
--offlineskips it. - Left null — an indel in an offline run, an unreachable sequence service, an off-assembly
contig, or an allele that names no sequence at all (
_sequence_free_reason). A missing id is the honest outcome; an unjustified one would be aga4gh:VA.…string that looks interoperable and silently is not.
The last of those is permanent, and keeping it apart from the rest is the point: an indel offline
is a re-run away, while a <DEL:4977> has no sequence for any run to justify, so a reason that offers
the re-run is a false remedy as well as a misdiagnosis.
All three are decided per ALT, not per row: vrs_id is a comma-joined parallel array of alts
(just_dna_format.vrs.split_vrs_ids), so a multi-allelic site names each of its alleles and a hole
marks the one that could not be minted. Deciding it per row is what left 909 of 1,613 rows in a real
module with no id at all while every input needed to mint them sat in the same row.
A source's own id (gnomAD serves one) is never written over a locally-minted value. It is compared against it, and a disagreement is logged loudly — that cross-check is free provenance, and it is how a bug in either implementation would surface.
MintResult
dataclass
¶
MintResult(
minted_stdlib: int = 0,
minted_normalized: int = 0,
skipped_unmintable: int = 0,
already_present: int = 0,
mismatches: list[str] = list(),
unmintable_reasons: dict[str, int] = dict(),
alleles: int = 0,
identified: int = 0,
)
What a minting pass did — counted by how, because the routes have different costs and different verifiability (the compiler can recompute a stdlib id and cannot recompute a normalized one).
The three minting counters are per allele, not per row: a multi-allelic site mints one id per
ALT, so a row can contribute to minted_stdlib and skipped_unmintable at once. already_present
is per row, because an existing vrs_id is left alone whole — the pass does not reach inside a
cell it did not write. For the single-alt rows that are most of any table the two granularities
coincide, which is why the counters kept their names.
complete
property
¶
Whether every allele in the table now carries an id.
The question a caller actually has, and the reason this pass reports more than counts: a VA is
becoming the identity these tables are keyed on, and an identity scheme covering an unstated
fraction of the table is not one anything can key on. alleles == 0 (an empty table) is
vacuously complete — there is no shortfall to report about nothing.
coverage_warnings ¶
The shortfall as text, one line plus one per reason — or nothing when coverage is complete.
Returned rather than logged so the CLI, enrich() and a library caller all say the same thing.
Source code in enricher/src/just_dna_enricher/vrs.py
VrsMinter
dataclass
¶
VrsMinter(
seqrepo_uri: str = DEFAULT_SEQREPO_URI,
offline: bool = False,
sequences: SequenceProxy | None = None,
)
Mints allele ids: substitutions always, indels whenever the run is allowed to reach the network.
Sequence access comes from a shared SequenceProxy, so a run that both mints indels and checks
reference alleles builds one proxy and shares one read cache between them.
mint ¶
mint(
chrom: str | None,
start: int | None,
ref: str | None,
alt: str | None,
*,
build: str = "GRCh38",
) -> tuple[str | None, str | None]
Return (vrs_id, how) where how is "stdlib", "normalized", or None.
Reporting how rather than only what is deliberate: a normalized id depended on a live sequence service and a substitution's did not, and the compiler's verify pass — which can recompute one class and not the other — needs to be able to tell them apart.
Source code in enricher/src/just_dna_enricher/vrs.py
why_not ¶
why_not(
chrom: str | None,
start: int | None,
ref: str | None,
alt: str | None,
*,
build: str = "GRCh38",
) -> str
Why mint returned nothing for this allele — for grouping, not for a per-row line.
Deliberately not shared with the compiler's _recompute_vrs_id, which answers a different
question with different answers: an indel is permanently unrecomputable there (P2 keeps the
reference sequence out of that tier) and merely offline here. Sharing the wording would make
one of the two lie about what a re-run could fix, which is the whole point of reporting a
reason at all.
The branches are in mint's order on purpose, which is what puts the permanent classes ahead
of the build and the offline checks: a symbolic allele on a GRCh37 row is not waiting on a
refget table, and a refget table would not make it mintable.
Source code in enricher/src/just_dna_enricher/vrs.py
mint_resolution_rows ¶
mint_resolution_rows(
rows: list[ResolutionRow],
*,
minter: VrsMinter | None = None,
offline: bool = False,
source_ids: dict[str, str] | None = None,
) -> MintResult
Stamp vrs_id/vrs_spec onto every mintable row, in place. Returns what it did.
An existing vrs_id is never overwritten — the same "already there is authoritative" rule
enrich() applies to the rest of the table, so a hand-corrected id survives a re-run.
source_ids maps variant_key to an id a source reported (gnomAD serves one per variant). It is
used only as a cross-check against what we minted, never as the value: trusting it would make our
identity depend on which sources happened to answer, and the whole point of a content-addressed id
is that it does not.
The result also reports coverage — allele slots seen, how many are named, and what the rest are
blocked on (MintResult.coverage_warnings). Callers are expected to surface a shortfall rather than
only the success count: a VA is on its way to being the key these tables are joined on, and a pass
that says "minted 237" while leaving 185 alleles anonymous has reported its own success and hidden
the number that decides whether anything can key on the result.