just_dna_enricher.civic_identities¶
just_dna_enricher.civic_identities ¶
Identities CIViC states in a variant's name and never puts in its identifier columns.
Why this table exists at all. civic build places a row from what the source publishes — an
rs-number, or a GRCh38 RefSeq accession it can parse. 53 variants in the 01-Aug-2026 release carry
neither and are dropped as unresolvable_identity, and for most of them the identity was published
the whole time, in the variant's own name: N150fs (c.448delA), IVS2+1G>A, D1709N. A c. or
protein fragment plus the gene's numbering frame is an allele, and an allele registry will hold it.
Why the answers are a shipped constant rather than a lookup. Resolving a name needs the network —
the ClinGen Allele Registry, NCBI, Ensembl — and civic build must stay byte-reproducible from a
pinned dated release, which is why the CAID pass runs at draft time and never in a build. But these
identities are not a lookup's output: four of them required a judgement no lookup makes. A legacy
IVS2 name that converts structurally to the wrong exon; a name pairing a missense protein label with
a synonymous cDNA change; a protein consequence standing over an intronic allele; an rs-number that is
position-level where two alleles spell the same substitution. Each was adjudicated by hand, against
evidence recorded per variant, and re-deriving that at build time would either be impossible or would
silently pick a side. So the answer is data, the procedure is written down separately, and the
build stays offline.
The procedure is docs/probes/CIVIC_IDENTITY_PROTOCOL.md; the per-variant evidence and the classes
that did not resolve are docs/probes/CIVIC_UNRESOLVED.md.
What is deliberately not here.
- CIViC 4968
TP53 R72P. It resolves — rs1042522, CA178298 — but its identity is the reference allele (NC_000017.11:g.7676154G=,c.215C=): codon 72 isCCC= Pro on GRCh38, so CIViC's name has reference and alternate inverted. A snapshot row ischrom/start/ref/alt, andref == altis not a variant row — the compiler drops such rows by design. The identity exists and this format cannot carry it, which is a property of the representation rather than of the record. - The two conjunction-named records (3298, 4210). Each names two alterations; all four alleles are identified, and minting one identity for a record naming two would assert a locus the source did not. A conjunction is a haplotype plus a diplotype, not a variant row.
- The 9 liftover-only, the 6 that name a class of event, and the 2 with two readings and nothing to choose between them. No identity to carry, or no way to pick one.
The name is the key, and that is the safety property. Every identity below was derived from the
name string quoted beside it. A build applies a row only when the release's name matches that string
exactly, so a curated answer can never outlive the record it was an answer to: if CIViC re-names a
variant, its row goes stale and is counted rather than applied. And a row whose identifier columns
CIViC has since filled goes superseded — the source always wins over this table, which also makes a
supersession the cheapest currency signal there is.
CivicNameIdentity
dataclass
¶
CivicNameIdentity(
variant_id: int,
gene: str,
name: str,
chrom: str,
start: int,
ref: str,
alt: str,
rsid: str | None = None,
caid: str | None = None,
note: str | None = None,
)
One curated identity, keyed to the CIViC variant id and the name it was derived from.