just_dna_enricher.litvar¶
just_dna_enricher.litvar ¶
LitVar2/PubTator3 — which papers name this allele, and at which tier (RM167).
NCBI's LitVar2 indexes the biomedical literature by variant, and it does so at several tiers that sit
beside each other as separate nodes: litvar@rs1800562## is the position node for an rsID, and
litvar@CA113795#rs1800562## is the allele node beside it, named by a ClinGen canonical allele
id. A gene carries one gene-level node (litvar@#3077#, NCBI GeneID 3077 = HFE) and a long tail of
litvar@#7428#p.P71fsX text mentions, which are an unnormalized string a miner saw rather than an
identity. Nothing here treats a mention as a variant.
The id is litvar@ plus #-separated fields whose count and meaning vary by tier, which is
exactly why nothing in this module parses one: an rsID node has three fields with the rsID first, an
allele node has four with the CAID first, and a gene node has three with the gene id in the middle.
Read the tier off the source's own flag_rsid_variant / flag_clingen_variant / flag_gene_variant
booleans, or off which keys a record carries where the endpoint omits them, and echo every id back
verbatim rather than building one.
The tier a locus is answerable at is a property of the locus, not of the source, and that is the
whole finding this pass exists to make. Measured 2026-09-01: BRAF rs113488022 carries 32,095 papers on
its position node and three allele nodes (31,276 / 99 / 41) whose three ALTs at one codon differ by
three orders of magnitude, leaving 801 papers (2.5 %) position-only. APOE rs429358 carries 3,945 on
the position node and 328 on its single allele node, so 92 % of the literature at that locus is
not allele-resolved. A pass that reported the allele node's count as the answer would understate
APOE twelvefold, and one that quietly substituted the position count would answer an allele-level
question with a position-level fact. So the answer names its tier: allele-resolved, position-only
and absent are three outcomes, and for once the source supplies the three states rather than the
schema imposing them (@refutation-withholds).
What this does NOT answer, and the bound is not a footnote. LitVar tells you which papers discuss
an allele that is already identified. It does not tell you which allele a name meant. Those read as
the same question and are not. Asked of the two records this workspace could not resolve — CIViC 1955
VHL P71fs (c.211insT) and 2131 VHL Q73fs (c.214insGCCC), worked down by hand to four candidate
alleles with registered CAIDs — LitVar returns no node for any of the four, no node for
c.211insT / 211insT / c.214insGCCC, and for VHL P71fs exactly one node,
litvar@#7428#p.P71fsX with 1 PMID: 19996202, which is none of the four source papers but an
unrelated paper that happens to write "P71fs" (@existence-not-identity). The reason is structural
rather than incidental: PubTator3's export for all four source papers returns title and abstract only,
two passages with zero variant annotations in every one, none of them in the PMC open-access
subset. The alleles live in Table 3 of a paywalled 1996–2007 paper, and text mining over abstracts
cannot reach a table. On precisely the class this workspace built an identity protocol for — a source
that names a variant without identifying it, in old literature — LitVar is the wrong instrument.
Terms are NCBI's policy, and a policy is not a licence. NCBI states it "places no restrictions on
the use or distribution" of molecular data and, in the same passage, that it "cannot provide comment
or unrestricted permission concerning the use, copying, or distribution" because submitters may hold
rights it cannot assess. ClinVar escapes this through its own maintenance_use page, which is why
CLINVAR_TERMS records public-domain; LitVar has no such page, so under @no-named-licence its
gating axes are unknown rather than permissive, and recording it as public domain by analogy with
ClinVar is exactly the move that rule forbids. This is NCBI's side only — nothing was read about
EMBL-EBI's terms for surfaces EBI co-hosts, and nothing here asserts anything about them.
No SourceRow is written, and that is the rule rather than an omission (@write-the-sourcerow,
its converse). sources.csv travels to the registry meaning this module uses this source, and this
pass contributes no cell to any table: it compares, it reports, and it writes nothing but an
attestation. identifiers.py is the precedent — it consults HGNC and OLS4 for the same kind of
read-only verdict and writes no row either — while civic_draft.py, which does put registry-derived
values into a module, writes its clingen_allele_registry row. The source is named on the
VerificationRecord instead, which is where a check's provenance belongs.
One real API defect, pinned before anyone writes a second client. variant/search/gene/GENE
returns line-delimited Python repr(), not JSON — single-quoted keys, one dict per line. httpx's
.json() raises on it. The other endpoints return proper JSON. parse_repr_lines below is a literal
parser (ast.literal_eval), never eval (@probe-the-real-file).
LitvarError ¶
Bases: RuntimeError
LitVar could not be consulted in a way the caller must handle.
LitvarUnavailable ¶
Bases: LitvarError
The service was not reachable, so the question was never put (RM101's shape).
A subclass rather than a second exception: every existing except LitvarError still fires, and a
caller that needs to separate the index said no from the index never answered can, without
reading __cause__. Only this one means nobody was asked.
LitvarNode
dataclass
¶
LitvarNode(
node_id: str,
tier: str,
rsid: str | None = None,
clingen_id: str | None = None,
name: str | None = None,
genes: tuple[str, ...] = (),
pmid_count: int | None = None,
)
One node in the index, at one tier.
parse
classmethod
¶
One listing record → a node, or LitvarError when it is not one.
Every field is coerced rather than trusted. A record with no _id used to raise a bare
KeyError past every handler in this tier, and a non-string rsid an AttributeError out of
the exact-match filter — two ways for a payload shape to arrive as a crash rather than as
this client's own error type.
Source code in enricher/src/just_dna_enricher/litvar.py
LitvarClient ¶
LitvarClient(
*,
base: str = LITVAR_API_BASE,
client: Client | None = None,
gate: PacingGate | None = None,
)
The four LitVar2 calls this lane needs, paced, retried and translated.
Node ids are never constructed from the grammar. Every id this client hands to get or
publications came back from autocomplete or search/gene verbatim. The first attempt at this
source read the trailing ## of litvar@rs1800562## as a suffix on the rsID, concluded there was
no allele tier, and was wrong about the whole source on the strength of one misread character; an
id that is only ever echoed cannot reproduce that.
Source code in enricher/src/just_dna_enricher/litvar.py
autocomplete ¶
Every node whose id, rsID or synonyms match query, at whatever tier each sits.
The match is a prefix search, so the caller must filter. ?query=rs429358 returns
rs42935848 as a second hit — a real node for a different variant whose number begins with
the one asked about. position_node and allele_node below are that filter; nothing should
take [0] off this list.
Source code in enricher/src/just_dna_enricher/litvar.py
node ¶
The node's own record — clingen_ids lives here and nowhere else — or None if absent.
Source code in enricher/src/just_dna_enricher/litvar.py
pmids ¶
Every PubMed id the index holds for this node. An absent node answers with an empty set.
The payload states its own total and this compares against it (@dont-discard-computed):
publications carries pmids_count beside pmids, the two agree on every recorded response,
and a paginated or truncated body is otherwise a confidently wrong number with the checker
sitting in the same dict. A disagreement warns rather than raising — the ids really were
served, and refusing them would turn a partial answer into no answer — but it never passes
silently.
A member that is not a PubMed id is a shape this client cannot read, and it raises: dropping it would report a short list as the node's literature.
Source code in enricher/src/just_dna_enricher/litvar.py
gene_nodes ¶
Every node LitVar holds under a gene symbol, across all four id shapes.
This is the endpoint that serves Python repr() rather than JSON. Nothing calls .json() on
it; parse_repr_lines does the reading.
Source code in enricher/src/just_dna_enricher/litvar.py
position_node ¶
The position node for exactly this rsID, or None when the index holds none.
Source code in enricher/src/just_dna_enricher/litvar.py
allele_node ¶
The allele node for exactly this CAID, or None when the index holds none.
Source code in enricher/src/just_dna_enricher/litvar.py
LocusRoster
dataclass
¶
LocusRoster(
alleles: dict[str, list[str]] = dict(),
refs: dict[str, str] = dict(),
starts: dict[str, int] = dict(),
caids: dict[str, str] = dict(),
read: list[str] = list(),
not_read: dict[str, str] = dict(),
)
The loci a module names, the alleles it names at each, and which tables that came out of.
LocusCoverage
dataclass
¶
LocusCoverage(
rsid: str,
asked_tier: str,
tier: str,
reason: str,
position_pmids: int | None = None,
allele_pmids: int | None = None,
position_only_pmids: int | None = None,
caids_at_locus: tuple[str, ...] = (),
caids_with_a_node: tuple[str, ...] = (),
matched_caids: tuple[str, ...] = (),
node_id: str | None = None,
)
What the index holds for one module locus, and at which tier it said it.
degraded
property
¶
An allele-level question answered at position level — the finding this pass exists for.
LiteratureCoverageReport
dataclass
¶
LiteratureCoverageReport(
loci: list[LocusCoverage] = list(),
tables_read: list[str] = list(),
tables_not_read: dict[str, str] = dict(),
offline: bool = False,
)
One module's literature coverage, per locus and per tier.
parse_repr_lines ¶
One dict per line of a Python repr() payload, which is what search/gene really serves.
ast.literal_eval is a literal parser, not eval: it walks the parsed syntax tree and refuses
anything that is not a literal container, so it cannot call, import or execute. Calling .json()
on this payload raises, and calling eval on it would be a remote-code path.
A line that will not parse is a defect in the response rather than in one record, so it raises; silently dropping it would make a short answer indistinguishable from a complete one.
Source code in enricher/src/just_dna_enricher/litvar.py
node_tier ¶
Which tier a node record sits at, from whichever evidence the endpoint supplied.
The flag_rsid_variant / flag_clingen_variant / flag_gene_variant booleans say it directly and
are present on autocomplete and get — but search/gene omits all three, carrying only the
keys a record has. So both routes are read, flags first: an id is parsed only where nothing states
the answer, because parsing an id is exactly how the first pass at this got it wrong.
Source code in enricher/src/just_dna_enricher/litvar.py
rsid_bearing_tables ¶
{filename: model} for every authored table whose model declares rsid.
Derived from DRAFTABLE rather than restated (@registry-completeness), the same walk
identifiers._id_bearing_tables does for its own columns — a table kind added later joins this
roster by existing, not by somebody remembering it.
Source code in enricher/src/just_dna_enricher/litvar.py
module_loci ¶
Every rsID a module names, with the alleles it names there and any CAID it already holds.
The allele columns come from the model (AuthoredModel.ALLELE_COLUMNS), never from a list
here: variants.csv states an allele in alts, genotype and effect_allele, haplotypes.csv
in allele, diplotypes.csv in genotype, and a hand-kept list would be the roster defect one
layer down. The ref column is skipped because it states the locus rather than a claim about
an allele — but a genotype cell names the reference base too, so the bag is the alleles the
module writes here, not the alternates. That is deliberate and the REF comparison in
_matching_caids is what keeps it honest: a registry allele whose reference base disagrees with
the row's is refused. Where no table at a locus states a ref there is nothing to refuse it with,
which is a real if narrow way for a reference base to match an alternate.
resolution.csv is read too, for its alts and caid columns — the CAID is the allele identity
RM153 puts there, and it is the shortest route from a module row to an allele node.
Source code in enricher/src/just_dna_enricher/litvar.py
coverage_reason ¶
Why this locus landed on its tier, one sentence per arm.
A verdict function with several arms owes a reason function with the same arms, pairwise distinct
(@answered-is-not-absent) — otherwise two situations a reader must tell apart share a sentence,
which is how a strand-flipped SNV was once reported as an event-size disagreement.
Source code in enricher/src/just_dna_enricher/litvar.py
check_literature_coverage ¶
check_literature_coverage(
spec_dir: Path,
*,
client: LitvarClient | None = None,
registry: ClingenAlleleClient | None = None,
offline: bool = False,
progress: Callable[[int, int], None] | None = None,
) -> LiteratureCoverageReport
Ask LitVar what it holds for each of a module's loci, and at which tier.
Reports and repairs nothing (@enrichment-is-validation). Nothing here is written into a module:
a PMID list per variant is not a table kind, literature.csv is keyed by article, and 32,095
PMIDs for one BRAF locus would be a row-writer arguing against itself. What lands is the
attestation and this report.
progress is called with (done, total) over loci, the unit @progress-unit-is-subjects
fixes, because the total has to be known before the first request goes out.
Source code in enricher/src/just_dna_enricher/litvar.py
verification_records ¶
The attestation, and it names the tier — a coverage answer that does not is the defect.
One record. subjects is the loci the index answered about, findings the loci where an
allele-level question came back position-level, and detail carries the whole tier breakdown,
because "12 checked, 3 flagged" says nothing about which of 328 and 3,945 a reader is holding.