Skip to content

just_dna_enricher.litvar

just_dna_enricher.litvar

LitVar2/PubTator3 — which papers name this allele, and at which tier (RM167).

NCBI's LitVar2 indexes the biomedical literature by variant, and it does so at several tiers that sit beside each other as separate nodes: litvar@rs1800562## is the position node for an rsID, and litvar@CA113795#rs1800562## is the allele node beside it, named by a ClinGen canonical allele id. A gene carries one gene-level node (litvar@#3077#, NCBI GeneID 3077 = HFE) and a long tail of litvar@#7428#p.P71fsX text mentions, which are an unnormalized string a miner saw rather than an identity. Nothing here treats a mention as a variant.

The id is litvar@ plus #-separated fields whose count and meaning vary by tier, which is exactly why nothing in this module parses one: an rsID node has three fields with the rsID first, an allele node has four with the CAID first, and a gene node has three with the gene id in the middle. Read the tier off the source's own flag_rsid_variant / flag_clingen_variant / flag_gene_variant booleans, or off which keys a record carries where the endpoint omits them, and echo every id back verbatim rather than building one.

The tier a locus is answerable at is a property of the locus, not of the source, and that is the whole finding this pass exists to make. Measured 2026-09-01: BRAF rs113488022 carries 32,095 papers on its position node and three allele nodes (31,276 / 99 / 41) whose three ALTs at one codon differ by three orders of magnitude, leaving 801 papers (2.5 %) position-only. APOE rs429358 carries 3,945 on the position node and 328 on its single allele node, so 92 % of the literature at that locus is not allele-resolved. A pass that reported the allele node's count as the answer would understate APOE twelvefold, and one that quietly substituted the position count would answer an allele-level question with a position-level fact. So the answer names its tier: allele-resolved, position-only and absent are three outcomes, and for once the source supplies the three states rather than the schema imposing them (@refutation-withholds).

What this does NOT answer, and the bound is not a footnote. LitVar tells you which papers discuss an allele that is already identified. It does not tell you which allele a name meant. Those read as the same question and are not. Asked of the two records this workspace could not resolve — CIViC 1955 VHL P71fs (c.211insT) and 2131 VHL Q73fs (c.214insGCCC), worked down by hand to four candidate alleles with registered CAIDs — LitVar returns no node for any of the four, no node for c.211insT / 211insT / c.214insGCCC, and for VHL P71fs exactly one node, litvar@#7428#p.P71fsX with 1 PMID: 19996202, which is none of the four source papers but an unrelated paper that happens to write "P71fs" (@existence-not-identity). The reason is structural rather than incidental: PubTator3's export for all four source papers returns title and abstract only, two passages with zero variant annotations in every one, none of them in the PMC open-access subset. The alleles live in Table 3 of a paywalled 1996–2007 paper, and text mining over abstracts cannot reach a table. On precisely the class this workspace built an identity protocol for — a source that names a variant without identifying it, in old literature — LitVar is the wrong instrument.

Terms are NCBI's policy, and a policy is not a licence. NCBI states it "places no restrictions on the use or distribution" of molecular data and, in the same passage, that it "cannot provide comment or unrestricted permission concerning the use, copying, or distribution" because submitters may hold rights it cannot assess. ClinVar escapes this through its own maintenance_use page, which is why CLINVAR_TERMS records public-domain; LitVar has no such page, so under @no-named-licence its gating axes are unknown rather than permissive, and recording it as public domain by analogy with ClinVar is exactly the move that rule forbids. This is NCBI's side only — nothing was read about EMBL-EBI's terms for surfaces EBI co-hosts, and nothing here asserts anything about them.

No SourceRow is written, and that is the rule rather than an omission (@write-the-sourcerow, its converse). sources.csv travels to the registry meaning this module uses this source, and this pass contributes no cell to any table: it compares, it reports, and it writes nothing but an attestation. identifiers.py is the precedent — it consults HGNC and OLS4 for the same kind of read-only verdict and writes no row either — while civic_draft.py, which does put registry-derived values into a module, writes its clingen_allele_registry row. The source is named on the VerificationRecord instead, which is where a check's provenance belongs.

One real API defect, pinned before anyone writes a second client. variant/search/gene/GENE returns line-delimited Python repr(), not JSON — single-quoted keys, one dict per line. httpx's .json() raises on it. The other endpoints return proper JSON. parse_repr_lines below is a literal parser (ast.literal_eval), never eval (@probe-the-real-file).

LitvarError

Bases: RuntimeError

LitVar could not be consulted in a way the caller must handle.

LitvarUnavailable

Bases: LitvarError

The service was not reachable, so the question was never put (RM101's shape).

A subclass rather than a second exception: every existing except LitvarError still fires, and a caller that needs to separate the index said no from the index never answered can, without reading __cause__. Only this one means nobody was asked.

LitvarNode dataclass

LitvarNode(
    node_id: str,
    tier: str,
    rsid: str | None = None,
    clingen_id: str | None = None,
    name: str | None = None,
    genes: tuple[str, ...] = (),
    pmid_count: int | None = None,
)

One node in the index, at one tier.

parse classmethod

parse(record: dict) -> LitvarNode

One listing record → a node, or LitvarError when it is not one.

Every field is coerced rather than trusted. A record with no _id used to raise a bare KeyError past every handler in this tier, and a non-string rsid an AttributeError out of the exact-match filter — two ways for a payload shape to arrive as a crash rather than as this client's own error type.

Source code in enricher/src/just_dna_enricher/litvar.py
@classmethod
def parse(cls, record: dict) -> "LitvarNode":
    """One listing record → a node, or `LitvarError` when it is not one.

    Every field is coerced rather than trusted. A record with no `_id` used to raise a bare
    `KeyError` past every handler in this tier, and a non-string `rsid` an `AttributeError` out of
    the exact-match filter — two ways for a payload shape to arrive as a crash rather than as
    this client's own error type.
    """
    node_id = record.get("_id")
    if not node_id:
        raise LitvarError("a LitVar record carries no `_id`, so it names no node")
    genes = record.get("gene") or []
    count = record.get("pmids_count")
    return cls(
        node_id=str(node_id),
        tier=node_tier(record),
        rsid=str(record["rsid"]) if record.get("rsid") else None,
        clingen_id=str(record["clingen_id"]) if record.get("clingen_id") else None,
        name=str(record["name"]) if record.get("name") else None,
        genes=tuple(str(g) for g in genes) if isinstance(genes, list) else (),
        pmid_count=int(count) if isinstance(count, int) else None,
    )

LitvarClient

LitvarClient(
    *,
    base: str = LITVAR_API_BASE,
    client: Client | None = None,
    gate: PacingGate | None = None,
)

The four LitVar2 calls this lane needs, paced, retried and translated.

Node ids are never constructed from the grammar. Every id this client hands to get or publications came back from autocomplete or search/gene verbatim. The first attempt at this source read the trailing ## of litvar@rs1800562## as a suffix on the rsID, concluded there was no allele tier, and was wrong about the whole source on the strength of one misread character; an id that is only ever echoed cannot reproduce that.

Source code in enricher/src/just_dna_enricher/litvar.py
def __init__(
    self,
    *,
    base: str = LITVAR_API_BASE,
    client: httpx.Client | None = None,
    gate: PacingGate | None = None,
) -> None:
    self._base = base.rstrip("/")
    self._client = client
    self._gate = gate or PacingGate(_REQUEST_INTERVAL)
    # Per-run, per-client: a locus reached from two tables must not be two requests, and a
    # module-level cache would make two callers in one process share state they never agreed on.
    self._nodes: dict[str, list[LitvarNode]] = {}
    self._details: dict[str, dict] = {}
    self._pmids: dict[str, frozenset[int]] = {}

autocomplete

autocomplete(query: str) -> list[LitvarNode]

Every node whose id, rsID or synonyms match query, at whatever tier each sits.

The match is a prefix search, so the caller must filter. ?query=rs429358 returns rs42935848 as a second hit — a real node for a different variant whose number begins with the one asked about. position_node and allele_node below are that filter; nothing should take [0] off this list.

Source code in enricher/src/just_dna_enricher/litvar.py
def autocomplete(self, query: str) -> list[LitvarNode]:
    """Every node whose id, rsID or synonyms match `query`, at whatever tier each sits.

    **The match is a prefix search, so the caller must filter.** `?query=rs429358` returns
    `rs42935848` as a second hit — a real node for a different variant whose number begins with
    the one asked about. `position_node` and `allele_node` below are that filter; nothing should
    take `[0]` off this list.
    """
    key = query.strip()
    if key not in self._nodes:
        status, body = self._request(f"/variant/autocomplete/?query={urllib.parse.quote(key)}")
        self._nodes[key] = (
            [] if status is None else [LitvarNode.parse(record) for record in _as_list(body)]
        )
    return list(self._nodes[key])

node

node(node_id: str) -> dict | None

The node's own record — clingen_ids lives here and nowhere else — or None if absent.

Source code in enricher/src/just_dna_enricher/litvar.py
def node(self, node_id: str) -> dict | None:
    """The node's own record — `clingen_ids` lives here and nowhere else — or `None` if absent."""
    if node_id not in self._details:
        status, body = self._request(f"/variant/get/{urllib.parse.quote(node_id, safe='')}")
        self._details[node_id] = {} if status is None else _as_dict(body)
    detail = self._details[node_id]
    return detail or None

pmids

pmids(node_id: str) -> frozenset[int]

Every PubMed id the index holds for this node. An absent node answers with an empty set.

The payload states its own total and this compares against it (@dont-discard-computed): publications carries pmids_count beside pmids, the two agree on every recorded response, and a paginated or truncated body is otherwise a confidently wrong number with the checker sitting in the same dict. A disagreement warns rather than raising — the ids really were served, and refusing them would turn a partial answer into no answer — but it never passes silently.

A member that is not a PubMed id is a shape this client cannot read, and it raises: dropping it would report a short list as the node's literature.

Source code in enricher/src/just_dna_enricher/litvar.py
def pmids(self, node_id: str) -> frozenset[int]:
    """Every PubMed id the index holds for this node. An absent node answers with an empty set.

    **The payload states its own total and this compares against it** (`@dont-discard-computed`):
    `publications` carries `pmids_count` beside `pmids`, the two agree on every recorded response,
    and a paginated or truncated body is otherwise a confidently wrong number with the checker
    sitting in the same dict. A disagreement warns rather than raising — the ids really were
    served, and refusing them would turn a partial answer into no answer — but it never passes
    silently.

    A member that is not a PubMed id is a shape this client cannot read, and it raises: dropping
    it would report a short list as the node's literature.
    """
    if node_id not in self._pmids:
        status, body = self._request(f"/variant/get/{urllib.parse.quote(node_id, safe='')}/publications")
        payload = {} if status is None else _as_dict(body)
        values = payload.get("pmids") or []
        pmids = frozenset(_as_pmid(value, node_id) for value in values)
        stated = payload.get("pmids_count")
        if isinstance(stated, int) and stated != len(pmids):
            logger.warning(
                "LitVar says %s holds %d paper(s) and served %d — the coverage number for this "
                "node is the served set, and it is short of what the index claims.",
                node_id,
                stated,
                len(pmids),
            )
        self._pmids[node_id] = pmids
    return self._pmids[node_id]

gene_nodes

gene_nodes(gene: str) -> list[LitvarNode]

Every node LitVar holds under a gene symbol, across all four id shapes.

This is the endpoint that serves Python repr() rather than JSON. Nothing calls .json() on it; parse_repr_lines does the reading.

Source code in enricher/src/just_dna_enricher/litvar.py
def gene_nodes(self, gene: str) -> list[LitvarNode]:
    """Every node LitVar holds under a gene symbol, across all four id shapes.

    This is the endpoint that serves Python `repr()` rather than JSON. Nothing calls `.json()` on
    it; `parse_repr_lines` does the reading.
    """
    status, body = self._request(f"/variant/search/gene/{urllib.parse.quote(gene, safe='')}")
    if status is None:
        return []
    return [LitvarNode.parse(record) for record in parse_repr_lines(body)]

position_node

position_node(rsid: str) -> LitvarNode | None

The position node for exactly this rsID, or None when the index holds none.

Source code in enricher/src/just_dna_enricher/litvar.py
def position_node(self, rsid: str) -> LitvarNode | None:
    """The position node for exactly this rsID, or `None` when the index holds none."""
    wanted = rsid.strip().lower()
    for node in self.autocomplete(rsid):
        if node.tier == "rsid" and (node.rsid or "").lower() == wanted:
            return node
    return None

allele_node

allele_node(caid: str) -> LitvarNode | None

The allele node for exactly this CAID, or None when the index holds none.

Source code in enricher/src/just_dna_enricher/litvar.py
def allele_node(self, caid: str) -> LitvarNode | None:
    """The allele node for exactly this CAID, or `None` when the index holds none."""
    wanted = caid.strip().upper()
    for node in self.autocomplete(caid):
        if node.tier == "clingen" and (node.clingen_id or "").upper() == wanted:
            return node
    return None

LocusRoster dataclass

LocusRoster(
    alleles: dict[str, list[str]] = dict(),
    refs: dict[str, str] = dict(),
    starts: dict[str, int] = dict(),
    caids: dict[str, str] = dict(),
    read: list[str] = list(),
    not_read: dict[str, str] = dict(),
)

The loci a module names, the alleles it names at each, and which tables that came out of.

LocusCoverage dataclass

LocusCoverage(
    rsid: str,
    asked_tier: str,
    tier: str,
    reason: str,
    position_pmids: int | None = None,
    allele_pmids: int | None = None,
    position_only_pmids: int | None = None,
    caids_at_locus: tuple[str, ...] = (),
    caids_with_a_node: tuple[str, ...] = (),
    matched_caids: tuple[str, ...] = (),
    node_id: str | None = None,
)

What the index holds for one module locus, and at which tier it said it.

degraded property

degraded: bool

An allele-level question answered at position level — the finding this pass exists for.

LiteratureCoverageReport dataclass

LiteratureCoverageReport(
    loci: list[LocusCoverage] = list(),
    tables_read: list[str] = list(),
    tables_not_read: dict[str, str] = dict(),
    offline: bool = False,
)

One module's literature coverage, per locus and per tier.

answered property

answered: list[LocusCoverage]

Loci the index gave an answer about — the denominator a coverage number is out of.

position_only_residue property

position_only_residue: int

Papers across the module's loci that sit on a position node and on no allele node.

parse_repr_lines

parse_repr_lines(text: str) -> list[dict]

One dict per line of a Python repr() payload, which is what search/gene really serves.

ast.literal_eval is a literal parser, not eval: it walks the parsed syntax tree and refuses anything that is not a literal container, so it cannot call, import or execute. Calling .json() on this payload raises, and calling eval on it would be a remote-code path.

A line that will not parse is a defect in the response rather than in one record, so it raises; silently dropping it would make a short answer indistinguishable from a complete one.

Source code in enricher/src/just_dna_enricher/litvar.py
def parse_repr_lines(text: str) -> list[dict]:
    """One dict per line of a Python `repr()` payload, which is what `search/gene` really serves.

    `ast.literal_eval` is a **literal parser**, not `eval`: it walks the parsed syntax tree and refuses
    anything that is not a literal container, so it cannot call, import or execute. Calling `.json()`
    on this payload raises, and calling `eval` on it would be a remote-code path.

    A line that will not parse is a defect in the response rather than in one record, so it raises;
    silently dropping it would make a short answer indistinguishable from a complete one.
    """
    rows: list[dict] = []
    for number, line in enumerate(text.splitlines(), start=1):
        stripped = line.strip()
        if not stripped:
            continue
        try:
            value = ast.literal_eval(stripped)
        except (ValueError, SyntaxError) as exc:
            raise LitvarError(
                f"line {number} of the gene search response is not a Python literal ({exc})"
            ) from exc
        if not isinstance(value, dict):
            raise LitvarError(f"line {number} of the gene search response is not a dict")
        rows.append(value)
    return rows

node_tier

node_tier(record: dict) -> str

Which tier a node record sits at, from whichever evidence the endpoint supplied.

The flag_rsid_variant / flag_clingen_variant / flag_gene_variant booleans say it directly and are present on autocomplete and get — but search/gene omits all three, carrying only the keys a record has. So both routes are read, flags first: an id is parsed only where nothing states the answer, because parsing an id is exactly how the first pass at this got it wrong.

Source code in enricher/src/just_dna_enricher/litvar.py
def node_tier(record: dict) -> str:
    """Which tier a node record sits at, from whichever evidence the endpoint supplied.

    The `flag_rsid_variant` / `flag_clingen_variant` / `flag_gene_variant` booleans say it directly and
    are present on `autocomplete` and `get` — but `search/gene` omits all three, carrying only the
    keys a record has. So both routes are read, flags first: an id is parsed only where nothing states
    the answer, because parsing an id is exactly how the first pass at this got it wrong.
    """
    if record.get("flag_clingen_variant"):
        return "clingen"
    if record.get("flag_rsid_variant"):
        return "rsid"
    if record.get("flag_gene_variant"):
        return "gene"
    if record.get("clingen_id"):
        return "clingen"
    if record.get("rsid"):
        return "rsid"
    # Everything left is `litvar@#<gene_id>#…`. The gene node leaves the last slot empty; anything in
    # it is a protein string a miner normalized nothing about.
    return "gene" if str(record.get("_id", "")).endswith("#") else "mention"

rsid_bearing_tables

rsid_bearing_tables() -> dict[str, type[AuthoredModel]]

{filename: model} for every authored table whose model declares rsid.

Derived from DRAFTABLE rather than restated (@registry-completeness), the same walk identifiers._id_bearing_tables does for its own columns — a table kind added later joins this roster by existing, not by somebody remembering it.

Source code in enricher/src/just_dna_enricher/litvar.py
def rsid_bearing_tables() -> dict[str, type[AuthoredModel]]:
    """`{filename: model}` for every authored table whose model declares `rsid`.

    Derived from `DRAFTABLE` rather than restated (`@registry-completeness`), the same walk
    `identifiers._id_bearing_tables` does for its own columns — a table kind added later joins this
    roster by existing, not by somebody remembering it.
    """
    return {
        name: model
        for name, model in DRAFTABLE.items()
        if isinstance(model, type) and issubclass(model, AuthoredModel) and "rsid" in model.model_fields
    }

module_loci

module_loci(spec_dir: Path) -> LocusRoster

Every rsID a module names, with the alleles it names there and any CAID it already holds.

The allele columns come from the model (AuthoredModel.ALLELE_COLUMNS), never from a list here: variants.csv states an allele in alts, genotype and effect_allele, haplotypes.csv in allele, diplotypes.csv in genotype, and a hand-kept list would be the roster defect one layer down. The ref column is skipped because it states the locus rather than a claim about an allele — but a genotype cell names the reference base too, so the bag is the alleles the module writes here, not the alternates. That is deliberate and the REF comparison in _matching_caids is what keeps it honest: a registry allele whose reference base disagrees with the row's is refused. Where no table at a locus states a ref there is nothing to refuse it with, which is a real if narrow way for a reference base to match an alternate.

resolution.csv is read too, for its alts and caid columns — the CAID is the allele identity RM153 puts there, and it is the shortest route from a module row to an allele node.

Source code in enricher/src/just_dna_enricher/litvar.py
def module_loci(spec_dir: Path) -> LocusRoster:
    """Every rsID a module names, with the alleles it names there and any CAID it already holds.

    **The allele columns come from the model** (`AuthoredModel.ALLELE_COLUMNS`), never from a list
    here: `variants.csv` states an allele in `alts`, `genotype` and `effect_allele`, `haplotypes.csv`
    in `allele`, `diplotypes.csv` in `genotype`, and a hand-kept list would be the roster defect one
    layer down. The `ref` **column** is skipped because it states the locus rather than a claim about
    an allele — but a `genotype` cell names the reference base too, so the bag is *the alleles the
    module writes here*, not *the alternates*. That is deliberate and the REF comparison in
    `_matching_caids` is what keeps it honest: a registry allele whose reference base disagrees with
    the row's is refused. Where no table at a locus states a `ref` there is nothing to refuse it with,
    which is a real if narrow way for a reference base to match an alternate.

    `resolution.csv` is read too, for its `alts` and `caid` columns — the CAID is the allele identity
    RM153 puts there, and it is the shortest route from a module row to an allele node.
    """
    roster = LocusRoster()
    for name, model in sorted(rsid_bearing_tables().items()):
        rows, why = _read_table(spec_dir, name, model)
        if why is not None:
            roster.not_read[name] = why
            continue
        roster.read.append(name)
        for row in rows:
            for rsid in _split_cell(getattr(row, "rsid", None)):
                bag = roster.alleles.setdefault(rsid, [])
                for column in getattr(type(row), "ALLELE_COLUMNS", ()):
                    if column == "ref":
                        continue
                    for allele in _split_alleles(getattr(row, column, None)):
                        if allele not in bag:
                            bag.append(allele)
                ref = getattr(row, "ref", None)
                if ref:
                    roster.refs.setdefault(rsid, str(ref).strip().upper())
                start = getattr(row, "start", None)
                if isinstance(start, int):
                    roster.starts.setdefault(rsid, start)
    rows, why = _read_table(spec_dir, RESOLUTION_CSV, ResolutionRow)
    if why is not None:
        roster.not_read[RESOLUTION_CSV] = why
    else:
        roster.read.append(RESOLUTION_CSV)
        # An **input** read, so the author's overlay applies (`@overlay-read-at-inputs-never-at-
        # baselines`): a corrected `alts` or `caid` in `overrides.csv` is the cell this pass should
        # be joining on, and reading the uncorrected one would have it ask about the wrong allele.
        rows = overlaid_input_rows(spec_dir, RESOLUTION_CSV, rows, error=ValueError)
        for row in rows:
            rsid = (getattr(row, "rsid", None) or "").strip()
            if not rsid or rsid not in roster.alleles:
                # A resolution row for a subject no authored table names is a coordinate-keyed row;
                # this pass is keyed on rsIDs, so it has nothing to add there.
                continue
            for allele in _split_alleles(getattr(row, "alts", None)):
                if allele not in roster.alleles[rsid]:
                    roster.alleles[rsid].append(allele)
            caid = (getattr(row, "caid", None) or "").strip()
            if caid:
                roster.caids.setdefault(rsid, caid.upper())
    return roster

coverage_reason

coverage_reason(coverage: LocusCoverage) -> str

Why this locus landed on its tier, one sentence per arm.

A verdict function with several arms owes a reason function with the same arms, pairwise distinct (@answered-is-not-absent) — otherwise two situations a reader must tell apart share a sentence, which is how a strand-flipped SNV was once reported as an event-size disagreement.

Source code in enricher/src/just_dna_enricher/litvar.py
def coverage_reason(coverage: LocusCoverage) -> str:
    """Why this locus landed on its tier, one sentence per arm.

    A verdict function with several arms owes a reason function with the same arms, pairwise distinct
    (`@answered-is-not-absent`) — otherwise two situations a reader must tell apart share a sentence,
    which is how a strand-flipped SNV was once reported as an event-size disagreement.
    """
    rsid = coverage.rsid
    # The ids the index really holds a node for — never the position node's whole `clingen_ids` list.
    # A sentence that says "the index holds allele nodes (CA1, CA2)" about a CAID it holds nothing for
    # is a claim about a lookup that never happened, which is the defect this function exists against.
    answerable = ", ".join(coverage.caids_with_a_node) or "none"
    listed = ", ".join(coverage.caids_at_locus) or "none"
    return {
        "allele_node_matched": (
            f"{rsid}: the index holds an allele node for the allele this module names "
            f"({', '.join(coverage.matched_caids)}), so the answer is allele-resolved."
        ),
        "row_names_no_allele": (
            f"{rsid}: no row names an allele at this locus, so the question is position-level and "
            f"the position node answers it exactly."
        ),
        "no_allele_node_at_locus": (
            f"{rsid}: the index holds a position node and no allele node at all, so the literature "
            f"here is not allele-resolved by the source. The count is position-level."
        ),
        "allele_nodes_name_other_alleles": (
            f"{rsid}: the index holds allele nodes ({answerable}) and none of them is the allele this "
            f"module names, so the allele-level answer is withheld and the count is position-level."
        ),
        "no_node_for_rsid": (
            f"{rsid}: the index holds no node for this rsID at any tier — an answered absence, not a "
            f"failed lookup."
        ),
        "offline": f"{rsid}: the run was offline, so the index was never asked.",
        "index_unreachable": (
            f"{rsid}: LitVar could not be consulted — unreachable, or an answer this client could "
            f"not read — so nothing is established about this locus."
        ),
        "registry_unreachable": (
            f"{rsid}: the ClinGen Allele Registry could not be reached for the allele nodes at this "
            f"locus ({answerable}), so whether one of them is this module's allele was never asked."
        ),
        "allele_not_comparable": (
            f"{rsid}: the registry answered for the allele nodes at this locus ({answerable}) and "
            f"holds nothing these columns can compare against the allele this module names, so the "
            f"tier is undecided rather than settled. The position node lists {listed}."
        ),
    }[coverage.reason]

check_literature_coverage

check_literature_coverage(
    spec_dir: Path,
    *,
    client: LitvarClient | None = None,
    registry: ClingenAlleleClient | None = None,
    offline: bool = False,
    progress: Callable[[int, int], None] | None = None,
) -> LiteratureCoverageReport

Ask LitVar what it holds for each of a module's loci, and at which tier.

Reports and repairs nothing (@enrichment-is-validation). Nothing here is written into a module: a PMID list per variant is not a table kind, literature.csv is keyed by article, and 32,095 PMIDs for one BRAF locus would be a row-writer arguing against itself. What lands is the attestation and this report.

progress is called with (done, total) over loci, the unit @progress-unit-is-subjects fixes, because the total has to be known before the first request goes out.

Source code in enricher/src/just_dna_enricher/litvar.py
def check_literature_coverage(
    spec_dir: Path,
    *,
    client: LitvarClient | None = None,
    registry: ClingenAlleleClient | None = None,
    offline: bool = False,
    progress: Callable[[int, int], None] | None = None,
) -> LiteratureCoverageReport:
    """Ask LitVar what it holds for each of a module's loci, and at which tier.

    Reports and repairs nothing (`@enrichment-is-validation`). Nothing here is written into a module:
    a PMID list per variant is not a table kind, `literature.csv` is keyed by article, and 32,095
    PMIDs for one BRAF locus would be a row-writer arguing against itself. What lands is the
    attestation and this report.

    `progress` is called with `(done, total)` over **loci**, the unit `@progress-unit-is-subjects`
    fixes, because the total has to be known before the first request goes out.
    """
    spec_dir = Path(spec_dir)
    roster = module_loci(spec_dir)
    report = LiteratureCoverageReport(
        tables_read=roster.read, tables_not_read=roster.not_read, offline=offline
    )
    rsids = dedupe(roster.rsids)
    total = len(rsids)
    if offline:
        report.loci = [
            LocusCoverage(rsid=rsid, asked_tier=_asked_tier(roster, rsid), tier="unchecked", reason="offline")
            for rsid in rsids
        ]
        return report
    index = client or LitvarClient()
    allele_registry = registry or ClingenAlleleClient()
    for done, rsid in enumerate(rsids, start=1):
        report.loci.append(_one_locus(index, allele_registry, roster, rsid))
        if progress is not None:
            progress(done, total)
    return report

verification_records

verification_records(
    report: LiteratureCoverageReport,
) -> list[VerificationRecord]

The attestation, and it names the tier — a coverage answer that does not is the defect.

One record. subjects is the loci the index answered about, findings the loci where an allele-level question came back position-level, and detail carries the whole tier breakdown, because "12 checked, 3 flagged" says nothing about which of 328 and 3,945 a reader is holding.

Source code in enricher/src/just_dna_enricher/litvar.py
def verification_records(report: LiteratureCoverageReport) -> list[VerificationRecord]:
    """The attestation, and it names the tier — a coverage answer that does not is the defect.

    One record. `subjects` is the loci the index **answered** about, `findings` the loci where an
    allele-level question came back position-level, and `detail` carries the whole tier breakdown,
    because "12 checked, 3 flagged" says nothing about which of 328 and 3,945 a reader is holding.
    """
    if not report.loci:
        return [
            skipped(
                "literature_coverage",
                "nothing_to_check",
                detail=("no authored table names an rsID, so there was no locus to ask LitVar about"),
                source=LITVAR_SOURCE,
            )
        ]
    if report.offline:
        return [
            skipped(
                "literature_coverage",
                "offline",
                detail=f"--offline, so none of the {len(report.loci)} locus/loci was looked up",
                source=LITVAR_SOURCE,
            )
        ]
    answered = report.answered
    if not answered:
        return [
            skipped(
                "literature_coverage",
                "unreachable",
                detail=(
                    f"none of the {len(report.loci)} locus/loci could be looked up: "
                    f"{examples([locus.rsid for locus in report.loci])}"
                ),
                source=LITVAR_SOURCE,
            )
        ]
    allele = report.at("allele")
    position = report.at("position")
    absent = report.at("absent")
    unchecked = report.at("unchecked")
    degraded = report.degraded
    detail = (
        f"{len(allele)} locus/loci answered at allele tier, {len(position)} at position tier only, "
        f"{len(absent)} absent from the index"
        + (f", {len(unchecked)} could not be asked" if unchecked else "")
        + f"; {report.position_only_residue} paper(s) sit on a position node that no allele node "
        f"claims"
        + (
            f"; allele-level questions answered position-level: "
            f"{examples([locus.rsid for locus in degraded])}"
            if degraded
            else ""
        )
    )
    return [
        ran(
            "literature_coverage",
            subjects=len(answered),
            findings=len(degraded),
            source=LITVAR_SOURCE,
            detail=detail,
        )
    ]