Skip to content

just_dna_enricher.literature

just_dna_enricher.literature

enrich-literature — the fourth pass: citations in, literature.csv out.

Closes the reachable part of the compiler's "does the cited study support the row?" blind spot. Three questions, in increasing ambition and decreasing coverage:

  1. Does the citation exist? PubMed esummary, batched. A nonexistent PMID comes back as a record carrying an error key, so this is a clean yes/no.
  2. Do the identifiers agree? The DOI and PMCID arrive free in the same response. An absent authored DOI is filled here in the sidecar; an authored DOI that contradicts the registry is a finding. Neither is ever written back into studies.csv — the enricher does not edit authored files, because content_signature is defined as reference-independent, and a network fetch that could change it would make that documented property false.
  3. Does the quoted passage appear in the article? Only for the open-access subset, and honestly labelled as such.

All three are attested (RM45): the pass writes citation_existence, citation_identifier and provenance_quote into verification.json on its way out, including on the offline return, so a module can say which of the three was put and over how many citations. Until 0.6 the answers reached a log line and the result object and died there, which left a module whose citations had been checked indistinguishable from one where the command was never run.

Coverage is partial by nature, and saying so is part of the check. A pass that reported "0 quotes found" for an article it could not read would be describing its own reach as if it were a property of the module. So the result separates checked and not found from never retrievable — quotes_found is null rather than zero when no fulltext could be read — and, since dogfooding caught the conflation, also from nothing to check: a citation with no authored quote asked no question and is not counted against coverage. The same distinctions are carried into manifest.literature.

Corrections to the drafted plan, made under probing rather than assumed:

  • The PMC ID converter is not used, though the plan budgeted for it. esummary already returns both doi and pmc in articleids, and Europe PMC's search returns doi/pmcid too — so the converter is a third request for data already in hand. Worse, it answers a different question: for PMID 12345678 (a real, indexed PubMed record) it replies status: error, "Identifier not found in PMC", which is about PMC membership, not existence. Wiring it in as an existence check would report every paywalled article as a broken citation.
  • Europe PMC is not an existence oracle either. Asked for three ids where one does not exist, it returns two results and simply omits the third — no error, no marker. Absence there is indistinguishable from "not indexed", so PubMed decides existence and Europe PMC only decides retrievability.

On running provenance_regex here. The charter requires a linear-time / ReDoS-safe engine for pattern matching, written when the match was specified as consumer-side — arbitrary patterns meeting arbitrary documents. Here the pattern comes from the module being enriched and the document from a public archive, on the author's own machine, so the threat model is a curator writing a slow pattern by accident rather than an attacker. That is worth a bound rather than a compiled dependency — so the match runs under a wall-clock timeout, and a timeout is recorded as not checked.

That bound is enforced with a child process, not a thread, and the reason is worth stating because the thread version looks correct: re cannot be interrupted, threads cannot be killed, and the interpreter joins pool threads at exit — so a thread-based timeout returns on schedule and then hangs the process on the way out. See regex_matches.

LiteratureEnrichmentError

Bases: RuntimeError

Raised in strict mode when a citation does not resolve, or contradicts its own identifiers.

LiteratureUnavailable

Bases: LiteratureEnrichmentError

PubMed could not be reached, so no citation question was put at all (RM101).

A subclass rather than a second exception, so every existing except LiteratureEnrichmentError still catches it (P3 — additive within a major). It separates the source was asked and never answered from a citation that genuinely does not resolve under strict. Only this one means nothing was established either way.

Before RM101 an EutilsError travelled straight out of enrich_literature through a try/finally with no except, so a caller's except LiteratureEnrichmentError was silent for exactly the failure it was written for.

Scoped to the eutils leg on purpose. EuropePmcClient.fulltext and CrossrefClient.exists already answer a transport failure with None rather than an exception — the tri-state withhold this codebase uses for "could not be retrieved" — and turning either into an error here would convert a withheld answer into a failed run.

DoiConflict dataclass

DoiConflict(pmid: str, authored: str, registry: str)

An authored DOI that disagrees with the one the registry reports for the same PMID.

PmcidConflict dataclass

PmcidConflict(pmid: str, authored: str, registry: str)

An authored PMC id that disagrees with the one PubMed reports for the same PMID (RM50).

The DoiConflict shape, for the other cross-registry identifier. It costs no request — the PMC id is already in the esummary articleids block that answered existence — and it catches the case the schema guard cannot see: a cell like 21551363 (PMC3110567) carries a real PubMed id, so nothing refuses it, while the two halves name different articles.

LiteratureResult dataclass

LiteratureResult(
    rows: list[LiteratureRow],
    missing: list[str] = list(),
    doi_conflicts: list[DoiConflict] = list(),
    pmcid_conflicts: list[PmcidConflict] = list(),
    cited: list[str] = list(),
    existence_checked: int = 0,
    unresolved_citations: int = 0,
    doi_verdicts_stale: int = 0,
    doi_never_checked: int = 0,
    noncommercial_quoted: list[str] = list(),
    titles_as_quotes: list[str] = list(),
    fulltext_requested: bool = True,
    quotes_authored: int = 0,
    quotes_found: int = 0,
    quotes_checked: int = 0,
    quotes_unchecked: int = 0,
    quotes_unexamined: int = 0,
    fulltext_checked: list[str] = list(),
    abstract_checked: list[str] = list(),
    doi_missing: list[str] = list(),
    identifiers_authored: int = 0,
    identifiers_compared: int = 0,
    identifiers_conflicting: int = 0,
    identifiers_unmatched: int = 0,
    identifiers_foreign: int = 0,
    sources: list[str] = list(),
    mode: str = "best_effort",
    skipped_offline: bool = False,
)

What the pass found, and — through subject_rows — what it found it about.

One subject set, read by the report, by the strict gates and by the attestation. The set is the citations the module makes now, answered by the rows literature.csv holds now, and it is the single rule the whole pass turns on. Three separate defects came from parts of this file reading three different sets:

  • the strict gates read lists appended inside the fetch loop, so a module enriched once with --best-effort and then with --strict was blessed on a citation PubMed has no record of — the refusal depended on the order the two runs happened in, and the attestation written by the same run said findings=1 while the gate beside it said nothing was wrong;
  • the identifier cross-check ran inside that loop too, so an existing literature.csv hid every DOI/PMC id disagreement;
  • the counts read rows, which is everything the sidecar carries and never shrinks, so a citation deleted from studies.csv went on being counted and no authored edit could clear the finding it produced.

Merge-not-clobber is what makes the distinction real: the sidecar keeps rows for citations the module has since dropped, and those rows are still written back (deleting a curator's row is not this pass's call) — they are simply not what any of this run's questions were about.

subject_rows property

subject_rows: list[LiteratureRow]

The rows this run's questions were about: a row for a citation the module still makes.

Not rows, which is the whole sidecar — see the class docstring. Not this run's fetch list either: literature.csv is the pin, so a row merged from an earlier run carries a verdict that still stands, and counting only what this run looked up would let a no-op re-run replace a true attestation with subjects=0, which reads as "the check ran and had nothing in scope".

coverage property

coverage: str

The sentence this pass exists to be able to say honestly.

The denominator is what had something to check — a citation with no authored quote was not skipped for lack of a fulltext, it simply asked no question. Counting it as unretrievable was a real bug found by running this against reference_examples/pathogenic_clinvar/, whose single citation is open access and carries no quote: the old wording claimed its fulltext could not be retrieved, which was the opposite of true.

Counted in quotes, off the tally, and not in citations off the run's own fetch lists. The earlier wording read fulltext_checked, which only ever holds PMIDs this run retrieved, so a no-op re-run over a pinned open-access citation printed "0 of 1 … 1 with nothing retrievable" one line above "1/1 found" — the same false retrievability claim the paragraph above records as a bug, arrived at from the other direction. Quotes also partition cleanly where citations do not: one citation can carry a settled quote and a never-examined one at once, so a citation-granular sentence has to put it in one bucket and be wrong about the other.

EuropePmcClient dataclass

EuropePmcClient(
    base_url: str = DEFAULT_EUROPEPMC_BASE,
    batch_size: int = 25,
    min_request_interval: float = 0.5,
    timeout: float = 30.0,
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

Open-access lookup and fulltext retrieval. Paced like every other client in this tier.

lookup

lookup(pmids: list[str]) -> dict[str, dict]

pmid -> {pmcid, doi, is_open_access, license, abstract} for the ids Europe PMC knows.

Ids it does not know are simply absent from the result, with no error marker — which is why this is not an existence check. A caller must read a miss here as "not retrievable", never as "does not exist".

The abstract comes back for paywalled records too, in this same response, and that is worth more than it looks: probed across a mix of non-open-access papers, four of five carried one (only a 1994 non-research document did not). It is the difference between checking a quote for the open-access minority and checking it for nearly everything.

So does the article's license (RM46), which is why per-article terms cost no extra request. Probed 2026-08-13 over 100 records: the values are lowercase CC spellings — cc by (64), cc by-nc (28), cc by-nc-nd (8). It is independent of isOpenAccess and must not be derived from it: PMID 28546431 comes back isOpenAccess: N with license: cc by, since the flag describes Europe PMC's OA subset while the licence describes the article. Stored verbatim; licensing.article_terms maps it to rights at read time.

Source code in enricher/src/just_dna_enricher/literature.py
def lookup(self, pmids: list[str]) -> dict[str, dict]:
    """`pmid -> {pmcid, doi, is_open_access, license, abstract}` for the ids Europe PMC knows.

    Ids it does not know are simply **absent from the result**, with no error marker — which is why
    this is not an existence check. A caller must read a miss here as "not retrievable", never as
    "does not exist".

    **The abstract comes back for paywalled records too**, in this same response, and that is worth
    more than it looks: probed across a mix of non-open-access papers, four of five carried one
    (only a 1994 non-research document did not). It is the difference between checking a quote for
    the open-access minority and checking it for nearly everything.

    **So does the article's `license`** (RM46), which is why per-article terms cost no extra
    request. Probed 2026-08-13 over 100 records: the values are lowercase CC spellings — `cc by`
    (64), `cc by-nc` (28), `cc by-nc-nd` (8). It is **independent of `isOpenAccess`** and must not
    be derived from it: PMID 28546431 comes back `isOpenAccess: N` with `license: cc by`, since
    the flag describes Europe PMC's OA subset while the licence describes the article. Stored
    verbatim; `licensing.article_terms` maps it to rights at read time.
    """
    out: dict[str, dict] = {}
    for batch in batched(dedupe(pmids), self.batch_size):
        query = " OR ".join(f"EXT_ID:{pmid}" for pmid in batch)
        # **Translate, do not leak** (`@client-exception-contract`, RM230). `_get` retries the
        # transport legs and re-raises what survives; this method used to call it bare and then
        # `.json()` the result, so all three failure legs escaped as their httpx/json types:
        # a persistent 503 as `httpx.HTTPStatusError`, a refused connection as
        # `httpx.ConnectError`, and a 200 that is not JSON as `json.JSONDecodeError`. The last is
        # the exact fourth leg `test_client_exception_contract.py` was written for, and every
        # sibling client in this tier closes it in its own module.
        #
        # It mattered where it landed: `enrich_literature` calls this inside a `try:` whose only
        # companion is `finally:` — no `except` — which is verbatim the shape
        # `test_pass_exception_contract.py` exists to refuse. The pass-level suite never drove
        # this leg because its stub raises earlier, on the eutils call.
        #
        # `fulltext` below is deliberately NOT changed: it catches httpx and returns `None`, the
        # tri-state withhold this tier uses for "could not be retrieved". The two methods answer
        # different questions, which is why the class's contract-suite exemption named `fulltext`
        # and silently covered `lookup` too.
        try:
            response = self._get("search", {"query": query, "resultType": "core", "format": "json"})
            payload = response.json()
        except httpx.HTTPError as exc:
            raise LiteratureUnavailable(f"Europe PMC could not be asked: {exc}") from exc
        except ValueError as exc:
            raise LiteratureUnavailable(
                f"Europe PMC answered {response.status_code} with a body that is not JSON: {exc}"
            ) from exc
        for record in (payload.get("resultList") or {}).get("result") or []:
            pmid = str(record.get("pmid") or "")
            if not pmid:
                continue
            out[pmid] = {
                "pmcid": record.get("pmcid"),
                "doi": record.get("doi"),
                # The API answers with the strings 'Y'/'N', not booleans.
                "is_open_access": str(record.get("isOpenAccess") or "").upper() == "Y",
                # `inEPMC` is NOT a fulltext signal: a record can be in Europe PMC with
                # `isOpenAccess=N`, and `fullTextXML` then answers 404. Probed on PMID 23788249.
                "in_epmc": str(record.get("inEPMC") or "").upper() == "Y",
                # Verbatim, and `None` when the field is absent — an article whose terms Europe
                # PMC does not state is unknown, not unlicensed.
                "license": (record.get("license") or None),
                "abstract": record.get("abstractText"),
            }
    return out

fulltext

fulltext(pmcid: str) -> str | None

Whitespace-normalized article text, or None when it cannot be retrieved.

None is a normal outcome, not an error: an embargoed or author-manuscript-only record answers 404, and the caller must record that as unchecked rather than as a failed match.

Source code in enricher/src/just_dna_enricher/literature.py
def fulltext(self, pmcid: str) -> str | None:
    """Whitespace-normalized article text, or `None` when it cannot be retrieved.

    `None` is a normal outcome, not an error: an embargoed or author-manuscript-only record answers
    404, and the caller must record that as *unchecked* rather than as a failed match.
    """
    try:
        response = self._get(f"{pmcid}/fullTextXML")
    except httpx.HTTPStatusError as exc:
        logger.info("No retrievable fulltext for %s (HTTP %s)", pmcid, exc.response.status_code)
        return None
    except httpx.HTTPError as exc:
        logger.warning("Fulltext fetch failed for %s (%s)", pmcid, exc)
        return None
    return extract_text(response.text)

CrossrefClient dataclass

CrossrefClient(
    base_url: str = DEFAULT_CROSSREF_BASE,
    min_request_interval: float = 0.1,
    timeout: float = 30.0,
    contact_email: str | None = None,
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

DOI existence, for the citations PubMed does not index.

PubMed answers "does this article exist" for anything it indexes — including paywalled work; a paywall governs the fulltext, not the record. What it cannot answer for is everything outside its scope: preprints, books, theses, datasets, standards. Those have DOIs and no PMID, and Crossref is the registry that mints and resolves them (a probed bioRxiv preprint, 10.1101/2024.06.17.599351, returns type: posted-content; a fabricated DOI returns a clean 404).

This is also what makes the 1.0 doi-first flip low-risk: once pmid becomes optional and a citation may carry only a DOI, existence checking has to work without PubMed, and it already does.

Crossref asks callers to identify themselves in the User-Agent and gives the "polite pool" in return; the contact address is sent only when one is configured, for the same reason eutils omits it rather than inventing one.

exists

exists(doi: str) -> bool | None

True/False, or None when Crossref could not be asked.

None rather than False on a transport failure or an unexpected status: "we could not check" and "this DOI does not exist" are different claims, and only the second is a finding against the module. The translation stays here and the retrying stays in _request above — reraise=True means the last failure arrives back at this except after the attempts are spent, so the three-valued contract is unchanged and only the number of tries moved.

Source code in enricher/src/just_dna_enricher/literature.py
def exists(self, doi: str) -> bool | None:
    """`True`/`False`, or `None` when Crossref could not be asked.

    `None` rather than `False` on a transport failure or an unexpected status: "we could not
    check" and "this DOI does not exist" are different claims, and only the second is a finding
    against the module. The translation stays here and the retrying stays in `_request` above —
    `reraise=True` means the last failure arrives back at this `except` after the attempts are
    spent, so the three-valued contract is unchanged and only the number of tries moved.
    """
    try:
        response = self._request(doi)
    except httpx.HTTPError as exc:
        logger.warning("Crossref lookup failed for %s (%s); not checked", doi, exc)
        return None
    if response.status_code == 404:
        return False
    if response.status_code != 200:
        logger.warning("Crossref answered HTTP %s for %s; not checked", response.status_code, doi)
        return None
    return True

PmcIdRecord dataclass

PmcIdRecord(
    pmcid: str,
    pmid: str | None = None,
    in_pmc: bool = True,
    error: str | None = None,
)

One answer from the PMC id converter. in_pmc=False is PMC saying it has no such record; an id the request never reached is absent from the result entirely, never one of these.

PmcIdConverterClient dataclass

PmcIdConverterClient(
    base_url: str = DEFAULT_PMC_IDCONV_URL,
    batch_size: int = 200,
    min_request_interval: float = 0.5,
    timeout: float = 30.0,
    contact_email: str | None = None,
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

PMCID → PMID, for a curator holding the wrong half of the pair (RM50).

It reports; it never fills. The PMID it returns is handed back as an advisory the author types themselves, because filling StudyRow.pmid from NCBI would make literature.exists compare NCBI against NCBI — the hints.REDUNDANCY_BEARING rule, already argued for doi (Crossref is asked about the authored DOI, since a derived one exists by construction).

Why this exists although the pass docstring says the converter is unused. That statement is true of the direction the pass needs — PMID → PMCID arrives free in the esummary articleids block, so calling the converter for it would be a third request for data already in hand, and it answers a different question (asked about a paywalled PMID it replies "Identifier not found in PMC", which is about PMC membership rather than existence). None of that says anything about PMCID → PMID, which is the direction the converter is actually for and the one a curator has no other route to.

Probed 2026-08-13: the www.ncbi.nlm.nih.gov/pmc/utils/idconv/v1.0/ path 301-redirects to the address below, an absent id comes back as a record with status: "error" and errmsg: "Identifier not found in PMC", and pmid arrives as a JSON number.

resolve

resolve(pmcids: list[str]) -> dict[str, PmcIdRecord]

PMCID -> PmcIdRecord for every id the converter answered about.

Four outcomes, and none of them is spelled the same way as another. An id it resolved carries a pmid; an id it knows with no PubMed id carries in_pmc=True, pmid=None (a real answer — some PMC records genuinely have none); an id it has no record of carries in_pmc=False with the service's own errmsg; and an id the request never reached is simply absent from the result. A caller must not read a missing key as "does not exist" — that conflation is exactly what S20 was about, one service over.

Source code in enricher/src/just_dna_enricher/literature.py
def resolve(self, pmcids: list[str]) -> dict[str, PmcIdRecord]:
    """`PMCID -> PmcIdRecord` for every id the converter answered about.

    **Four outcomes, and none of them is spelled the same way as another.** An id it resolved
    carries a `pmid`; an id it knows with no PubMed id carries `in_pmc=True, pmid=None` (a real
    answer — some PMC records genuinely have none); an id it has no record of carries
    `in_pmc=False` with the service's own `errmsg`; and an id the request never reached is simply
    **absent from the result**. A caller must not read a missing key as "does not exist" — that
    conflation is exactly what S20 was about, one service over.
    """
    out: dict[str, PmcIdRecord] = {}
    for batch in batched(dedupe(pmcids), self.batch_size):
        params = {"ids": ",".join(batch), "format": "json", "tool": "just-dna-enricher"}
        if self.contact_email:
            params["email"] = self.contact_email
        try:
            payload = self._get(params).json()
        except httpx.HTTPError as exc:
            logger.warning("PMC id converter could not be asked for %s (%s)", batch, exc)
            continue
        for record in payload.get("records") or []:
            requested = str(record.get("requested-id") or record.get("pmcid") or "")
            if not requested:
                continue
            pmid = record.get("pmid")
            out[requested] = PmcIdRecord(
                pmcid=requested,
                pmid=str(pmid) if pmid is not None else None,
                in_pmc=record.get("status") != "error",
                error=record.get("errmsg"),
            )
    return out

extract_text

extract_text(xml: str) -> str | None

JATS XML → one normalized string, or None if it does not parse.

itertext() rather than a tag-stripping regex: real JATS carries nested inline markup inside sentences (<italic>, <xref>, <sup>), and a regex that removes tags without joining their text would silently glue words together and break matches that should succeed.

Source code in enricher/src/just_dna_enricher/literature.py
def extract_text(xml: str) -> str | None:
    """JATS XML → one normalized string, or `None` if it does not parse.

    `itertext()` rather than a tag-stripping regex: real JATS carries nested inline markup inside
    sentences (`<italic>`, `<xref>`, `<sup>`), and a regex that removes tags without joining their
    text would silently glue words together and break matches that should succeed.
    """
    try:
        root = ET.fromstring(xml)
    except ET.ParseError as exc:
        logger.warning("Fulltext did not parse as XML (%s)", exc)
        return None
    return _WHITESPACE.sub(" ", " ".join(root.itertext())).strip()

quote_matches

quote_matches(quote: str, fulltext: str) -> bool

Literal, whitespace- and case-insensitive containment. No engine, so nothing to bound.

Source code in enricher/src/just_dna_enricher/literature.py
def quote_matches(quote: str, fulltext: str) -> bool:
    """Literal, whitespace- and case-insensitive containment. No engine, so nothing to bound."""
    return _normalize(quote) in _normalize(fulltext)

regex_matches

regex_matches(
    pattern: str,
    fulltext: str,
    *,
    timeout: float = DEFAULT_REGEX_TIMEOUT,
) -> bool | None

True/False, or None when the match could not be completed in time.

The three-way return is the point: a pattern that runs long has not failed to match, it has failed to be checked, and reporting it as "not found" would send an author to fix a quote that is probably there.

Why a subprocess and not a thread. The obvious implementation — submit to a ThreadPoolExecutor and take future.result(timeout=...) — does not work, and fails in a way that looks like it works: re never releases the GIL to a cancellation, threads cannot be killed, and the interpreter joins non-daemon pool threads at exit. So a runaway pattern returns None on time and then hangs the process on the way out. Verified by writing it that way first and watching the test suite stop. A child process is the only bound in the standard library that can actually be enforced, and it is cheap here: it runs only for a row that has a provenance_regex and a retrievable fulltext, which is a small subset of a small subset.

Source code in enricher/src/just_dna_enricher/literature.py
def regex_matches(pattern: str, fulltext: str, *, timeout: float = DEFAULT_REGEX_TIMEOUT) -> bool | None:
    """`True`/`False`, or **`None` when the match could not be completed in time**.

    The three-way return is the point: a pattern that runs long has not failed to match, it has failed
    to be *checked*, and reporting it as "not found" would send an author to fix a quote that is
    probably there.

    **Why a subprocess and not a thread.** The obvious implementation — submit to a
    `ThreadPoolExecutor` and take `future.result(timeout=...)` — does not work, and fails in a way that
    looks like it works: `re` never releases the GIL to a cancellation, threads cannot be killed, and
    the interpreter joins non-daemon pool threads at exit. So a runaway pattern returns `None` on time
    and then hangs the process on the way out. Verified by writing it that way first and watching the
    test suite stop. A child process is the only bound in the standard library that can actually be
    enforced, and it is cheap here: it runs only for a row that has a `provenance_regex` *and* a
    retrievable fulltext, which is a small subset of a small subset.
    """
    try:
        re.compile(pattern)
    except re.error as exc:
        # The compiler already grammar-checks `provenance_regex`, so this is a corrupt sidecar rather
        # than an authoring slip; abstain rather than claiming the quote is absent.
        logger.warning("provenance_regex %r did not compile (%s); not checked", pattern, exc)
        return None

    sink: multiprocessing.Queue = multiprocessing.Queue()
    worker = multiprocessing.Process(target=_regex_worker, args=(pattern, fulltext, sink))
    worker.daemon = True
    worker.start()
    try:
        worker.join(timeout)
        if worker.is_alive():
            worker.terminate()
            worker.join(1.0)
            logger.warning(
                "provenance_regex %r exceeded %.1fs against a %d-character fulltext; recorded as "
                "NOT CHECKED (never as not-found)",
                pattern,
                timeout,
                len(fulltext),
            )
            return None
        return sink.get_nowait() if not sink.empty() else None
    except Exception as exc:
        logger.warning("provenance_regex %r could not be evaluated (%s); not checked", pattern, exc)
        return None
    finally:
        sink.close()
        if worker.is_alive():
            worker.kill()

enrich_literature

enrich_literature(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    offline: bool = False,
    check_fulltext: bool = True,
    check_doi: bool = True,
    regex_timeout: float = DEFAULT_REGEX_TIMEOUT,
    write: bool = True,
    eutils: EutilsClient | None = None,
    europepmc: EuropePmcClient | None = None,
    crossref: CrossrefClient | None = None,
) -> LiteratureResult

Fill literature.csv from the citations a module makes.

studies.csv is one citation site of several (RM47, RM132): MeasureBinRow.pmid grounds the threshold its row states, and PharmVariantRow.pmid grounds the drug/genotype claim its row makes — a claim studies.csv cannot ground, because a study row attaches to the whole variant. A module whose only citations come from those tables is enriched exactly like one with a studies.csv; reading fewer than all the sites would leave the rest unchecked, which is worse than the honest gap the columns replaced. The kinds come from the compiler's own derived registry through load_citing_rows/table_citations, so a kind that gains the column is read here with no edit to this tier.

Existing rows are authoritative and merged, never clobbered — the same rule enrich() applies to resolution.csv, with the same consequence: to regenerate after a machinery change you must delete the file first. That includes the licence columns (RM46): rows written before 0.6 carry no license, and re-running will not back-fill them, because merge-not-clobber cannot tell an absent value from a curator's deliberate blank. Delete the sidecar to re-derive.

--offline makes this a no-op with a warning. There is no offline literature snapshot and there will not be one; once literature.csv is written it is the pin, and later compiles read it offline and deterministically.

Three checks are attested on the way out (_attest), the offline return included: a run that did not put a question has to say so, or the manifest cannot tell it from a run that put one and found nothing.

Source code in enricher/src/just_dna_enricher/literature.py
def enrich_literature(
    spec_dir: Path,
    *,
    mode: str = "best_effort",
    offline: bool = False,
    check_fulltext: bool = True,
    check_doi: bool = True,
    regex_timeout: float = DEFAULT_REGEX_TIMEOUT,
    write: bool = True,
    eutils: EutilsClient | None = None,
    europepmc: EuropePmcClient | None = None,
    crossref: CrossrefClient | None = None,
) -> LiteratureResult:
    """Fill `literature.csv` from the citations a module makes.

    **`studies.csv` is one citation site of several** (RM47, RM132): `MeasureBinRow.pmid` grounds the
    threshold its row states, and `PharmVariantRow.pmid` grounds the drug/genotype claim its row
    makes — a claim `studies.csv` cannot ground, because a study row attaches to the whole variant. A
    module whose only citations come from those tables is enriched exactly like one with a
    `studies.csv`; reading fewer than all the sites would leave the rest unchecked, which is worse
    than the honest gap the columns replaced. The kinds come from the compiler's own derived registry
    through `load_citing_rows`/`table_citations`, so a kind that gains the column is read here with no
    edit to this tier.

    Existing rows are authoritative and merged, never clobbered — the same rule `enrich()` applies to
    `resolution.csv`, with the same consequence: to regenerate after a machinery change you must delete
    the file first. **That includes the licence columns** (RM46): rows written before 0.6 carry no
    `license`, and re-running will not back-fill them, because merge-not-clobber cannot tell an
    absent value from a curator's deliberate blank. Delete the sidecar to re-derive.

    `--offline` makes this a no-op with a warning. There is no offline literature snapshot and there
    will not be one; once `literature.csv` is written it *is* the pin, and later compiles read it
    offline and deterministically.

    **Three checks are attested on the way out** (`_attest`), the offline return included: a run that
    did not put a question has to say so, or the manifest cannot tell it from a run that put one and
    found nothing.
    """
    spec_dir = Path(spec_dir)
    # `studies.csv` stays a plain join: it is an AUTHORED table and lives in the spec root, and
    # `sidecar_path` is for the machine-written files RM49 allowed under `derived/`.
    studies_path = spec_dir / "studies.csv"
    # Through the resolver, never joined by hand (RM99): a module keeping its sidecars under
    # `derived/` (RM49) has this file there, and a pass with its own literal would read the split copy
    # and write a flat one, leaving the module with both -- the collision RM49 made an error rather
    # than a preference. `@sidecar-name-and-place`: write to the file you read.
    output_path = sidecar_path(spec_dir, "literature.csv", error=LiteratureEnrichmentError)

    studies: list[StudyRow] = []
    if studies_path.exists():
        studies, errors, _ = load_csv_rows(studies_path, StudyRow, "studies.csv")
        if errors:
            raise LiteratureEnrichmentError(f"studies.csv is invalid: {errors[0]}")
    # The table pointers, through the compiler's own loader: importing its private citing-kind tuple
    # or keeping a second list of the kinds here is the RM41 shape, and the copy goes stale the next
    # time a model declares a `pmid`. Its `ValueError` is re-raised as this pass's own error so a bad
    # citing row is diagnosed the way a bad `studies.csv` two lines above already is, rather than
    # tracebacking out of the CLI.
    try:
        citing_rows = load_citing_rows(spec_dir)
    except ValueError as exc:
        raise LiteratureEnrichmentError(f"a citing table is invalid: {exc}") from exc
    table_pmids = table_citations(citing_rows)
    if not studies and not table_pmids:
        raise LiteratureEnrichmentError(
            f"no citations in {spec_dir} — the literature pass checks the citations a module makes, "
            f"and this one has neither studies.csv rows nor a `pmid` on any binning or "
            f"pharm_variants row."
        )

    existing: dict[tuple, LiteratureRow] = {}
    if output_path.exists():
        rows, errors, _ = load_csv_rows(output_path, LiteratureRow, "literature.csv")
        if errors:
            raise LiteratureEnrichmentError(f"existing literature.csv is invalid: {errors[0]}")
        for row in rows:
            existing[merge_key(row)] = row

    citations = _citations(studies, table_pmids)
    authored_total = sum(
        1 for rows in citations.values() for s in rows if s.provenance_quote or s.provenance_regex
    )

    if offline:
        logger.warning(
            "Literature enrichment skipped: --offline and there is no offline PubMed/Europe PMC "
            "snapshot. Any existing literature.csv is kept as the pin; compiles stay reproducible."
        )
        out = sorted(existing.values(), key=lambda r: int(r.pmid))
        if write and existing:
            _write_literature_csv(out, output_path)
        return _attest(
            LiteratureResult(
                rows=out,
                mode=mode,
                skipped_offline=True,
                sources=sorted({r.source for r in out if r.source}),
                cited=sorted(citations, key=int),
                quotes_authored=authored_total,
                fulltext_requested=check_fulltext,
            ),
            spec_dir,
            write=write,
            check_fulltext=check_fulltext,
            check_doi=check_doi,
        )

    # **A pin whose quote count no longer matches `studies.csv` is re-fetched (RM277).** The row's own
    # input changed, so it describes a different set of quotes: its `quotes_found` cannot be paired
    # with the new count (`_tally_quotes` reads exactly that mismatch as "unexamined"), and rewriting
    # the count alone would attach the old verdict to the new quotes. Re-deriving the row is the
    # corrected derivation, and it is what makes the compiler's `quote_counter_stale` remedy, "re-run
    # the literature pass", true. Offline has returned above and keeps the pin.
    stale_pins = {
        row.pmid
        for row in existing.values()
        if row.pmid in citations
        and (row.quotes_authored or 0)
        != sum(1 for s in citations[row.pmid] if s.provenance_quote or s.provenance_regex)
    }
    if stale_pins:
        logger.info(
            "Literature: re-fetching %d pinned citation(s) whose quote count changed in studies.csv: %s",
            len(stale_pins),
            ", ".join(sorted(stale_pins, key=int)),
        )
    # Against the rows rather than the dict's keys, which are merge-key tuples and not bare PMIDs.
    have = {row.pmid for row in existing.values()} - stale_pins
    wanted = [pmid for pmid in citations if pmid not in have]
    fetched_at = now_utc_iso()
    result = LiteratureResult(
        rows=[row for row in existing.values() if row.pmid not in stale_pins],
        mode=mode,
        cited=sorted(citations, key=int),
        quotes_authored=authored_total,
        fulltext_requested=check_fulltext,
    )

    # **Titles are needed for MERGED rows too, which is the correction S54's reporter filed against
    # their own report.** A pinned row is not in `wanted`, so the fetch loop never sees it — and on
    # the four modules that motivated the title check, every row is pinned, so the check would not
    # have fired on a single one of the 3,668 quotes it was written for. `esummary` batches, so
    # adding the already-pinned PMIDs that carry a quote costs no extra round trip in the common
    # case and nothing at all when there are none.
    titled = [pmid for pmid in citations if pmid not in wanted and _has_quote(citations[pmid])]
    if wanted or titled:
        owned_eutils = eutils is None
        client = eutils or EutilsClient()
        try:
            summaries = client.esummary("pubmed", wanted + titled)
        except EutilsError as exc:
            raise LiteratureUnavailable(f"PubMed could not be reached: {exc}") from exc
        finally:
            if owned_eutils:
                client.close()

        owned_epmc = europepmc is None
        epmc = europepmc or EuropePmcClient()
        owned_crossref = crossref is None
        crossref = crossref or CrossrefClient()
        try:
            indexed = epmc.lookup(wanted)
            for pmid in wanted:
                summary = summaries.get(pmid, {})
                exists = not is_missing(summary)
                ids = _identifiers(summary)
                epmc_record = indexed.get(pmid, {})
                doi = ids.get("doi") or epmc_record.get("doi")
                pmcid = ids.get("pmcid") or epmc_record.get("pmcid")
                is_open = epmc_record.get("is_open_access") if epmc_record else None
                # Verbatim from Europe PMC; rights derived at read time so a mapping correction
                # reaches rows already written (`licensing.article_terms`).
                license_name = epmc_record.get("license") if epmc_record else None
                terms = article_terms(license_name)

                # Nothing is tallied in this loop, and that is the rule rather than a preference:
                # every count this pass reports is taken afterwards over `subject_rows`, so the
                # `strict` gates, the CLI report and the attestation cannot end up reading three
                # different sets. A citation already pinned in `literature.csv` is not fetched here,
                # so a count made in this loop is a count of *this run's requests* — which made the
                # existence gate, the identifier cross-check and the record disagree with each other
                # depending on the order two runs happened in.

                # Crossref checks the **authored** DOI in preference to the derived one. Checking the
                # registry's own DOI would be circular — it exists by construction, since the registry
                # just handed it over. The authored cell is the one nobody has verified, and it is the
                # only one that exists at all for a citation PubMed does not index (a preprint, book or
                # dataset), which is the case this whole client is here for.
                target_doi = _doi_to_check(citations[pmid], doi)
                doi_exists: bool | None = None
                if check_doi and target_doi:
                    doi_exists = crossref.exists(target_doi)

                quotes = [s for s in citations[pmid] if s.provenance_quote or s.provenance_regex]
                found: int | None = None
                quote_source: str | None = None
                if check_fulltext and quotes:
                    text = epmc.fulltext(pmcid) if (is_open and pmcid) else None
                    if text is not None:
                        result.fulltext_checked.append(pmid)
                        quote_source = "fulltext"
                    else:
                        # Fall back to the abstract, which Europe PMC serves for paywalled records
                        # too. A HIT here is as conclusive as one in the body; a MISS is not, because
                        # the body was never searched — which is what `quote_source` records, and why
                        # this row still counts as unchecked below.
                        text = epmc_record.get("abstract")
                        if text:
                            result.abstract_checked.append(pmid)
                            quote_source = "abstract"
                    if text:
                        found = sum(
                            1 for s in quotes if _study_quote_found(s, text, regex_timeout=regex_timeout)
                        )
                # Costs no request: the title arrived in the same `esummary` response that answered
                # existence. Outside the `check_fulltext` guard deliberately — this compares the quote
                # against metadata already held, so it answers even for a paywalled article whose
                # fulltext was never retrievable, which is where the defect hides.
                if quotes and _quote_is_the_title(quotes, summary):
                    result.titles_as_quotes.append(pmid)
                # What the retrieved text settled is tallied once, after the loop, over every row —
                # see `_tally_quotes`. The per-row facts it reads (`quotes_authored`, `quotes_found`,
                # `quote_source`) are written right here, so the tally is the same arithmetic applied
                # to merged rows as well as fresh ones.
                result.rows.append(
                    LiteratureRow(
                        pmid=pmid,
                        doi=doi,
                        pmcid=pmcid,
                        exists=exists,
                        is_open_access=is_open,
                        license=license_name,
                        share_alike=terms.share_alike,
                        commercial_use=terms.commercial_use,
                        redistribution=terms.redistribution,
                        quotes_authored=len(quotes),
                        quotes_found=found,
                        quote_source=quote_source,
                        doi_exists=doi_exists,
                        # Which DOI that verdict is about. Without it the pin cannot say, and a
                        # re-run pairs the stored answer with whatever the author writes next.
                        doi_checked=target_doi if doi_exists is not None else None,
                        # PubMed is the row's source: it decides existence and supplies the
                        # identifiers. Europe PMC contributes `is_open_access` and the fulltext, but
                        # it cannot originate a row (it silently omits ids it does not know), so it
                        # is not a `source` in the sense the other sidecars use the word.
                        source="pubmed",
                        status="resolved" if exists else "not_found",
                        fetched_at=fetched_at,
                    )
                )
            # The pinned rows: no row is written for them (the sidecar is authoritative) and the
            # fulltext is never fetched, but the title comparison needs only the summary — so it
            # answers here where `quotes_found` cannot.
            for pmid in titled:
                if _quote_is_the_title(
                    [s for s in citations[pmid] if s.provenance_quote or s.provenance_regex],
                    summaries.get(pmid, {}),
                ):
                    result.titles_as_quotes.append(pmid)
        finally:
            if owned_epmc:
                epmc.close()
            if owned_crossref:
                crossref.close()

    result.titles_as_quotes = sorted(set(result.titles_as_quotes), key=int)
    result.rows.sort(key=lambda r: int(r.pmid))
    result.fulltext_checked = sorted(set(result.fulltext_checked), key=int)
    result.sources = sorted({r.source for r in result.rows if r.source})
    # Every tally runs here, once, over the sorted subject rows — so each list is ordered by PMID
    # rather than by whatever order the fetch took, and the gates below refuse on exactly what the
    # attestation records.
    _tally_existence(result, citations)
    _tally_quotes(result, citations)
    _compare_identifiers(result, citations)
    logger.info("Literature: %s", result.coverage)

    if mode == "strict" and result.missing:
        raise LiteratureEnrichmentError(
            f"strict literature enrichment: PubMed has no record of {len(result.missing)} cited "
            f"PMID(s): {result.missing}. Either the identifier is a typo or the article was pulled "
            f"from the index; both mean the row's grounding evidence does not resolve. Fix the "
            f"citation, or enrich with mode='best_effort' to record it as a warning."
        )
    if mode == "strict" and result.doi_missing:
        raise LiteratureEnrichmentError(
            f"strict literature enrichment: Crossref has no record of {len(result.doi_missing)} "
            f"cited DOI(s): {result.doi_missing}. A DOI that does not resolve is a citation that "
            f"cannot be followed."
        )
    if mode == "strict" and result.doi_conflicts:
        raise LiteratureEnrichmentError(
            f"strict literature enrichment: {len(result.doi_conflicts)} authored DOI(s) disagree with "
            f"the registry: {[str(c) for c in result.doi_conflicts]}. One of the two identifiers "
            f"points at the wrong paper."
        )
    if mode == "strict" and result.pmcid_conflicts:
        raise LiteratureEnrichmentError(
            f"strict literature enrichment: {len(result.pmcid_conflicts)} authored PMC id(s) disagree "
            f"with PubMed's for the same record: {[str(c) for c in result.pmcid_conflicts]}. One of "
            f"the two identifiers points at the wrong paper."
        )
    if write:
        _write_literature_csv(result.rows, output_path)
    return _attest(result, spec_dir, write=write, check_fulltext=check_fulltext, check_doi=check_doi)

bibliographic

bibliographic(summary: dict) -> dict[str, str | None]

Pull the fields that say which paper this is out of an esummary record.

Public, unlike _identifiers, because two tiers need it and the alternative is a consumer re-implementing a parse of a payload we already hold — the RM41 lesson. lookup.CitationHint reads it so "does this PMID exist" can become "does this PMID name the paper you meant": existence alone cannot catch a fabricated citation, because PMIDs are densely allocated and an invented number is usually a real record for a different article.

Every value is None when the field is absent rather than empty-string, so a caller can tell "PubMed did not say" from "PubMed said nothing is there" — the house tri-state, applied to metadata. year is the leading four digits of pubdate (which is free-form: 2017 Nov 20, 2017, 2017 Nov-Dec), and nothing is invented when it does not start with a year.

Source code in enricher/src/just_dna_enricher/literature.py
def bibliographic(summary: dict) -> dict[str, str | None]:
    """Pull the fields that say *which paper this is* out of an esummary record.

    Public, unlike `_identifiers`, because two tiers need it and the alternative is a consumer
    re-implementing a parse of a payload we already hold — the RM41 lesson. `lookup.CitationHint`
    reads it so "does this PMID exist" can become "does this PMID name the paper you meant":
    existence alone cannot catch a fabricated citation, because PMIDs are densely allocated and an
    invented number is usually a real record for a different article.

    Every value is `None` when the field is absent rather than empty-string, so a caller can tell
    "PubMed did not say" from "PubMed said nothing is there" — the house tri-state, applied to
    metadata. `year` is the leading four digits of `pubdate` (which is free-form: `2017 Nov 20`,
    `2017`, `2017 Nov-Dec`), and nothing is invented when it does not start with a year.
    """
    out: dict[str, str | None] = {"title": None, "journal": None, "year": None, "first_author": None}
    for key, target in (
        ("title", "title"),
        ("fulljournalname", "journal"),
        ("sortfirstauthor", "first_author"),
    ):
        value = summary.get(key)
        if isinstance(value, str) and value.strip():
            out[target] = value.strip()
    pubdate = summary.get("pubdate")
    if isinstance(pubdate, str):
        leading = pubdate.strip()[:4]
        if leading.isdigit():
            out["year"] = leading
    return out