just_dna_enricher.literature¶
just_dna_enricher.literature ¶
enrich-literature — the fourth pass: citations in, literature.csv out.
Closes the reachable part of the compiler's "does the cited study support the row?" blind spot. Three questions, in increasing ambition and decreasing coverage:
- Does the citation exist? PubMed
esummary, batched. A nonexistent PMID comes back as a record carrying anerrorkey, so this is a clean yes/no. - Do the identifiers agree? The DOI and PMCID arrive free in the same response. An absent
authored DOI is filled here in the sidecar; an authored DOI that contradicts the registry is a
finding. Neither is ever written back into
studies.csv— the enricher does not edit authored files, becausecontent_signatureis defined as reference-independent, and a network fetch that could change it would make that documented property false. - Does the quoted passage appear in the article? Only for the open-access subset, and honestly labelled as such.
All three are attested (RM45): the pass writes citation_existence, citation_identifier and
provenance_quote into verification.json on its way out, including on the offline return, so a
module can say which of the three was put and over how many citations. Until 0.6 the answers reached
a log line and the result object and died there, which left a module whose citations had been checked
indistinguishable from one where the command was never run.
Coverage is partial by nature, and saying so is part of the check. A pass that reported "0 quotes
found" for an article it could not read would be describing its own reach as if it were a property of
the module. So the result separates checked and not found from never retrievable — quotes_found
is null rather than zero when no fulltext could be read — and, since dogfooding caught the
conflation, also from nothing to check: a citation with no authored quote asked no question and is
not counted against coverage. The same distinctions are carried into manifest.literature.
Corrections to the drafted plan, made under probing rather than assumed:
- The PMC ID converter is not used, though the plan budgeted for it.
esummaryalready returns bothdoiandpmcinarticleids, and Europe PMC'ssearchreturnsdoi/pmcidtoo — so the converter is a third request for data already in hand. Worse, it answers a different question: for PMID 12345678 (a real, indexed PubMed record) it repliesstatus: error, "Identifier not found in PMC", which is about PMC membership, not existence. Wiring it in as an existence check would report every paywalled article as a broken citation. - Europe PMC is not an existence oracle either. Asked for three ids where one does not exist, it returns two results and simply omits the third — no error, no marker. Absence there is indistinguishable from "not indexed", so PubMed decides existence and Europe PMC only decides retrievability.
On running provenance_regex here. The charter requires a linear-time / ReDoS-safe engine for
pattern matching, written when the match was specified as consumer-side — arbitrary patterns meeting
arbitrary documents. Here the pattern comes from the module being enriched and the document from a
public archive, on the author's own machine, so the threat model is a curator writing a slow pattern by
accident rather than an attacker. That is worth a bound rather than a compiled dependency — so the
match runs under a wall-clock timeout, and a timeout is recorded as not checked.
That bound is enforced with a child process, not a thread, and the reason is worth stating because
the thread version looks correct: re cannot be interrupted, threads cannot be killed, and the
interpreter joins pool threads at exit — so a thread-based timeout returns on schedule and then hangs
the process on the way out. See regex_matches.
LiteratureEnrichmentError ¶
Bases: RuntimeError
Raised in strict mode when a citation does not resolve, or contradicts its own identifiers.
LiteratureUnavailable ¶
Bases: LiteratureEnrichmentError
PubMed could not be reached, so no citation question was put at all (RM101).
A subclass rather than a second exception, so every existing except LiteratureEnrichmentError
still catches it (P3 — additive within a major). It separates the source was asked and never
answered from a citation that genuinely does not resolve under strict. Only this one means
nothing was established either way.
Before RM101 an EutilsError travelled straight out of enrich_literature through a
try/finally with no except, so a caller's except LiteratureEnrichmentError was silent for
exactly the failure it was written for.
Scoped to the eutils leg on purpose. EuropePmcClient.fulltext and CrossrefClient.exists
already answer a transport failure with None rather than an exception — the tri-state withhold
this codebase uses for "could not be retrieved" — and turning either into an error here would
convert a withheld answer into a failed run.
DoiConflict
dataclass
¶
An authored DOI that disagrees with the one the registry reports for the same PMID.
PmcidConflict
dataclass
¶
An authored PMC id that disagrees with the one PubMed reports for the same PMID (RM50).
The DoiConflict shape, for the other cross-registry identifier. It costs no request — the PMC id
is already in the esummary articleids block that answered existence — and it catches the case
the schema guard cannot see: a cell like 21551363 (PMC3110567) carries a real PubMed id, so
nothing refuses it, while the two halves name different articles.
LiteratureResult
dataclass
¶
LiteratureResult(
rows: list[LiteratureRow],
missing: list[str] = list(),
doi_conflicts: list[DoiConflict] = list(),
pmcid_conflicts: list[PmcidConflict] = list(),
cited: list[str] = list(),
existence_checked: int = 0,
unresolved_citations: int = 0,
doi_verdicts_stale: int = 0,
doi_never_checked: int = 0,
noncommercial_quoted: list[str] = list(),
titles_as_quotes: list[str] = list(),
fulltext_requested: bool = True,
quotes_authored: int = 0,
quotes_found: int = 0,
quotes_checked: int = 0,
quotes_unchecked: int = 0,
quotes_unexamined: int = 0,
fulltext_checked: list[str] = list(),
abstract_checked: list[str] = list(),
doi_missing: list[str] = list(),
identifiers_authored: int = 0,
identifiers_compared: int = 0,
identifiers_conflicting: int = 0,
identifiers_unmatched: int = 0,
identifiers_foreign: int = 0,
sources: list[str] = list(),
mode: str = "best_effort",
skipped_offline: bool = False,
)
What the pass found, and — through subject_rows — what it found it about.
One subject set, read by the report, by the strict gates and by the attestation. The set is
the citations the module makes now, answered by the rows literature.csv holds now, and it is
the single rule the whole pass turns on. Three separate defects came from parts of this file
reading three different sets:
- the
strictgates read lists appended inside the fetch loop, so a module enriched once with--best-effortand then with--strictwas blessed on a citation PubMed has no record of — the refusal depended on the order the two runs happened in, and the attestation written by the same run saidfindings=1while the gate beside it said nothing was wrong; - the identifier cross-check ran inside that loop too, so an existing
literature.csvhid every DOI/PMC id disagreement; - the counts read
rows, which is everything the sidecar carries and never shrinks, so a citation deleted fromstudies.csvwent on being counted and no authored edit could clear the finding it produced.
Merge-not-clobber is what makes the distinction real: the sidecar keeps rows for citations the module has since dropped, and those rows are still written back (deleting a curator's row is not this pass's call) — they are simply not what any of this run's questions were about.
subject_rows
property
¶
The rows this run's questions were about: a row for a citation the module still makes.
Not rows, which is the whole sidecar — see the class docstring. Not this run's fetch list
either: literature.csv is the pin, so a row merged from an earlier run carries a verdict
that still stands, and counting only what this run looked up would let a no-op re-run replace
a true attestation with subjects=0, which reads as "the check ran and had nothing in scope".
coverage
property
¶
The sentence this pass exists to be able to say honestly.
The denominator is what had something to check — a citation with no authored quote was
not skipped for lack of a fulltext, it simply asked no question. Counting it as unretrievable
was a real bug found by running this against reference_examples/pathogenic_clinvar/, whose
single citation is open access and carries no quote: the old wording claimed its fulltext
could not be retrieved, which was the opposite of true.
Counted in quotes, off the tally, and not in citations off the run's own fetch lists. The
earlier wording read fulltext_checked, which only ever holds PMIDs this run retrieved, so
a no-op re-run over a pinned open-access citation printed "0 of 1 … 1 with nothing
retrievable" one line above "1/1 found" — the same false retrievability claim the paragraph
above records as a bug, arrived at from the other direction. Quotes also partition cleanly
where citations do not: one citation can carry a settled quote and a never-examined one at
once, so a citation-granular sentence has to put it in one bucket and be wrong about the
other.
EuropePmcClient
dataclass
¶
EuropePmcClient(
base_url: str = DEFAULT_EUROPEPMC_BASE,
batch_size: int = 25,
min_request_interval: float = 0.5,
timeout: float = 30.0,
gate: PacingGate | None = None,
_client: Client | None = None,
)
Open-access lookup and fulltext retrieval. Paced like every other client in this tier.
lookup ¶
pmid -> {pmcid, doi, is_open_access, license, abstract} for the ids Europe PMC knows.
Ids it does not know are simply absent from the result, with no error marker — which is why this is not an existence check. A caller must read a miss here as "not retrievable", never as "does not exist".
The abstract comes back for paywalled records too, in this same response, and that is worth more than it looks: probed across a mix of non-open-access papers, four of five carried one (only a 1994 non-research document did not). It is the difference between checking a quote for the open-access minority and checking it for nearly everything.
So does the article's license (RM46), which is why per-article terms cost no extra
request. Probed 2026-08-13 over 100 records: the values are lowercase CC spellings — cc by
(64), cc by-nc (28), cc by-nc-nd (8). It is independent of isOpenAccess and must not
be derived from it: PMID 28546431 comes back isOpenAccess: N with license: cc by, since
the flag describes Europe PMC's OA subset while the licence describes the article. Stored
verbatim; licensing.article_terms maps it to rights at read time.
Source code in enricher/src/just_dna_enricher/literature.py
fulltext ¶
Whitespace-normalized article text, or None when it cannot be retrieved.
None is a normal outcome, not an error: an embargoed or author-manuscript-only record answers
404, and the caller must record that as unchecked rather than as a failed match.
Source code in enricher/src/just_dna_enricher/literature.py
CrossrefClient
dataclass
¶
CrossrefClient(
base_url: str = DEFAULT_CROSSREF_BASE,
min_request_interval: float = 0.1,
timeout: float = 30.0,
contact_email: str | None = None,
gate: PacingGate | None = None,
_client: Client | None = None,
)
DOI existence, for the citations PubMed does not index.
PubMed answers "does this article exist" for anything it indexes — including paywalled work;
a paywall governs the fulltext, not the record. What it cannot answer for is everything outside
its scope: preprints, books, theses, datasets, standards. Those have DOIs and no PMID, and Crossref
is the registry that mints and resolves them (a probed bioRxiv preprint, 10.1101/2024.06.17.599351,
returns type: posted-content; a fabricated DOI returns a clean 404).
This is also what makes the 1.0 doi-first flip low-risk: once pmid becomes optional and a
citation may carry only a DOI, existence checking has to work without PubMed, and it already does.
Crossref asks callers to identify themselves in the User-Agent and gives the "polite pool" in
return; the contact address is sent only when one is configured, for the same reason eutils omits
it rather than inventing one.
exists ¶
True/False, or None when Crossref could not be asked.
None rather than False on a transport failure or an unexpected status: "we could not
check" and "this DOI does not exist" are different claims, and only the second is a finding
against the module. The translation stays here and the retrying stays in _request above —
reraise=True means the last failure arrives back at this except after the attempts are
spent, so the three-valued contract is unchanged and only the number of tries moved.
Source code in enricher/src/just_dna_enricher/literature.py
PmcIdRecord
dataclass
¶
One answer from the PMC id converter. in_pmc=False is PMC saying it has no such record;
an id the request never reached is absent from the result entirely, never one of these.
PmcIdConverterClient
dataclass
¶
PmcIdConverterClient(
base_url: str = DEFAULT_PMC_IDCONV_URL,
batch_size: int = 200,
min_request_interval: float = 0.5,
timeout: float = 30.0,
contact_email: str | None = None,
gate: PacingGate | None = None,
_client: Client | None = None,
)
PMCID → PMID, for a curator holding the wrong half of the pair (RM50).
It reports; it never fills. The PMID it returns is handed back as an advisory the author types
themselves, because filling StudyRow.pmid from NCBI would make literature.exists compare NCBI
against NCBI — the hints.REDUNDANCY_BEARING rule, already argued for doi (Crossref is asked
about the authored DOI, since a derived one exists by construction).
Why this exists although the pass docstring says the converter is unused. That statement is
true of the direction the pass needs — PMID → PMCID arrives free in the esummary articleids
block, so calling the converter for it would be a third request for data already in hand, and it
answers a different question (asked about a paywalled PMID it replies "Identifier not found in
PMC", which is about PMC membership rather than existence). None of that says anything about
PMCID → PMID, which is the direction the converter is actually for and the one a curator has
no other route to.
Probed 2026-08-13: the www.ncbi.nlm.nih.gov/pmc/utils/idconv/v1.0/ path 301-redirects to the
address below, an absent id comes back as a record with status: "error" and
errmsg: "Identifier not found in PMC", and pmid arrives as a JSON number.
resolve ¶
PMCID -> PmcIdRecord for every id the converter answered about.
Four outcomes, and none of them is spelled the same way as another. An id it resolved
carries a pmid; an id it knows with no PubMed id carries in_pmc=True, pmid=None (a real
answer — some PMC records genuinely have none); an id it has no record of carries
in_pmc=False with the service's own errmsg; and an id the request never reached is simply
absent from the result. A caller must not read a missing key as "does not exist" — that
conflation is exactly what S20 was about, one service over.
Source code in enricher/src/just_dna_enricher/literature.py
extract_text ¶
JATS XML → one normalized string, or None if it does not parse.
itertext() rather than a tag-stripping regex: real JATS carries nested inline markup inside
sentences (<italic>, <xref>, <sup>), and a regex that removes tags without joining their
text would silently glue words together and break matches that should succeed.
Source code in enricher/src/just_dna_enricher/literature.py
quote_matches ¶
Literal, whitespace- and case-insensitive containment. No engine, so nothing to bound.
regex_matches ¶
regex_matches(
pattern: str,
fulltext: str,
*,
timeout: float = DEFAULT_REGEX_TIMEOUT,
) -> bool | None
True/False, or None when the match could not be completed in time.
The three-way return is the point: a pattern that runs long has not failed to match, it has failed to be checked, and reporting it as "not found" would send an author to fix a quote that is probably there.
Why a subprocess and not a thread. The obvious implementation — submit to a
ThreadPoolExecutor and take future.result(timeout=...) — does not work, and fails in a way that
looks like it works: re never releases the GIL to a cancellation, threads cannot be killed, and
the interpreter joins non-daemon pool threads at exit. So a runaway pattern returns None on time
and then hangs the process on the way out. Verified by writing it that way first and watching the
test suite stop. A child process is the only bound in the standard library that can actually be
enforced, and it is cheap here: it runs only for a row that has a provenance_regex and a
retrievable fulltext, which is a small subset of a small subset.
Source code in enricher/src/just_dna_enricher/literature.py
enrich_literature ¶
enrich_literature(
spec_dir: Path,
*,
mode: str = "best_effort",
offline: bool = False,
check_fulltext: bool = True,
check_doi: bool = True,
regex_timeout: float = DEFAULT_REGEX_TIMEOUT,
write: bool = True,
eutils: EutilsClient | None = None,
europepmc: EuropePmcClient | None = None,
crossref: CrossrefClient | None = None,
) -> LiteratureResult
Fill literature.csv from the citations a module makes.
studies.csv is one citation site of several (RM47, RM132): MeasureBinRow.pmid grounds the
threshold its row states, and PharmVariantRow.pmid grounds the drug/genotype claim its row
makes — a claim studies.csv cannot ground, because a study row attaches to the whole variant. A
module whose only citations come from those tables is enriched exactly like one with a
studies.csv; reading fewer than all the sites would leave the rest unchecked, which is worse
than the honest gap the columns replaced. The kinds come from the compiler's own derived registry
through load_citing_rows/table_citations, so a kind that gains the column is read here with no
edit to this tier.
Existing rows are authoritative and merged, never clobbered — the same rule enrich() applies to
resolution.csv, with the same consequence: to regenerate after a machinery change you must delete
the file first. That includes the licence columns (RM46): rows written before 0.6 carry no
license, and re-running will not back-fill them, because merge-not-clobber cannot tell an
absent value from a curator's deliberate blank. Delete the sidecar to re-derive.
--offline makes this a no-op with a warning. There is no offline literature snapshot and there
will not be one; once literature.csv is written it is the pin, and later compiles read it
offline and deterministically.
Three checks are attested on the way out (_attest), the offline return included: a run that
did not put a question has to say so, or the manifest cannot tell it from a run that put one and
found nothing.
Source code in enricher/src/just_dna_enricher/literature.py
744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 | |
bibliographic ¶
Pull the fields that say which paper this is out of an esummary record.
Public, unlike _identifiers, because two tiers need it and the alternative is a consumer
re-implementing a parse of a payload we already hold — the RM41 lesson. lookup.CitationHint
reads it so "does this PMID exist" can become "does this PMID name the paper you meant":
existence alone cannot catch a fabricated citation, because PMIDs are densely allocated and an
invented number is usually a real record for a different article.
Every value is None when the field is absent rather than empty-string, so a caller can tell
"PubMed did not say" from "PubMed said nothing is there" — the house tri-state, applied to
metadata. year is the leading four digits of pubdate (which is free-form: 2017 Nov 20,
2017, 2017 Nov-Dec), and nothing is invented when it does not start with a year.