Skip to content

just_dna_enricher.eutils

just_dna_enricher.eutils

NCBI E-utilities client — batched, paced, and honest about the shape of a missing record.

Two callers with nothing else in common share this: the literature pass asks db=pubmed whether a cited PMID exists, and the identifier-hygiene pass asks db=snp whether an rsID is live, merged or gone. Both want the same three things — many ids per request, a rate limit obeyed, and a per-record failure that does not sink its batch — so the transport lives here and the interpretation lives with each caller.

The per-record error shape is the interesting part. Unlike the GraphQL services, eutils reports a missing record inside the result, as a normal-looking entry carrying an error key:

{"uid": "999999999", "error": "cannot get document summary"}

That is a clean, documented signal, and it is why existence-checking works at all. It also means a parser that only looks for HTTP failures would read "this id does not exist" as "this id is fine".

Rate limit. NCBI allows 3 requests/second without an API key and 10 with one, so the gate is 1/3 s by default and tightens to 1/10 s when NCBI_API_KEY is present in the environment. NCBI also asks callers to identify themselves with tool and email; both are sent when known. email is read from the environment and simply omitted when unset — inventing a plausible-looking address would be worse than sending none, because it would misattribute the traffic to someone.

EutilsError

Bases: RuntimeError

An eutils request failed outright (transport, HTTP status, or an unparseable body).

EutilsRateLimitedError

Bases: EutilsError

NCBI answered HTTP 429. Retried with backoff; the pacing gate should normally prevent it.

EutilsSettings dataclass

EutilsSettings(
    base_url: str = DEFAULT_EUTILS_BASE,
    tool: str = DEFAULT_TOOL,
    email: str | None = None,
    api_key: str | None = None,
    batch_size: int = 200,
    min_request_interval: float | None = None,
    timeout: float = 30.0,
)

identity_params

identity_params() -> dict[str, str]

tool/email/api_key, omitting whatever is genuinely unknown.

Source code in enricher/src/just_dna_enricher/eutils.py
def identity_params(self) -> dict[str, str]:
    """`tool`/`email`/`api_key`, omitting whatever is genuinely unknown."""
    params = {"tool": self.tool}
    if self.email:
        params["email"] = self.email
    if self.api_key:
        params["api_key"] = self.api_key
    return params

EutilsClient dataclass

EutilsClient(
    settings: EutilsSettings = EutilsSettings(),
    gate: PacingGate | None = None,
    _client: Client | None = None,
)

Batched, paced esummary access. Reused across a pass, so one gate covers every request.

esummary

esummary(
    db: str, ids: list[str]
) -> dict[str, dict[str, Any]]

uid -> record across as many batches as ids needs, in first-occurrence order.

Records carrying an error key are kept, not dropped: "NCBI has no summary for this uid" is the answer both callers came for, and discarding it would turn a detectable absence into an indistinguishable silence. Use is_missing to read it.

Source code in enricher/src/just_dna_enricher/eutils.py
def esummary(self, db: str, ids: list[str]) -> dict[str, dict[str, Any]]:
    """`uid -> record` across as many batches as `ids` needs, in first-occurrence order.

    Records carrying an `error` key are **kept, not dropped**: "NCBI has no summary for this uid"
    is the answer both callers came for, and discarding it would turn a detectable absence into an
    indistinguishable silence. Use `is_missing` to read it.
    """
    wanted = dedupe(i for i in ids if i)
    out: dict[str, dict[str, Any]] = {}
    for batch in batched(wanted, self.settings.batch_size):
        payload = self._get("esummary.fcgi", {"db": db, "id": ",".join(batch), "retmode": "json"})
        result = payload.get("result") or {}
        for uid in batch:
            record = result.get(uid)
            if isinstance(record, dict):
                out[uid] = record
            else:
                # Absent from `result` entirely — rarer than the `error` record, and it means the
                # same thing to a caller, so it is normalized into the same shape rather than
                # silently omitted.
                out[uid] = {"uid": uid, "error": NO_SUMMARY}
    return out

is_missing

is_missing(record: dict[str, Any]) -> bool

Whether an esummary record is NCBI's "no such document" answer rather than real data.

Source code in enricher/src/just_dna_enricher/eutils.py
def is_missing(record: dict[str, Any]) -> bool:
    """Whether an esummary record is NCBI's "no such document" answer rather than real data."""
    error = record.get("error")
    return bool(error) and NO_SUMMARY in str(error).lower()