just_dna_enricher.lookup¶
just_dna_enricher.lookup ¶
Authoring lookups — answer an author's question about a variant or a citation (0.5).
The network half of the authoring surface. just_dna_compiler.hints inspects rows using nothing but
the module's own bytes; this module adds the facts that need a reference, in the same report shape.
It writes nothing — not a file, not a cell. Every pass in this package that does write does so
into a sidecar (resolution.csv, frequencies.csv, …) after a deliberate command. A lookup is a
question, and its answers come back as hints.Alterations with applied=False and a refusal
naming why the value is the author's to type.
That refusal is not fastidiousness. Almost every fact here is cross-examined later by a check that
only works because the author wrote the value independently: resolution._verify compares an
authored coordinate against the table, sequences.verify_reference_alleles compares an authored
ref against the genome, literature._doi_conflicts compares an authored DOI against the registry.
Supplying the cell from the same oracle the checker consults turns each into a tautology — and for an
rsid-only row _verify does not run at all, so the row would go from honestly unverified to
apparently verified. literature already reasons this way about one field: it asks Crossref about
the authored DOI because the derived one "exists by construction".
Clients are injected and reused. Each carries its own PacingGate — gnomAD is one request per
six seconds, ten per minute — so constructing a client per question throws away both the pacing state
and the connection pool. Hold one LookupClients for a session.
Offline is a first-class answer, not a failure: a check that could not run reports unchecked, never
absent. None is not False anywhere in this file.
LookupClients
dataclass
¶
LookupClients(
gnomad: GnomadClient | None = None,
eutils: EutilsClient | None = None,
europepmc: EuropePmcClient | None = None,
crossref: CrossrefClient | None = None,
pmc_idconv: PmcIdConverterClient | None = None,
ontology: OntologyClient | None = None,
ensembl: EnsemblResolver | None = None,
grch37: Grch37Client | None = None,
_lock: Lock = threading.Lock(),
)
The network clients a lookup session reuses. Every one is optional and lazily built.
Held by the caller rather than made per call, because each owns a PacingGate and an
httpx.Client: a fresh one per question discards the rate-limit state that keeps gnomAD from
refusing us, and reopens a connection for a single request.
One lazy path, and it is the same for every field (S92). A leg that finds its field unset
calls ensure, which builds the client under a lock, stores it on the bundle and returns it — so
an unfilled field is paced from the first call on, exactly as a filled one is, and the bundle
closes it in close(). Two of the eight legs used to do that and six built a per-request client
they closed in a finally, which discarded the pacing state this docstring says to keep; a host
filling six fields and leaving pmc_idconv unset had unpaced egress on one leg and nothing at
the call site to say which. The distinction is gone rather than documented.
A bundle a lookup_* call builds for itself is that call's to close. Hold one yourself for a
session; pass none for a one-shot question and the call closes what it opened.
A hosted CPIC draft is not in here on purpose: pgx_draft.draft_gene takes an injected
client=, so a host shares pacing by holding one CpicClient and passing it, and a field here
that nothing in this module reads would be a promise lookup cannot keep.
ensure ¶
The client in field name, built by factory on first use and kept on the bundle.
name is checked against the dataclass's own fields before anything is set: a plain
dataclass accepts setattr(self, "gnomd", …) silently, and a typo here would build a client
per call forever while looking exactly like the lazy path.
Source code in enricher/src/just_dna_enricher/lookup.py
close ¶
Close every client on the bundle — walked off the fields, never a hand-kept list.
Source code in enricher/src/just_dna_enricher/lookup.py
VariantHint
dataclass
¶
VariantHint(
rsid: str | None = None,
rsid_status: RsidStatus | None = None,
loci: list[dict] = list(),
rsid_candidates: list[str] = list(),
populations: list[dict] = list(),
clin_sig: list[dict] = list(),
pubmind: list[dict] = list(),
vrs_id: str | None = None,
findings: list[Finding] = list(),
alterations: list[Alteration] = list(),
checked: set[str] = set(),
snapshots: dict[str, str] = dict(),
)
What is known about one variant, and what of it is the author's to write.
loci is the uniform locus shape every link in this package returns —
{chrom, start, ref, alts} with start 1-based and alts comma-joined.
ambiguous
property
¶
More than one locus, or more than one rsID at a position — the author must choose.
CitationHint
dataclass
¶
CitationHint(
pmid: str | None = None,
doi: str | None = None,
pmid_exists: bool | None = None,
doi_exists: bool | None = None,
registry_doi: str | None = None,
pmcid: str | None = None,
open_access: bool | None = None,
abstract_available: bool | None = None,
title: str | None = None,
journal: str | None = None,
year: str | None = None,
first_author: str | None = None,
findings: list[Finding] = list(),
alterations: list[Alteration] = list(),
)
What is known about one citation. Every existence answer is tri-state.
Existence is not identity, which is why the bibliographic fields are here (S12). PMIDs are
densely allocated, so a recalled or invented 8-digit number is very likely to be a real record —
for a different paper — and pmid_exists=True alone therefore cannot catch a fabricated citation.
Fabrication is a failure of identity, so the answer has to name the paper it found and let the
caller compare it against the one they meant. title/journal/year/first_author all arrive in
the same esummary response that answers existence, so this costs no extra request.
OldAssemblyHint
dataclass
¶
OldAssemblyHint(
recovery: RsidRecovery,
findings: list[Finding] = list(),
alterations: list[Alteration] = list(),
)
What GRCh37 dbSNP records at an old coordinate, and why it is not written for you (RM48).
lookup_variant ¶
lookup_variant(
*,
rsid: str | None = None,
chrom: str | None = None,
start: int | None = None,
ref: str | None = None,
alts: str | None = None,
ambiguity: bool = False,
frequencies: bool = False,
offline: bool = False,
ensembl_cache: Path | None = None,
clinvar_cache: Path | None = None,
pubmind_cache: Path | None = None,
clients: LookupClients | None = None,
) -> VariantHint
Answer "what is this variant?" — validity, coordinates, alleles, frequencies, clinical calls.
Give an rsid, a coordinate, or both. Nothing is written and nothing is decided: a one-to-many
rsID returns every locus rather than picking one, and a position matching several rsIDs returns
every candidate. ambiguity=True adds the explicit warning; the candidates are always returned,
because hiding them would be a decision too.
frequencies=True is opt-in because it costs a paced gnomAD round trip (one per six seconds).
Source code in enricher/src/just_dna_enricher/lookup.py
lookup_old_assembly ¶
lookup_old_assembly(
*,
chrom: str,
start: int,
ref: str | None = None,
alts: str | None = None,
offline: bool = False,
clients: LookupClients | None = None,
) -> OldAssemblyHint
Answer "I have an hg19 coordinate — what is its rs-number?".
The authoring answer to a paper that predates GRCh38. It is recovery, not liftover, and the
difference is the whole design: an rs-number authored into variants.csv resolves through the
ordinary chain into a coordinate resolution._verify can cross-examine, while a lifted-over
position becomes the row's sole identity with nothing to check it against — an
unverifiable-by-construction identity, which is the hazard class behind this tree's 3,038-row
off-by-one.
Nothing is written, and the refusal is the sharpest one in _REFUSAL_BY_COLUMN: rsid is
identity_bearing, so a machine filling it would be performing an identity migration by network
lookup, exactly as _check_rsid_currency refuses to do for a merged id.
Several candidates are reported, never picked — --ref/--alts narrow them, and a pick among
equals is not a finding.
Source code in enricher/src/just_dna_enricher/lookup.py
lookup_citation ¶
lookup_citation(
*,
pmid: str | None = None,
doi: str | None = None,
pmcid: str | None = None,
offline: bool = False,
clients: LookupClients | None = None,
) -> CitationHint
Answer "does this citation exist, and what is its other identifier?".
A paywall hides the fulltext, never the PubMed record, so existence is answerable for paywalled work. Crossref covers what PubMed does not index at all (preprints, books, datasets).
pmcid= is the reverse direction, and it reports rather than fills (RM50). PubMed and PubMed
Central number articles independently, StudyRow.pmid requires the PubMed one, and a curator
holding only a PMC id previously had no route to it — the schema refused the cell and named no
remedy. The resolved PMID comes back as an advisory (applied=False) the author types
themselves: writing it into pmid would make literature.exists compare NCBI against NCBI, the
same argument that keeps doi unfilled. When both are given, the PMC id is additionally checked
against the PMID's own record, which is where a mismatched pair shows up.
Source code in enricher/src/just_dna_enricher/lookup.py
lookup_trait ¶
Is this trait CURIE current, obsolete, or unknown? (OLS4; unchecked when it cannot be asked.)
Source code in enricher/src/just_dna_enricher/lookup.py
lookup_gene ¶
Is this gene symbol approved or retired? (HGNC exact endpoints, never the fuzzy search.)
Source code in enricher/src/just_dna_enricher/lookup.py
as_report_rows ¶
A hint's advisory alterations as plain dicts, for a JSON surface or a table.
Deliberately not a HintReport: that type carries csv_out, and a lookup has no CSV to emit —
it answers a question about a value, not about a row a caller handed over.