pmid |
str |
required |
|
PubMed id, digits only. Normalized from StudyRow.pmid, which is free-form and may carry several ids or a [PMID: N] wrapper (spec.extract_pmids does the extraction). |
doi |
str | None |
optional |
|
Digital Object Identifier for this PMID, as the registry reports it. Filled here rather than written back into studies.csv: the enricher does not edit authored files, because content_signature is defined as reference-independent. |
pmcid |
str | None |
optional |
|
PubMed Central id (PMC…) when the article is in PMC — the key to fulltext. |
exists |
bool | None |
optional |
|
Whether PubMed returned a document summary for this id. False is a FACT (the citation does not resolve), distinct from null (never checked). |
is_open_access |
bool | None |
optional |
|
Whether Europe PMC reports retrievable open-access fulltext. Outside the fact set: embargoes lift, so this describes the world's state rather than the module's. |
license |
str | None |
optional |
|
The article's own licence, verbatim as the source spells it (Europe PMC writes cc by, cc by-nc, cc by-nc-nd, cc0). Per ARTICLE, never per source: PubMed's metadata and the publisher's text are different property, and the open subset spans all of those. |
share_alike |
bool | None |
optional |
|
Whether the article's licence is viral (the -SA family). Null when the terms could not be established — never false, which would state that they permit something. |
commercial_use |
bool | None |
optional |
|
Whether the article's licence permits commercial reuse (-NC makes it false). This is what a module quoting the article in provenance_quote has to answer for, since that quote is publisher text sitting in the module's own annotation layer. |
redistribution |
bool | None |
optional |
|
Whether the article's licence permits passing the text on. A third axis, not a reading of the other two: CC BY-NC forbids sale and allows sharing. |
quotes_authored |
int | None |
optional |
|
How many study rows cite this article with a provenance_quote/provenance_regex. Derivable from studies.csv; carried here so the coverage figure is readable in one place. |
quotes_found |
int | None |
optional |
|
How many of those were located. Null means not checked (nothing retrievable), which is materially different from 0 (something was read and the quote was not in it). Read it with quote_source: a 0 against abstract means only the abstract was searched, so the quote may still be in the body. |
quote_source |
str | None |
optional |
one of: abstract, fulltext |
What the quotes were matched against: fulltext|abstract. Null when neither could be retrieved. A hit is conclusive from either, but a miss is only conclusive against fulltext — which is why the two are recorded rather than collapsed. |
doi_exists |
bool | None |
optional |
|
Whether the authored/derived DOI resolves in Crossref. Independent of exists, which is PubMed's answer: a preprint, book or dataset has a DOI and no PMID, so this is the check that covers citations PubMed does not index at all. Read it with doi_checked, which names which DOI it is about. |
doi_checked |
str | None |
optional |
|
The DOI doi_exists is a verdict about — the authored one where a study row supplied it, else the registry's. Recorded because this table is a pin: a re-run does not refetch a row it already has, so without it a stored 'does not resolve' would be re-attributed to whatever DOI the author writes next, and correcting the citation could not clear the finding. Null when Crossref was not asked. |
source |
str | None |
optional |
|
Which service answered: pubmed|pmc-idconv|europepmc (open, like every source column). |
status |
str | None |
optional |
one of: ambiguous, not_found, resolved |
Lookup outcome: resolved|not_found|ambiguous |
fetched_at |
str | None |
optional |
|
ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-03T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the source published anything |