just-dna-enricher¶
Fill the source-independent resolution table (cache + live Ensembl) the compiler consumes.
Generated from the command tree itself at build time, so a flag added or renamed in the code appears here without anyone editing a table.
just-dna-enricher acmg¶
Build the ACMG secondary-findings snapshot from ACMG's published workbook (dev surface).
just-dna-enricher acmg build¶
Convert ACMG's SF workbook into the snapshot check-acmg --sf-list reads.
Why this exists: NCBI's page serves v3.2 and ACMG published v3.3 in June 2025, so the live scrape reports correctly authored rows as wrong. Nothing is downloaded here — the workbook is ACMG/Elsevier supplementary material and the author supplies their own copy, which is the same inject-only shape every other reference in this repo uses.
| Argument | Type | Default | Says |
|---|---|---|---|
workbook |
file |
required | ACMG SF supplementary workbook (.xlsx), downloaded by you. |
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/acmg_sf |
Output snapshot directory (writes acmg_sf.csv + release.json). |
--source-url |
str |
Where the workbook came from, recorded in release.json. | |
--doi |
str |
DOI of the statement the workbook accompanies, recorded in release.json. |
just-dna-enricher alphagenome¶
Re-encode AlphaGenome's AVI variant-impact scores into a cache lane. Reads a file you already hold — the artifact is 88.5 GB behind a sign-in whose eligibility clause bars classes of holder, so this never downloads.
just-dna-enricher alphagenome build¶
Build the AVI snapshot from a local copy of the artifact.
raw_score is stored as Int32 at a scale of 105 — exactly lossless, since the artifact
prints at most five decimals — and PHRED is not** stored: it is a rank, a function of
raw_score, and the 466 KB knot table beside the data reconstructs it while also carrying the
per-value ambiguity interval a threshold has to be checked against.
| Option | Type | Default | Says |
|---|---|---|---|
--input |
file |
required | The extracted alphagenome_variant_impact_score_snvs.tsv.gz (its .tbi must be beside it). Required, and there is no default URL: acquisition is yours, under your own acceptance of the AlphaGenome Services Additional Terms. |
--out |
directory |
data/repro/alphagenome_avi |
Output snapshot directory (writes data/alphagenome_avi-*.parquet, avi_knots.parquet, release.json, LICENSE.txt). |
--contig |
str |
Build only these contigs, repeatable. Omit for every contig the .tbi index knows. | |
--workers |
int range |
12 |
How many contigs to read at once. Twelve ran 24 contigs in 41-46 minutes; one takes about four times as long. |
--no-hash |
flag | Skip the source sha256. It is a few minutes over 88.5 GB; release.json then records null, which is unknown rather than unpinned. |
just-dna-enricher alphagenome check¶
Cross-check a module's variants against AlphaGenome's AVI scores. Reports, never repairs.
The local snapshot answers most of it. The Atlas is asked only where the knot table says the local data genuinely cannot decide — a threshold falling inside a printed score's PHRED interval — and that set is computed offline, before any request is spent.
| Argument | Type | Default | Says |
|---|---|---|---|
spec |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--reference |
directory |
An AVI snapshot directory. Omit to use $JUST_DNA_ALPHAGENOME_AVI_CACHE. | |
--threshold |
float |
A PHRED cut to check the module's variants against. Without one the pass is entirely offline: there is no question the local artifact cannot answer. | |
--offline |
flag | Never reach the Atlas. Straddling variants are recorded as nobody-asked, not as decided. | |
--refinement-cap |
int range |
200 |
Refuse rather than refine more than this many variants over the network in one run. |
--strict |
flag | Carried for the report; see the docstring. |
just-dna-enricher alphagenome expression¶
Fill expression_effects.csv with AlphaGenome's per-gene expression effects for one gene.
Example — and the --use is not decoration, an undeclared run is a no-op:
just-dna-enricher alphagenome expression ./my_module --gene TBX1 --use non-commercial
One row per (variant, gene): which way the variant moves that gene's predicted expression, how many of the 371 tissue tracks agree, and how far it sits from the gene. The interval is the gene's MANE span widened by the model's measured +/-512 kb attribution horizon, unless --chrom/--start/--end supply one; either way the gene names the server-side filter, and the MANE lane is still consulted for the distance, which an explicit interval cannot supply.
A whole gene is ~3.3 M SNVs at the measured 1,091 SNVs/s — about 50 minutes — and the cost is printed before the query runs rather than discovered during it.
| Argument | Type | Default | Says |
|---|---|---|---|
spec |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--gene |
str |
required | HGNC symbol. REQUIRED even with an explicit interval: the server-side gene filter is not an optimisation, and an unfiltered interval query is refused before it is sent. |
--chrom |
str |
Contig of an explicit interval. Wins over the gene's MANE span. | |
--start |
int range |
1-based start of that interval. | |
--end |
int range |
1-based end of that interval. | |
--min-score |
float |
Keep only pairs whose magnitude reaches this. Distal scores run ~10x lower than scores at the gene, so a flat bar keeps the proximal rows and looks like it filtered on effect. | |
--max-rows |
int range |
50000 |
Refuse rather than write more rows than this. Raising it is a deliberate act. |
--offline |
flag | No-op with a warning: this pass reads the Atlas, not a snapshot. | |
--dry-run |
flag | Report what would be written without writing it. | |
--use |
str |
unstated |
Declared use recorded on the licence row: unstated|non-commercial|commercial. AlphaGenome Output is NON-COMMERCIAL ONLY, so an undeclared run writes nothing and says so — pass --use non-commercial. |
just-dna-enricher alphagenome publish¶
Publish the AVI snapshot to HuggingFace.
Its own command because no other one can reach this lane (RM202). cache rebuild --publish
walks lanes that have a rebuild adapter, and this lane cannot have one — its source is behind an
eligibility gate, so there is nothing for an unattended rebuild to fetch. RM198 gave the lane a
publish_repo and left it unreachable.
The upload is two commits and, above 5 GB, goes through the resumable uploader (RM199): the
payload first, then release.json — the description must never arrive before the bytes it
describes.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot |
directory |
The built snapshot directory. Omit to use the resolved cache ($JUST_DNA_ALPHAGENOME_AVI_CACHE, then the cache base), falling back to where alphagenome build writes. |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
Target HF dataset. Default: just-dna-seq/alphagenome_avi. | |
--dry-run |
flag | Show what would be uploaded. Reads the repo's file list; sends nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher assertions¶
Fill clinical_assertions.csv from the coordinates already in resolution.csv.
Records what ClinVar says about each allele and how much review sits behind it — the star rating a compiled module previously discarded, so a one-star single submission and a practice guideline stopped being the same claim. Offline-capable: with a snapshot provisioned this pass never touches the network, and with none reachable it is a no-op rather than a failure.
It records; it does not adjudicate. Whether the module's own clin_sig agrees with ClinVar's is the
enrich cross-check's question, and that one warns in both modes on purpose.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every resolved allele has a ClinVar record. | |
--offline |
flag | Snapshot only: never touch the network. | |
--clinvar-cache |
path |
Explicit ClinVar snapshot directory. |
just-dna-enricher atlas¶
The AlphaGenome Atlas — precomputed variant scores over gRPC. Needs the [atlas] extra (grpcio + protobuf, 19 MB); the bindings are generated from pinned Apache-2.0 .proto sources rather than committed, so atlas generate runs once per checkout.
just-dna-enricher atlas generate¶
Fetch the pinned Atlas .proto sources and generate the gRPC bindings from them.
The sources are not vendored: the repository carries a commit id and a sha256 per file, and a
file that does not match its pin is refused rather than used (RM196). Needs grpcio-tools,
which is in \[dev] and deliberately not in \[atlas] — the runtime imports the bindings without
it. A released wheel carries both the sources and the bindings already, so this is a checkout
command.
| Option | Type | Default | Says |
|---|---|---|---|
--refetch |
flag | Re-download the pinned sources even if they are already on disk and match. |
just-dna-enricher cache¶
Pre-provision, rebuild and report the snapshot caches.
just-dna-enricher cache prepare¶
Leave this machine with every cache it can have — pull what is published, build what is not.
The complement of cache pull, and the one command a deployment actually wants. pull
fetches the published snapshots and stops; four lanes are not published for recorded reasons —
PharmVar's personal key, PubMind's absent terms, NCBI's policy over MANE, ACMG's supplementary
material — so a machine that only pulled is missing four caches and the checks that read them
skip themselves. This runs each lane by the route it has.
The route is a property of the lane, never a flag. A published lane pulls, because building it would spend an operator's bandwidth re-deriving bytes somebody already made; an unpublished one builds, because that is the only route there will ever be. Asking for the choice would be asking an operator to restate the licensing story.
A cache that is already present is left alone, exactly as cache pull leaves one alone, so
this is idempotent and cheap to re-run. Re-cutting a snapshot that exists is cache rebuild,
which writes somewhere else on purpose — a build straight into a live cache is visible half-done
to anything reading it, and a short parquet still has a footer.
The Python counterpart is just_dna_enricher.caches.prepare_caches, which this calls.
| Option | Type | Default | Says |
|---|---|---|---|
--only |
str |
[] |
Prepare just these caches (repeatable). Default: every one. |
--use |
str |
unstated |
Declared use: one of ['commercial', 'non_commercial', 'unstated']. |
--pin |
str |
[] |
lane=release, repeatable, for the lanes that are built rather than pulled. |
--source |
str |
[] |
lane=path, repeatable: build from a file you already hold. |
just-dna-enricher cache prune¶
Say what a published snapshot repo carries that its lane is not made of, and offer to delete it.
Deletion is never a side effect of publishing, and this is the command that makes that
affordable (RM186). A published repo accumulates: the publisher adds and does not remove, so a
layout change leaves the old spelling in place, and just-dna-seq/clinvar still carries the
159 MB single-file clinvar.parquet from before the per-chromosome split. Provisioning already
refuses to download it — the glob is what defends this tier — but any consumer globbing
data/*.parquet, the dataset viewer included, still gets two schemas under one relation.
Nothing here is a sweep. A file is a candidate only if the lane's own glob excludes it or a
LayoutShift declares it retired; README.md, .gitattributes, release.json, LICENSE.txt
and sidecar directories are never touched. Without --yes this reads and prints and does nothing
else, which is the mode to run first.
| Option | Type | Default | Says |
|---|---|---|---|
--only |
str |
[] |
Prune just these caches (repeatable). Default: every published one. |
--yes |
flag | Delete without asking. Without it this prints the plan and stops. |
just-dna-enricher cache pull¶
Download the published parquet snapshots from HuggingFace into the local caches.
The provisioning step a hosted deployment runs once, so no pass ever reaches a source live per request. Already-complete caches are trusted without touching the network, so this is re-runnable and cheap; a truncated file is removed and refetched.
A lane with nothing published says so and names its reason, which is a field on the registry
rather than a comment: PharmVar's and PubMind's are refusals, ACMG's and MANE's are permissions
nobody has established. Build those with cache rebuild.
| Option | Type | Default | Says |
|---|---|---|---|
--only |
str |
[] |
Pull just these caches (repeatable). Default: every publishable one. |
--use |
str |
unstated |
Declared use for the licence-gated snapshots. They forbid sale, so they are SKIPPED when unstated and REFUSED when commercial — downloading is taking the data. |
just-dna-enricher cache rebuild¶
Rebuild every cache this tier builds — acquire, convert, and optionally publish (RM176).
The one endpoint over eleven builders. Each per-lane X build command stays, and this calls
the same download_*/build_* functions they do, so there is one conversion algorithm with two
callers rather than two that have to agree. What differs is only flag plumbing: a per-lane command
offers the local-file inputs an operator holds, and a rebuild pass by definition holds none.
Every lane is built into <base>/<lane>/, never in place over a resolved cache. A rebuild
takes minutes and an enrich reading a half-written snapshot mid-flight would see a real but
incomplete table — the failure a resolver cannot detect, because a short parquet is still a
parquet. Point the caches at the new base when the run is done, or copy each directory across.
An outcome is three-valued. ACMG needs a workbook that is Elsevier supplementary material, PharmVar a personal key, CIViC a release date to pin — none of those is a failure, and a nightly rebuild reporting errors for them would be reporting the licences working as designed. They are printed as not run, with the reason, and the exit code counts only real failures.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/caches |
Base directory. Each lane is built into |
--only |
str |
[] |
Rebuild just these caches (repeatable). Default: every one that can be. |
--use |
str |
unstated |
Declared use: one of ['commercial', 'non_commercial', 'unstated']. |
--pin |
str |
[] |
lane=release, repeatable. e.g. --pin mane=1.5 --pin civic=2026-08-01. |
--source |
str |
[] |
lane=path, repeatable: build from a file you already hold instead of downloading. Required for acmg; the offline off-switch for clinvar, constraint, clinpgx, drug_labels, pubmind and strchive. mane and civic take three files each and refuse it. |
--publish |
flag | Also upload each rebuilt snapshot to its HuggingFace repo. | |
--dry-run |
flag | With --publish: show what would be uploaded, send nothing. |
just-dna-enricher cache status¶
Say which snapshots are present, where, and which release each holds.
Reads only: nothing is downloaded, so this is safe on a machine with no network and it is the first thing to run when a pass reports that a source was skipped.
just-dna-enricher check-acmg¶
Check each row's acmg_sf against the ACMG secondary-findings list (reports only).
Writes no authored cell, and records that the question was put — the same two halves as
check-identifiers. acmg_sf is an authored cell this asks a registry about, not a fact this
pass contributes, and filling it here would break the check (see hints.REDUNDANCY_BEARING). The
verification.json record is an attestation, never a value: it says the list was consulted and
over how many rows, which is the one thing a downstream reader cannot reconstruct from the
artifact (RM45/RM72).
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Exit 1 if any acmg_sf disagrees. | |
--offline |
flag | No network. Needs --sf-list, else nothing is checked. | |
--url |
str |
https://www.ncbi.nlm.nih.gov/clinvar/docs/acmg/ |
ACMG secondary-findings page URL (fallback). |
--sf-list |
directory |
Built ACMG SF snapshot (see acmg build). Preferred: NCBI's page still serves v3.2. Omit it and a snapshot in $JUST_DNA_ACMG_CACHE (or the shared cache base) is used; the page is scraped only when neither is there. |
just-dna-enricher check-identifiers¶
Report obsolete trait terms, retired gene symbols and unrecognised PGS accessions (online).
Writes no authored cell, and records that the question was put. Unlike the rsID check (whose
verdict lands on resolution.csv), these are module-level identifiers with no sidecar column to
record, and filling one from the registry being asked about it would make the comparison vacuous
— see hints.REDUNDANCY_BEARING. What this does write is verification.json: an attestation that
the five checks ran and over how many rows, never a value. A consumer holding the artifact has no
other way to tell "asked and clean" from "never asked" (RM45/RM72).
The PGS leg also writes sources.csv (RM163), and that is not an exception to the sentence above:
the Catalog's license is a field on each score record and it varies, so a module carrying an
academic-research-use-only score must not compile claiming the generic terms. The rows are the
terms, never a value in an authored cell.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Exit 1 if any identifier is stale. | |
--traits / --no-traits |
flag | on | Check trait_efo_id against OLS4. |
--genes / --no-genes |
flag | on | Check gene symbols against HGNC. |
--pgs / --no-pgs |
flag | on | Check pgs_id against the PGS Catalog, and the two authored cells beside it. |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. A PGS score licensed for academic research only bars sale, so a module citing one compiles ONLY with a declaration — and this flag is the one the compile's own refusal tells you to re-run with. |
just-dna-enricher check-repeat-bands¶
Compare a module's repeat_alleles.csv bands against STRchive's, and report the differences.
Writes no authored cell and never fails on a difference. Where a catalogue and an expert
author draw a repeat threshold in different places, both are claims by an authority, and a compile
that refused would make this format pick the winner — the rule the ClinVar clin_sig and PGx
allele-function checks already follow. --strict is accepted so the flag means one thing across
the tier, and it changes nothing here but the mode recorded in the report.
The catalogue's pathogenic_max is reported as its own finding and is never written: it is the
longest allele the literature records, not a clinical ceiling, and a module that imported it would
silently answer nothing at all for a longer one.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--catalogue |
path |
Built STRchive snapshot directory (see strchive build), or a STRchive-loci.json. |
|
--strict / --best-effort |
flag | Carried into the report. A band difference NEVER fails, in either mode. |
just-dna-enricher civic¶
Build the CIViC snapshot from a dated bulk release. CC0, so unlike the PubMind snapshot this one may be published.
just-dna-enricher civic build¶
Reduce a dated CIViC release to the parquet snapshot the direction-axis drafter reads.
The bulk release, not the GraphQL API, and the two are not interchangeable. Every row of
ClinicalEvidenceSummaries.tsv is status accepted; the API defaults to NON_REJECTED and
serves roughly 2.35x as many evidence items. A snapshot has to be reproducible from a pinned
input, so this reads the dated files and records the basis in release.json.
--submitted widens that basis without leaving the dated release (RM169). CIViC publishes
<date>-civic_accepted_and_submitted.vcf beside the TSVs, so unreviewed evidence is pinnable too
and no API read is needed. The TSVs stay primary — the VCF cannot carry a variant with no GRCh37
coordinate, which is exactly the class whose identity had to be read out of its name — and the VCF
supplies the curation status, the submitted evidence, and the identity for the 112 variants
VariantSummaries.tsv (itself accepted-only) does not describe. Those rows are stamped
identity_derivation="vcf_csq", and nothing is placed from the VCF's own GRCh37 position.
There is no --use flag. CIViC is CC0 on every axis, so a declared-use gate would permit
every build unconditionally, and a flag feeding a gate that never gates is a flag that does
nothing (@acquisition-gate-is-not-a-read-gate).
| Option | Type | Default | Says |
|---|---|---|---|
--release |
str |
Dated CIViC release to download, e.g. 01-Aug-2026. A DATED release, never the nightly: a snapshot that cannot name its input is one nothing can reproduce. | |
--evidence |
file |
Local ClinicalEvidenceSummaries.tsv. Use instead of --release to build offline. | |
--variants |
file |
Local VariantSummaries.tsv. | |
--profiles |
file |
Local MolecularProfileSummaries.tsv. | |
--out |
directory |
data/repro/civic |
Output snapshot directory (writes data/civic.parquet + release.json). |
--submitted |
flag | Also read the release's civic_accepted_and_submitted.vcf, so evidence a curator entered but no editor signed off joins the snapshot. Over 01-Aug-2026 that widens the direction corpus from 507 rows on 270 variants to 1149 on 397, adding 642 submitted rows and 127 variants, and every row carries the status CIViC gave it. Dated and pinnable like the TSVs, so the build stays reproducible. | |
--vcf |
file |
Local civic_accepted_and_submitted.vcf. Use with the local TSV flags to build offline. |
just-dna-enricher civic citations¶
Append the citations a CIViC variant carries that the dated bulk release cannot reach (RM160).
Why this is not part of civic build. The builder reads a dated release and is byte-
reproducible from it; civic reproduce proves it by building twice. The wider basis RM169 adopted
comes from a VCF, and a VCF record needs a POS — so submitted evidence attached to a variant with
no GRCh37 coordinate is published on exactly one surface, the GraphQL API, which has no release to
pin. The read is also one request per variant by construction, because evidenceItems takes a
single variantId. Batching it into the builder is the first repair anyone proposes and it is
exactly the reproducibility bargain this shape refused.
A recovered citation lands in studies.csv; literature.csv is derived from those PMIDs by the
literature command, and an article row nothing cites is dropped from the artifact. CIViC's own
curation status rides in confidence/confidence_unit, unconverted, so an accepted row and a
submitted row are not the same row. Appending only — a second run over an unchanged API adds
nothing — and enrich re-asks later and reports what has moved since.
CIViC is CC0, so there is no --use flag: a declared-use gate would permit every call
unconditionally, and a flag feeding a gate that never gates is a flag that does nothing.
| Argument | Type | Default | Says |
|---|---|---|---|
spec |
directory |
required | Module spec directory. |
| Option | Type | Default | Says |
|---|---|---|---|
--snapshot |
directory |
CIViC snapshot to map authored rows through. Default: the provisioned cache. | |
--variant-id |
int |
[] |
Ask about a CIViC variant id directly, repeatable. Its citations ground the MODULE rather than a variant, which is the only route to a record CIViC publishes no identity for — variant 1955 is the case this exists for. |
--offline |
flag | Do not fetch. Every subject is recorded as not-asked; no row is written. | |
--dry-run |
flag | Report what would be appended, write nothing. |
just-dna-enricher civic publish¶
Upload a built CIViC snapshot to a HuggingFace dataset repo (publisher/dev).
This one does not refuse, and the contrast with pubmind publish is the point. CIViC's content
is CC0 1.0 — a public-domain dedication with no share-alike, no bar on sale and attribution
requested rather than required — so there is no permission to establish before passing the bytes
on. PubMind's command exists in order to say no because its terms are unstated; PharmVar's cache is
unpublishable because its terms forbid it. Nothing here is in either position.
What the snapshot carries is a derivation of CIViC's release, not a copy of it: the germline
direction rows, placed on GRCh38 through identifiers CIViC itself publishes. release.json records
which dated release it came from and the accepted status basis, so a consumer can tell what they
are looking at without re-deriving it.
The ClinGen Allele Registry's answers are NOT in here, and that is deliberate rather than an oversight: the registry states no terms, and a lookup performed at draft time is a read, while baking its responses into a published file would be redistribution of bytes nobody has established we may pass on. The snapshot carries the CAID; resolving it stays the consumer's own fetch.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/civic.parquet + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
Target HF dataset (owner/name). Default: just-dna-seq/civic. | |
--message / -m |
str |
Commit message. | |
--dry-run |
flag | Show what would be uploaded without contacting HuggingFace. |
just-dna-enricher civic reproduce¶
Build the CIViC snapshot from a dated release and check it, end to end.
Five checks, and the third is the one worth the network. The first two are about us; the third is about whether the coordinates we produced are real.
- The release downloads and its bytes are recorded — a sha256 per file, so a rerun that disagrees is a finding about the source rather than a mystery.
- Two independent builds are byte-identical (Principle 7). A parquet has no inherent row order, so this is the check that the sort is doing its job.
- Every placed coordinate is cross-checked against the GRCh38 reference sequence. This is the external validation: the snapshot's positions come from RefSeq accessions inside ClinVar HGVS, and this asks an unrelated service (refget/seqrepo) whether the reference base at each of those positions is what we wrote. A wrong-build or off-by-one placement fails here and nowhere else.
- The drop registry closes — every input row kept or counted, an equality over a walked set.
- The published file list is exactly what the publisher would upload.
Exits non-zero if any check fails, so it is usable in CI.
| Option | Type | Default | Says |
|---|---|---|---|
--release |
str |
01-Aug-2026 |
Dated CIViC release to reproduce, e.g. 01-Aug-2026. |
--out |
directory |
data/repro/civic_reproduce |
Working directory. The release files and two independent builds land here. The default is under data/, which this workspace git-ignores wholesale. |
--keep |
flag | Leave the downloaded release files in place for inspection. | |
--offline |
flag | Skip the reference cross-check. The build and determinism checks still run. | |
--submitted |
flag | Reproduce the wider basis: also download the release's civic_accepted_and_submitted.vcf and build with it, so the submitted rows and their coordinates go through every check below rather than only the accepted ones. |
just-dna-enricher clinpgx¶
Build the ClinPGx clinical-annotation and drug-label snapshots, and cross-check against them.
just-dna-enricher clinpgx build¶
Download + build the ClinPGx snapshot (dev surface; needs polars).
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/clinpgx |
Snapshot output directory. |
--zip |
path |
An existing summaryAnnotations.zip (else downloaded). | |
--url |
str |
https://api.clinpgx.org/v1/download/file/data/summaryAnnotations.zip |
ClinPGx bulk download URL. |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. |
just-dna-enricher clinpgx build-labels¶
Download + build the regulator drug-label snapshot (dev surface; needs polars).
A second archive from a source this tier already adopted, with its own release.json: ClinPGx
publishes at least twelve downloads on this endpoint and they do not refresh in lockstep, so the
label snapshot is dated from its own CREATED_*.txt rather than from the annotation lane's.
There is no --offline: a builder's off-switch is passing --zip instead of downloading.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/drug_labels |
Snapshot output directory. |
--zip |
file |
A drugLabels.zip you already have. Without it the archive is downloaded. | |
--url |
str |
https://api.clinpgx.org/v1/download/file/data/drugLabels.zip |
ClinPGx bulk download URL. |
--use |
str |
unstated |
Declared use: one of ['commercial', 'non_commercial', 'unstated']. |
just-dna-enricher clinpgx check¶
Cross-check pharm_variants.csv against the ClinPGx snapshot.
The snapshot no longer has to be handed over by hand (RM38): explicit path → $JUST_DNA_CLINPGX_CACHE
/ the default cache → downloaded from HuggingFace. --offline stops at the second step.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--snapshot |
path |
Explicit ClinPGx snapshot dir. Omit it and the cache is used, or one is downloaded. | |
--offline |
flag | Use a local snapshot only: never download one. | |
--strict / --best-effort |
flag | Fail on a stale evidence level. | |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. |
just-dna-enricher clinpgx check-labels¶
Compare a module's gene/allele/drug claims against the drug labels five regulators publish.
Two join tiers, reported apart, and the tier belongs to the question. What do the agencies say about this gene and this medicine is the gene-tier subject; …and this star allele or rsID is the allele-tier one. A label naming both answers both, because they are two questions rather than one asked twice, and a gene-level agreement is not an allele-level agreement.
Writes no authored cell and never fails on a difference. Five agencies genuinely disagree with
each other — clopidogrel and CYP2C19 is Actionable PGx at four of them and Informative PGx at
the EMA — and a compile that refused would make this format pick the winner. A blank Testing
Level is a third of the file and is reported as unknown, never as No Clinical PGx.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--snapshot |
directory |
Built drug-label snapshot directory (see clinpgx build-labels). Omit it and $JUST_DNA_DRUG_LABELS_CACHE (or the shared cache base) is used. |
|
--strict / --best-effort |
flag | Carried into the report. A label difference NEVER fails, in either mode. | |
--use |
str |
unstated |
Declared use: one of ['commercial', 'non_commercial', 'unstated']. |
just-dna-enricher clinpgx publish¶
Publish a built ClinPGx snapshot so clinpgx check can provision it (publisher/dev).
LICENSE.txt travels with the parquet — the terms ClinPGx ships inside its own archive are what
license_sha256 pins, and a published snapshot without them pins nothing for whoever downloads it.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/*.parquet + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
just-dna-seq/clinpgx |
Target HuggingFace dataset repo (owner/name). |
--dry-run |
flag | Show what would be uploaded; send nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher clinpgx publish-labels¶
Publish a built drug-label snapshot so clinpgx check-labels can provision it (publisher/dev).
A second repo rather than a second table in just-dna-seq/clinpgx, for the reason the builder
already gives its own release.json: the two ClinPGx archives do not refresh in lockstep, and one
repo holding both would date the pair from whichever was published last.
Same grounds as clinpgx publish — CC BY-SA permits redistribution, forbids sale, and requires
attribution, which sources.csv carries. LICENSE.txt travels with the parquet: a share-alike
snapshot whose terms did not travel pins nothing for whoever downloads it.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/drug_labels.parquet + LICENSE.txt + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
just-dna-seq/clinpgx_drug_labels |
Target HuggingFace dataset repo (owner/name). |
--dry-run |
flag | Show what would be uploaded; send nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher clinvar¶
Build and publish the ClinVar reference snapshot (publisher/dev surface).
just-dna-enricher clinvar build¶
Convert a ClinVar VCF into the per-chromosome parquet snapshot the resolver reads.
| Option | Type | Default | Says |
|---|---|---|---|
--vcf |
file |
Local ClinVar VCF (.vcf.gz). Omit and pass --download to fetch from NCBI. | |
--download |
flag | Download the NCBI ClinVar GRCh38 VCF into --out first. | |
--out |
directory |
data/repro/clinvar |
Output snapshot directory (writes data/*.parquet + release.json). |
just-dna-enricher clinvar citations¶
Add ClinVar's literature links to a snapshot: data/citations.parquet ([dev], needs polars).
Separate from clinvar build because ClinVar publishes citations separately from the VCF — which
is precisely why a drafted gene panel could not compile without this: studies.csv is mandatory
and the VCF carries no PMIDs. Written beside the snapshot, so an existing cache keeps its bytes.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
required | Existing ClinVar snapshot dir. |
--citations |
file |
Local var_citations.txt. | |
--download |
flag | Fetch var_citations.txt first. | |
--url |
str |
https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt |
Source for --download. |
just-dna-enricher clinvar publish¶
Create-or-update the dataset repo and upload the built ClinVar snapshot (publisher/dev).
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/*.parquet + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
Target HF dataset (owner/name). Default: just-dna-seq/clinvar. | |
--message / -m |
str |
Commit message. | |
--dry-run |
flag | Show what would be uploaded. Reads the repo's file list; uploads nothing. |
just-dna-enricher cpic¶
Build and publish the CPIC snapshot, so a hosted enricher never reaches CPIC per request.
just-dna-enricher cpic build¶
Fetch CPIC whole into data/*.parquet + release.json (dev surface; needs polars).
No gene filter, deliberately: the whole database is ~120k narrow rows, and a snapshot covering only the genes the operator thought of answers "CPIC has nothing" for the next one.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/cpic |
Snapshot output directory. |
--endpoint |
str |
https://api.cpicpgx.org/v1 |
CPIC PostgREST base URL. |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. |
just-dna-enricher cpic publish¶
Create-or-update the dataset repo and upload the built CPIC snapshot (publisher/dev).
Publishable because CPIC's recorded terms permit redistribution — CC BY-SA grants sharing under
share-alike plus attribution, which sources.csv carries. PharmVar has no equivalent command, and
that is the design rather than an omission.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/*.parquet + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
just-dna-seq/cpic |
Target HuggingFace dataset repo (owner/name). |
--dry-run |
flag | Show what would be uploaded; send nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher dosage¶
Add ClinGen dosage-sensitivity rows to gene_metrics.csv (haploinsufficiency/triplosensitivity).
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every gene is ClinGen-curated. | |
--offline |
flag | No-op with a warning: ClinGen's curation list is a live download with no snapshot. | |
--url |
str |
https://ftp.clinicalgenome.org/ClinGen_gene_curation_list_GRCh38.tsv |
ClinGen gene-curation list URL. |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. ClinGen is CC0, so no declaration is refused here — it is recorded into sources.csv beside the rows it justifies. |
just-dna-enricher draft¶
Draft PGx tables for one or more genes from CPIC — appends rows, never overwrites one.
Re-runnable and additive, so a multi-gene module is built up a gene at a time. A row whose key is
already in the file is reported, never replaced: what CPIC now says about a row you already wrote
is a finding for pgx, not an edit for this command to make.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--gene |
str |
required | Gene to draft from CPIC (repeatable). |
--drug |
str |
[] |
Also draft CPIC's prescribing recommendations for this drug (repeatable). |
--allele |
str |
[] |
Draft only these star alleles, in all three tables (repeatable; *1 is always kept). A caller emits a bounded allele set, and n alleles is n(n+1)/2 pairs — CYP2D6 is 16,290 diplotypes unfiltered. Requires a single --gene, since a star name is gene-scoped. |
--population |
str |
Draft only this CPIC clinical context (e.g. 'NVI'). Default: every context, as rows. | |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. CPIC forbids sale, so a draft is SKIPPED when unstated and REFUSED when commercial. |
--offline |
flag | Draft from a built CPIC snapshot only; never reach CPIC live. | |
--cpic-cache |
path |
Explicit CPIC snapshot dir. | |
--dry-run |
flag | Report what would be added; write nothing. |
just-dna-enricher draft-clinpgx¶
Draft pharm_variants.csv rows from the ClinPGx snapshot — appends, never overwrites a row.
Narrow with --drug and re-run as the module grows. A row already in the file is reported, never
replaced: drift against ClinPGx is clinpgx check's finding, not this command's edit to make.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--snapshot |
directory |
required | Built ClinPGx snapshot (see clinpgx build). Inject-only; nothing is downloaded. |
--drug |
str |
[] |
Only annotations naming this drug (repeatable). |
--gene |
str |
[] |
Only annotations naming this gene (repeatable). |
--min-evidence-level |
str |
Keep annotations at least this strong: 1A|1B|2A|2B|3|4. | |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. ClinPGx forbids sale. |
--dry-run |
flag | Report what would be added; write nothing. |
just-dna-enricher draft-panel¶
Draft a gene panel's variants.csv rows from an authority — appends, never overwrites a row.
The drafted rows carry a genotype placeholder, so the module will not compile until you decide what each finding is about. That is deliberate: ClinVar publishes alleles, and whether carrying one is a carrier state or an affected one follows from the condition's inheritance mode, which the source does not say. Rows land in their gene's block, and a re-run leaves anything already there — stub or filled — exactly as it is.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--gene |
str |
[] |
Gene to draft rows for (repeatable). Required for every source but mitomap-miss, whose increment is asked for as a whole and where --gene only filters. |
--source |
str |
clinvar |
Which authority to draft the calls from: clinvar (the default); pubmind — an LLM's reading of the literature, which needs an operator-built snapshot and still reads the ClinVar one for its gene attribution; civic — curated cancer interpretations, which writes the DIRECTION axis rather than clin_sig and needs a civic build snapshot; or mitomap-miss — the curated mtDNA calls MITOMAP publishes and the ClinVar cache does not, which needs a mitomap miss snapshot. |
--mitomap-miss-cache |
directory |
Built MITOMAP-miss snapshot (see mitomap miss). Only read under --source mitomap-miss; omit it and $JUST_DNA_MITOMAP_MISS_CACHE is used. |
|
--civic-cache |
directory |
Built CIViC snapshot (see civic build). Only read under --source civic; omit it and $JUST_DNA_CIVIC_CACHE is used. |
|
--snapshot |
directory |
Built ClinVar snapshot (see clinvar build). Omit it and the cache is used, or the published snapshot downloaded — the citations table comes with it, which is what a panel needs to compile. Read for its gene attribution under --source pubmind, which publishes no gene column of its own. |
|
--pubmind-cache |
directory |
Built PubMind snapshot (see pubmind build), for --source pubmind. Omit it and $JUST_DNA_PUBMIND_CACHE is read; there is no published one to download. |
|
--offline |
flag | Use a local snapshot only: never download one. | |
--download / --no-download |
flag | on | Provision the published snapshot when no local one is found. Fetching it is this command's only network use, so --no-download coincides with --offline today; it is a separate switch because it says 'do not go and get one', not 'make no request'. |
--clin-sig |
str |
Comma-separated calls to include. Default: pathogenic,likely_pathogenic. | |
--min-review-stars |
int range |
2 |
Review-status floor, --source clinvar only. 2 = multiple submitters, no conflicts. |
--max-citations |
int range |
3 |
Study rows to draft per variant from ClinVar's literature links. 0 disables. --source clinvar only: PubMind's channel carries no PMID. |
--min-confidence |
int range |
1 |
Evidence-depth floor, --source pubmind only. PubMind's confidence counts how much of the literature spoke, 0-3; 1 means more than a single mention. |
--use |
str |
unstated |
Declared use (ClinVar is public domain). |
--dry-run |
flag | Report what would be added; write nothing. |
just-dna-enricher draft-repeats¶
Draft repeat_alleles.csv identity rows from STRchive — appends, never overwrites a row.
The bands are not drafted, and that is the design rather than a limitation. A drafted row
carries the gene, the motif as the catalogue spells it, the trait CURIE where the locus names
exactly one disease, and a conclusion placeholder — so the table cannot compile until a human
has filled in what each band means. measure_min/measure_max stay empty: run
check-repeat-bands once you have written them and it will report where the catalogue disagrees.
The catalogue's coordinates, ref_copies and locus_structure have no authored column to land
in; the run counts them and says so rather than dropping them silently.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--gene / -g |
str |
[] |
Restrict to these genes. Repeatable; omit for every catalogue locus. |
--catalogue |
path |
Built STRchive snapshot directory (see strchive build), or a STRchive-loci.json. Omit it and $JUST_DNA_STRCHIVE_CACHE (or the shared cache base) is used. |
|
--use |
str |
unstated |
Declared use: one of ['commercial', 'non_commercial', 'unstated']. |
--dry-run |
flag | Report what would be added; write nothing. |
just-dna-enricher enrich¶
Resolve a spec's variants into resolution.csv beside the spec. Exit 1 in strict mode if unresolved.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every variant resolves. | |
--offline |
flag | Cache-only: never touch the network. | |
--ensembl-cache |
path |
Explicit Ensembl cache dir/.duckdb. | |
--clinvar-cache |
path |
Explicit ClinVar snapshot dir. | |
--pubmind-cache |
path |
Built PubMind snapshot dir (from pubmind build) — the second authority in the clinical-significance concordance check. Omit it and $JUST_DNA_PUBMIND_CACHE is read; with neither, PubMind's leg reads unchecked rather than agreement. |
|
--clinvar / --no-clinvar |
flag | on | Use the ClinVar link (after the Ensembl cache). |
--gnomad / --no-gnomad |
flag | on | Use the gnomAD link (last, after live Ensembl). |
--vrs / --no-vrs |
flag | on | Mint GA4GH VRS allele ids onto resolved rows. |
--verify-ref / --no-verify-ref |
flag | on | Check each authored ref against the reference sequence and report disagreements. |
--verify-clinsig / --no-verify-clinsig |
flag | on | Check each authored clin_sig against the ClinVar snapshot's own (warns, never fails). |
--verify-rsids / --no-verify-rsids |
flag | on | Check each authored rsID against dbSNP for merges/withdrawals (online only). |
--verify-datasets / --no-verify-datasets |
flag | on | Check each release recorded in sources.csv against the one that source publishes now, and report the gap. One request per source, and the cheap question to put before --rederive: it tells you whether re-asking every subject is worth the run. |
--keep-par-twin |
flag | Record both contigs of a pseudoautosomal locus. Default keeps only the X spelling, which is the one every annotation source uses and the only one a hard-masked GRCh38 analysis set can match. | |
--rederive |
flag | Re-ask every source about every subject, including the ones already recorded, and report which of them changed value. An ordinary run gap-fills and never re-asks, so a source that quietly revised an answer moves nothing you could notice. | |
--keep-staging |
flag | Leave the staged answers beside resolution.csv after a successful run. They are removed by default; a killed run leaves them either way, and the next run resumes from them. |
just-dna-enricher enrich-and-compile¶
Enrich, then compile from the produced resolution.csv (offline, deterministic). Exit 1 on failure.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
output_dir |
directory |
required | Output dir for parquet + manifest.json |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every variant resolves. | |
--offline |
flag | Cache-only: never touch the network. | |
--ensembl-cache |
path |
Explicit Ensembl cache dir/.duckdb. | |
--clinvar-cache |
path |
Explicit ClinVar snapshot dir. | |
--clinvar / --no-clinvar |
flag | on | Use the ClinVar link (after the Ensembl cache). |
--gnomad / --no-gnomad |
flag | on | Use the gnomAD link (last, after live Ensembl). |
--frequencies |
flag | Also run the frequency pass (writes frequencies.csv). | |
--gene-metrics |
flag | Also run the gene-constraint pass (writes gene_metrics.csv). |
just-dna-enricher frequencies¶
Fill frequencies.csv from the coordinates already in resolution.csv (pass 2, online only).
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every resolved allele has a frequency. | |
--offline |
flag | No-op with a warning: gnomAD frequency has no offline snapshot. | |
--populations |
str |
Comma-separated ancestry groups to keep (e.g. 'global' for one row per allele). Default: all. | |
--dataset |
str |
Override the dataset label recorded on each row. |
just-dna-enricher gene-metrics¶
Fill gene_metrics.csv for the genes variants.csv mentions (pass 3, snapshot then live API).
With no local snapshot the v4.1 one is downloaded from HuggingFace first, exactly as enrich
provisions the Ensembl and ClinVar snapshots — --offline is what turns that off, and then the pass
is snapshot-only. Reaching the live API instead means v2.1.1 numbers, which the row's dataset
records; provisioning is what keeps a plain install on v4.1.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail unless every gene has constraint metrics. | |
--offline |
flag | Snapshot only: never touch the network. | |
--constraint-cache |
path |
Explicit gnomAD constraint snapshot dir. |
just-dna-enricher gene-validity¶
Fill gene_validity.csv with curated gene-disease assertions for the genes variants.csv names.
One row per (gene, disease, mode of inheritance, submitter) — the source's own grain. Mode of
inheritance is in the key because 59 ClinGen (gene, disease) pairs carry two curations that differ
only there, and submitter is in it because GenCC publishes the disagreement between submitters,
which is the thing it exists to publish.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--source |
str |
clingen |
Which submitter to read: clingen (expert panels) or gencc (an aggregate of nineteen). |
--strict / --best-effort |
flag | Fail unless every gene carries a curated assertion. | |
--offline |
flag | No-op with a warning: neither ClinGen nor GenCC publishes an offline snapshot. | |
--url |
str |
Override the submitter's export URL. |
just-dna-enricher gnomad¶
Build and publish the gnomAD gene-constraint snapshot (publisher/dev surface).
just-dna-enricher gnomad constraint¶
The gene-level constraint snapshot the gene-metrics pass reads offline.
just-dna-enricher gnomad constraint build¶
Reduce the per-transcript constraint TSV to the gene-level parquet the resolver reads.
| Option | Type | Default | Says |
|---|---|---|---|
--tsv |
file |
Local gnomAD constraint metrics TSV. Omit and pass --download to fetch it. | |
--download |
flag | Download the gnomAD v4.1 constraint TSV (95.5 MB) into --out first. | |
--out |
directory |
data/repro/gnomad_constraint |
Output snapshot directory (writes data/gnomad_constraint.parquet + release.json). |
just-dna-enricher gnomad constraint publish¶
Create-or-update the dataset repo and upload the built constraint snapshot (publisher/dev).
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/*.parquet + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
Target HF dataset (owner/name). Default: just-dna-seq/gnomad_constraint. | |
--message / -m |
str |
Commit message. | |
--dry-run |
flag | Show what would be uploaded without contacting HuggingFace. |
just-dna-enricher gwas¶
Fill gwas_effects.csv with the GWAS Catalog's published effect sizes for this module's rsIDs.
One row per published association, not per variant — a well-studied variant carries dozens across
different traits and papers. It does NOT fill weight: an authored weight is the author's model
of the finding, and no tool writes one. The two sit side by side and a consumer picks.
Reads effect_unit verbatim, including the Catalog's uninformative "unit", because a beta whose
scale is unknown must not look like one whose scale is shared. An association the Catalog
published without establishing which allele carries the effect keeps a null effect_allele and is
counted in the manifest, never dropped.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Severity ladder for findings; see the pass docstring. | |
--offline |
flag | No-op with a warning: this pass reads the REST API, not a snapshot. | |
--use |
str |
unstated |
Declared use recorded on the licence row: unstated|non-commercial|commercial. |
--study-facts / --no-study-facts |
flag | on | Follow each association's study and trait links. Costs 2 requests per association; measured at 382 requests for one real module. Off keeps effects, drops pmid/trait/ancestry PERMANENTLY for the rows it writes: the merge is keyed on association_id, so a later run with study facts on skips those rows rather than back-filling. Delete gwas_effects.csv to re-derive them. |
just-dna-enricher hint¶
Look up what is known about a variant or citation. Writes nothing.
just-dna-enricher hint citation¶
Does this citation exist, and is it the paper you meant?
A paywall hides the fulltext, never the PubMed record, so existence is answerable for paywalled
work; Crossref covers what PubMed does not index at all. Every answer is tri-state — unknown
means the registry could not be asked, which is not the same as "no such paper".
Existence is not identity. PMIDs are densely allocated, so a recalled or invented number is
very likely to be a real record for a different article, and pmid_exists alone cannot catch a
fabricated citation. The title, journal, year and first author come back in the same response and
are printed for exactly that comparison (S12).
--pmcid goes the other way. Every pmid in the schema keys on the PubMed id — studies.csv,
a binning row's and a pharm_variants.csv row's alike — and a curator holding only a PMC… id
had no route to it: the schema refused the cell and named no remedy. This resolves it and then asks PubMed which paper that is. The id is reported,
never written: filling pmid from NCBI would make the existence check compare NCBI with itself.
| Option | Type | Default | Says |
|---|---|---|---|
--pmid |
str |
PubMed id to check. | |
--doi |
str |
DOI to check (the one you authored). | |
--pmcid |
str |
PubMed Central id (PMC…) to resolve to the PubMed id tables key on. | |
--offline |
flag | Skip the check and say so. | |
--json |
flag | Emit the full machine answer. |
just-dna-enricher hint gene¶
Is this gene symbol approved or retired?
| Argument | Type | Default | Says |
|---|---|---|---|
symbol |
str |
required | Gene symbol, e.g. MTHFR |
just-dna-enricher hint recover¶
Which rs-number GRCh37 dbSNP records at an hg19/GRCh37 coordinate.
For a paper that predates GRCh38. Author the rs-number, not a converted position: an
rs-number resolves into a coordinate the compiler can cross-examine, where a lifted-over position
becomes the row's only witness to itself. Nothing is written — the rs-number is the row's
identity, and a machine filling one migrates variant_key with no authored edit anywhere.
| Option | Type | Default | Says |
|---|---|---|---|
--chrom |
str |
required | Chromosome of the old coordinate. |
--start |
int |
required | 1-based GRCh37 position. |
--ref |
str |
Reference allele, to narrow the answer. | |
--alts |
str |
Alt allele(s), comma-separated. | |
--offline |
flag | Skip the lookup and say so. | |
--json |
flag | Emit the full machine answer. |
just-dna-enricher hint trait¶
Is this trait id current, obsolete, or unknown?
| Argument | Type | Default | Says |
|---|---|---|---|
curie |
str |
required | Trait CURIE, e.g. EFO_0004340 |
just-dna-enricher hint variant¶
Validity, coordinates, alleles, populations and clinical calls for one variant.
Nothing is decided for you: a one-to-many rsID returns every locus and a position matching
several rsIDs returns every candidate. The coordinate is reported, never written into
variants.csv — resolution puts it in resolution.csv, which is where it belongs.
| Option | Type | Default | Says |
|---|---|---|---|
--rsid |
str |
dbSNP id to look up. | |
--chrom |
str |
Chromosome (with --start). | |
--start |
int |
1-based position (with --chrom). | |
--ref |
str |
Reference allele, for an allele-exact lookup. | |
--alts |
str |
Alt allele(s), comma-separated. | |
--ambiguity |
flag | Warn when the answer is not unique. | |
--frequencies |
flag | Add gnomAD populations (paced: ~6s). | |
--offline |
flag | Snapshots only; never touch the network. | |
--ensembl-cache |
path |
Explicit Ensembl cache. | |
--clinvar-cache |
path |
Explicit ClinVar snapshot. | |
--pubmind-cache |
path |
Explicit PubMind snapshot (see pubmind build); $JUST_DNA_PUBMIND_CACHE otherwise. |
|
--json |
flag | Emit the full machine answer. |
just-dna-enricher literature¶
Fill literature.csv from a module's citations (pass 4, online only).
studies.csv is one citation site of several: a pmid on a binning row grounds the threshold it
sits on, and one on a pharm_variants.csv row grounds that row's own drug and genotype claim. The
pass reads every site, so a module citing only from those tables is enriched rather than refused.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail if a cited PMID does not resolve. | |
--offline |
flag | No-op with a warning: there is no offline PubMed snapshot. | |
--fulltext / --no-fulltext |
flag | on | Also match provenance quotes against fulltext, falling back to the abstract. |
--doi / --no-doi |
flag | on | Also confirm the authored DOI resolves in Crossref (covers preprints/books). |
just-dna-enricher litvar¶
Which papers a variant-literature index holds for a module's alleles, and at which tier. Reports only; writes no authored cell and no table row.
just-dna-enricher litvar coverage¶
Report LitVar's literature coverage per locus, naming the tier that answered.
It answers which papers discuss an allele that is already identified. It does not answer which allele a name meant — those read as the same question and are not. Measured against the two hardest records in this repository (CIViC 1955 and 2131, four candidate alleles with registered CAIDs), the index returns no node for any of them, because PubTator3 mines titles and abstracts and those alleles live in a table inside a paywalled paper. Do not reach for this to recover an identity.
Writes no row and no sources.csv entry: nothing here reaches a module's tables, so the module
does not use this source. What it does write is verification.json — an attestation that the
question was put, over how many loci, and at which tier each was answered.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--offline |
flag | No network; every locus is recorded as unchecked. | |
--quiet |
flag | Only the tier summary, not a line per locus. |
just-dna-enricher litvar gene¶
List every node LitVar holds under a gene symbol, grouped by tier. Writes nothing.
This is the endpoint that serves line-delimited Python repr() rather than JSON, and the tier
split is the reason to look: of 588 HFE nodes on 2026-09-01, 220 are rsID-only, 69 carry a
ClinGen allele id, exactly one is the gene node, and the remaining 298 are unnormalized protein
strings — text a miner saw, not an identity anything should join on.
| Argument | Type | Default | Says |
|---|---|---|---|
gene |
str |
required | Gene symbol, e.g. HFE |
just-dna-enricher mane¶
Build the MANE transcript snapshot — the numbering frame, cached and pinned instead of downloaded by hand and cited in prose.
just-dna-enricher mane build¶
Reduce one pinned MANE release to the three parquet tables the numbering frame needs.
All three files, in one pass. The summary is the frame; changed_select_accessions is the
currency check and Update_Affects_CDS is the numbering-frame axis stated by the source; the
negative roster is what makes "MANE has no answer for this gene" distinguishable from "nobody
asked", with the reason attached. Shipping the cache without the thing that notices it going
stale is the defect this command exists to close.
There is no --offline flag: the off-switch is passing the local files instead of
--download. And there is no --use flag, because NCBI states a policy rather than a licence —
every gating axis is unknown, and check_declared_use returns a skip for an unknown whatever
the declaration says, so the gate would silently skip every build. A flag feeding a gate that
never gates is a flag that does nothing (@acquisition-gate-is-not-a-read-gate).
MANE is the default, not the answer. A gene with two rows carries two CDS numbering frames
and MANE_status says which; a gene with one row says nothing about the isoforms MANE does not
carry, so a pass treating this table as an oracle would be wrong in a way the table cannot report.
| Option | Type | Default | Says |
|---|---|---|---|
--download |
flag | Fetch the release from NCBI. Without --release the newest version is discovered from current/README_versions.txt and then pinned to its versioned directory. | |
--release |
str |
MANE version to pin, e.g. 1.5. Resolves to release_ |
|
--summary |
file |
Local MANE.GRCh38.v |
|
--changed |
file |
Local MANE.GRCh38.v |
|
--not-in-mane |
file |
Local MANE.GRCh38.v |
|
--versions |
file |
Local README_versions.txt. Optional, and the only way an offline build can name its release: a filename is never parsed for one. | |
--out |
directory |
data/repro/mane |
Output snapshot directory (writes data/*.parquet + release.json). |
just-dna-enricher mitomap¶
Build the MITOMAP snapshot and the derived miss lane. CC BY 3.0 with commercial use stated free, so a deployment may publish the snapshot.
just-dna-enricher mitomap build¶
Cut the two curated mtDNA variant tables, their citations and the references out of the dump.
602 mmutation rows and 494 rtmutation rows out of 6.76 million lines. The snapshot records
every count this build computes — rows per table, the dump's own per-table edit dates, how much of
reference.nlmid is a PMID, the alleles that cannot be spelled as VCF and the brackets that are
not a documented VCEP class — because a number computed and dropped is one every reader has to
recompute.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/mitomap |
Output snapshot directory (writes data/mitomap-*.parquet + release.json). |
--dump |
file |
A mitomap.dump.sql.gz you already have. Without it the dump is downloaded — the data surface answers plain curl, unlike the web surface. A local dump carries no Last-Modified, so that snapshot is honestly unlabelled. | |
--url |
str |
https://mitomap.org/downloads/mitomap.dump.sql.gz |
Source URL for the dump (used only when --dump is absent). |
just-dna-enricher mitomap miss¶
Join MITOMAP against the ClinVar chrMT parquet and write the increment (RM171).
A derived lane, not a download. Its acquire stage is both parents being on disk, and a parent that is absent is reported as could-not-run rather than as an empty increment — a miss set computed without ClinVar would say MITOMAP publishes a thousand alleles nobody else has, from a comparison that never ran.
Exact (start, ref, alt) on chrMT, upper-cased both sides, no position-level fallback. Four
buckets, and draft-panel --source mitomap-miss writes only one of them.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/mitomap_miss |
Output snapshot directory (writes data/mitomap_miss.parquet + release.json). |
--mitomap-cache |
directory |
Built MITOMAP snapshot (see mitomap build). Omit it and $JUST_DNA_MITOMAP_CACHE is used. |
|
--clinvar-cache |
directory |
Built ClinVar snapshot (see clinvar build). Omit it and $JUST_DNA_CLINVAR_CACHE is used. |
just-dna-enricher mitomap publish¶
Create-or-update the dataset repo and upload the built MITOMAP snapshot (publisher/dev).
Publishable on the source's own terms: CC BY 3.0, with commercial and clinical use stated free and
attribution the one condition — which the snapshot's SourceRow carries. Only the parent lane
publishes. The derived miss snapshot pins two parent digests, so a pulled copy would be an
increment whose own currency check cannot be run by whoever pulled it; it is rebuilt locally from
the parents instead, which is cheaper than the download and cannot be stale.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (data/ + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
just-dna-seq/mitomap |
Target HuggingFace dataset repo (owner/name). |
--dry-run |
flag | Show what would be uploaded; send nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher pgx¶
Cross-check star-allele tables against PharmVar/CPIC and record terms into sources.csv.
Snapshot first, live second (RM38). A built snapshot serves the check without egress and without
spending a shared per-IP budget; --offline says snapshot-only, and a leg with neither is skipped
with a reason rather than silently passing.
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--strict / --best-effort |
flag | Fail on an allele-function discrepancy. | |
--offline |
flag | Snapshots only: never reach PharmVar or CPIC live. | |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. Sources that forbid sale are SKIPPED when unstated and REFUSED when commercial. |
--pharmvar / --no-pharmvar |
flag | on | Consult PharmVar (needs PHARMVAR_API_KEY). |
--cpic / --no-cpic |
flag | on | Consult CPIC (open, no key). |
--cpic-cache |
path |
Explicit CPIC snapshot dir. | |
--pharmvar-cache |
path |
Explicit PharmVar snapshot dir. |
just-dna-enricher pharmvar¶
Build the PharmVar snapshot with your own key. Operator-built and inject-only; never published.
just-dna-enricher pharmvar build¶
Fetch PharmVar whole into data/*.parquet + release.json (dev surface; needs polars + a key).
There is no pharmvar publish, and there will not be. The data is pulled under a key
PharmVar's terms §2 make personal and non-transferable, and no axis SourceTerms records covers
passing that on — an unestablished permission is not a permission. Build your own; point at it with
$JUST_DNA_PHARMVAR_CACHE or --pharmvar-cache.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/pharmvar |
Snapshot output directory. |
--use |
str |
unstated |
Declared use: unstated | non-commercial | commercial. |
just-dna-enricher pubmind¶
Build the PubMind literature-derived snapshot. Operator-built and inject-only; never published.
just-dna-enricher pubmind build¶
Reduce the ANNOVAR-distributed PubMind table to the parquet snapshot the checks read.
There is no --use flag, and its absence is the design. PubMind's data terms could not be
established, and unknown terms warn rather than gate (commercial_use=None never taints a
module). A declared-use gate here would refuse every build unconditionally, and a flag feeding a
gate that never gates is a flag that does nothing. What the unknown terms do gate is publishing
a snapshot or a module carrying these bytes — see pubmind publish, which refuses.
| Option | Type | Default | Says |
|---|---|---|---|
--table |
file |
Local hg38_pubmind_db.txt.gz. Omit and pass --download to fetch it from ANNOVAR. | |
--download |
flag | Download the ANNOVAR-distributed PubMind table into --out first. | |
--out |
directory |
data/repro/pubmind |
Output snapshot directory (writes data/pubmind.parquet + release.json). |
just-dna-enricher pubmind publish¶
Refuse to publish the PubMind snapshot, and say why.
The command exists in order to refuse. A missing one reads as an oversight somebody will helpfully add, and the reason belongs where a reader looks for it rather than only in a design document.
just-dna-enricher strchive¶
Build the STRchive repeat-locus snapshot. MIT-licensed, so a deployment may publish it.
just-dna-enricher strchive build¶
Fetch (or copy in) the STRchive catalogue and record its provenance beside it.
Pin a release: the default branch moves, so a comparison whose reference is "whatever was there
that afternoon" cannot be re-run, and only a pinned build gets a dataset label the verification
record can name.
| Option | Type | Default | Says |
|---|---|---|---|
--out |
directory |
data/repro/strchive |
Output snapshot directory (writes STRchive-loci.json + release.json). |
--catalogue |
file |
A STRchive-loci.json you already have. Without it the file is downloaded. | |
--release |
str |
Upstream release tag to pin, e.g. v2.26.0. Without it, the default branch, unlabelled. |
just-dna-enricher strchive publish¶
Create-or-update the dataset repo and upload the built STRchive catalogue (publisher/dev).
Publishable on the source's own terms: STRchive is MIT, which grants redistribution outright.
This lane had no publish command because it grew from a check rather than from a cache, not
because anything withheld the permission — the same distinction the roster draws between CIViC's
absent ensure_* (a gap) and PharmVar's (a refusal).
Publish a pinned build. An unlabelled snapshot carries no dataset, so whoever pulls it can
run the comparison and cannot say which release they compared against — build with --release
first, and this refuses nothing but says so.
| Argument | Type | Default | Says |
|---|---|---|---|
snapshot_dir |
directory |
required | Built snapshot directory (STRchive-loci.json + release.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
just-dna-seq/strchive |
Target HuggingFace dataset repo (owner/name). |
--dry-run |
flag | Show what would be uploaded; send nothing. | |
--message / -m |
str |
Commit message. |
just-dna-enricher template¶
Print a header-only CSV for one authored table kind, generated from the live models.
Kept working here, but just-dna-compiler template is canonical: this needs no network, and an
author who installed only the tier that owns the CSV shape should not have to add the network
tier to get a header. See just-dna-compiler stub for a template with rows to replace.
| Argument | Type | Default | Says |
|---|---|---|---|
kind |
str |
required | Authored CSV to emit a header for, e.g. repeat_alleles.csv |
just-dna-enricher upload¶
Upload a compiled module to a HuggingFace dataset collection (publisher/dev surface).
Writes data/
Refuses when the versioned path already holds a different artifact (compare by artifact.digest), unless --force. The flat path means latest and is overwritten either way.
| Argument | Type | Default | Says |
|---|---|---|---|
module_dir |
directory |
required | Compiled module directory (at least one annotation parquet + manifest.json). |
| Option | Type | Default | Says |
|---|---|---|---|
--repo |
str |
Target HF dataset (owner/name). Default: just-dna-seq/annotators. | |
--name |
str |
Module name under data/ |
|
--message / -m |
str |
Commit message. Default: 'Add |
|
--dry-run |
flag | Show what would be uploaded without contacting HuggingFace. | |
--force |
flag | Overwrite data/ |
just-dna-enricher vrs¶
GA4GH VRS allele identity for an already-resolved module.
just-dna-enricher vrs mint¶
Stamp ga4gh:VA.… allele ids onto resolution.csv (substitutions offline, indels online).
| Argument | Type | Default | Says |
|---|---|---|---|
spec_dir |
directory |
required | Module spec directory |
| Option | Type | Default | Says |
|---|---|---|---|
--offline |
flag | Substitutions only: indels need the reference sequence, which means a network call. |