Skip to content

just-dna-enricher

Fill the source-independent resolution table (cache + live Ensembl) the compiler consumes.

Generated from the command tree itself at build time, so a flag added or renamed in the code appears here without anyone editing a table.

just-dna-enricher acmg

Build the ACMG secondary-findings snapshot from ACMG's published workbook (dev surface).

just-dna-enricher acmg build

Convert ACMG's SF workbook into the snapshot check-acmg --sf-list reads.

Why this exists: NCBI's page serves v3.2 and ACMG published v3.3 in June 2025, so the live scrape reports correctly authored rows as wrong. Nothing is downloaded here — the workbook is ACMG/Elsevier supplementary material and the author supplies their own copy, which is the same inject-only shape every other reference in this repo uses.

Argument Type Default Says
workbook file required ACMG SF supplementary workbook (.xlsx), downloaded by you.
Option Type Default Says
--out directory data/repro/acmg_sf Output snapshot directory (writes acmg_sf.csv + release.json).
--source-url str Where the workbook came from, recorded in release.json.
--doi str DOI of the statement the workbook accompanies, recorded in release.json.

just-dna-enricher alphagenome

Re-encode AlphaGenome's AVI variant-impact scores into a cache lane. Reads a file you already hold — the artifact is 88.5 GB behind a sign-in whose eligibility clause bars classes of holder, so this never downloads.

just-dna-enricher alphagenome build

Build the AVI snapshot from a local copy of the artifact.

raw_score is stored as Int32 at a scale of 105 — exactly lossless, since the artifact prints at most five decimals — and PHRED is not** stored: it is a rank, a function of raw_score, and the 466 KB knot table beside the data reconstructs it while also carrying the per-value ambiguity interval a threshold has to be checked against.

Option Type Default Says
--input file required The extracted alphagenome_variant_impact_score_snvs.tsv.gz (its .tbi must be beside it). Required, and there is no default URL: acquisition is yours, under your own acceptance of the AlphaGenome Services Additional Terms.
--out directory data/repro/alphagenome_avi Output snapshot directory (writes data/alphagenome_avi-*.parquet, avi_knots.parquet, release.json, LICENSE.txt).
--contig str Build only these contigs, repeatable. Omit for every contig the .tbi index knows.
--workers int range 12 How many contigs to read at once. Twelve ran 24 contigs in 41-46 minutes; one takes about four times as long.
--no-hash flag Skip the source sha256. It is a few minutes over 88.5 GB; release.json then records null, which is unknown rather than unpinned.

just-dna-enricher alphagenome check

Cross-check a module's variants against AlphaGenome's AVI scores. Reports, never repairs.

The local snapshot answers most of it. The Atlas is asked only where the knot table says the local data genuinely cannot decide — a threshold falling inside a printed score's PHRED interval — and that set is computed offline, before any request is spent.

Argument Type Default Says
spec directory required Module spec directory
Option Type Default Says
--reference directory An AVI snapshot directory. Omit to use $JUST_DNA_ALPHAGENOME_AVI_CACHE.
--threshold float A PHRED cut to check the module's variants against. Without one the pass is entirely offline: there is no question the local artifact cannot answer.
--offline flag Never reach the Atlas. Straddling variants are recorded as nobody-asked, not as decided.
--refinement-cap int range 200 Refuse rather than refine more than this many variants over the network in one run.
--strict flag Carried for the report; see the docstring.

just-dna-enricher alphagenome expression

Fill expression_effects.csv with AlphaGenome's per-gene expression effects for one gene.

Example — and the --use is not decoration, an undeclared run is a no-op:

just-dna-enricher alphagenome expression ./my_module --gene TBX1 --use non-commercial

One row per (variant, gene): which way the variant moves that gene's predicted expression, how many of the 371 tissue tracks agree, and how far it sits from the gene. The interval is the gene's MANE span widened by the model's measured +/-512 kb attribution horizon, unless --chrom/--start/--end supply one; either way the gene names the server-side filter, and the MANE lane is still consulted for the distance, which an explicit interval cannot supply.

A whole gene is ~3.3 M SNVs at the measured 1,091 SNVs/s — about 50 minutes — and the cost is printed before the query runs rather than discovered during it.

Argument Type Default Says
spec directory required Module spec directory
Option Type Default Says
--gene str required HGNC symbol. REQUIRED even with an explicit interval: the server-side gene filter is not an optimisation, and an unfiltered interval query is refused before it is sent.
--chrom str Contig of an explicit interval. Wins over the gene's MANE span.
--start int range 1-based start of that interval.
--end int range 1-based end of that interval.
--min-score float Keep only pairs whose magnitude reaches this. Distal scores run ~10x lower than scores at the gene, so a flat bar keeps the proximal rows and looks like it filtered on effect.
--max-rows int range 50000 Refuse rather than write more rows than this. Raising it is a deliberate act.
--offline flag No-op with a warning: this pass reads the Atlas, not a snapshot.
--dry-run flag Report what would be written without writing it.
--use str unstated Declared use recorded on the licence row: unstated|non-commercial|commercial. AlphaGenome Output is NON-COMMERCIAL ONLY, so an undeclared run writes nothing and says so — pass --use non-commercial.

just-dna-enricher alphagenome publish

Publish the AVI snapshot to HuggingFace.

Its own command because no other one can reach this lane (RM202). cache rebuild --publish walks lanes that have a rebuild adapter, and this lane cannot have one — its source is behind an eligibility gate, so there is nothing for an unattended rebuild to fetch. RM198 gave the lane a publish_repo and left it unreachable.

The upload is two commits and, above 5 GB, goes through the resumable uploader (RM199): the payload first, then release.json — the description must never arrive before the bytes it describes.

Argument Type Default Says
snapshot directory The built snapshot directory. Omit to use the resolved cache ($JUST_DNA_ALPHAGENOME_AVI_CACHE, then the cache base), falling back to where alphagenome build writes.
Option Type Default Says
--repo str Target HF dataset. Default: just-dna-seq/alphagenome_avi.
--dry-run flag Show what would be uploaded. Reads the repo's file list; sends nothing.
--message / -m str Commit message.

just-dna-enricher assertions

Fill clinical_assertions.csv from the coordinates already in resolution.csv.

Records what ClinVar says about each allele and how much review sits behind it — the star rating a compiled module previously discarded, so a one-star single submission and a practice guideline stopped being the same claim. Offline-capable: with a snapshot provisioned this pass never touches the network, and with none reachable it is a no-op rather than a failure.

It records; it does not adjudicate. Whether the module's own clin_sig agrees with ClinVar's is the enrich cross-check's question, and that one warns in both modes on purpose.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail unless every resolved allele has a ClinVar record.
--offline flag Snapshot only: never touch the network.
--clinvar-cache path Explicit ClinVar snapshot directory.

just-dna-enricher atlas

The AlphaGenome Atlas — precomputed variant scores over gRPC. Needs the [atlas] extra (grpcio + protobuf, 19 MB); the bindings are generated from pinned Apache-2.0 .proto sources rather than committed, so atlas generate runs once per checkout.

just-dna-enricher atlas generate

Fetch the pinned Atlas .proto sources and generate the gRPC bindings from them.

The sources are not vendored: the repository carries a commit id and a sha256 per file, and a file that does not match its pin is refused rather than used (RM196). Needs grpcio-tools, which is in \[dev] and deliberately not in \[atlas] — the runtime imports the bindings without it. A released wheel carries both the sources and the bindings already, so this is a checkout command.

Option Type Default Says
--refetch flag Re-download the pinned sources even if they are already on disk and match.

just-dna-enricher cache

Pre-provision, rebuild and report the snapshot caches.

just-dna-enricher cache prepare

Leave this machine with every cache it can have — pull what is published, build what is not.

The complement of cache pull, and the one command a deployment actually wants. pull fetches the published snapshots and stops; four lanes are not published for recorded reasons — PharmVar's personal key, PubMind's absent terms, NCBI's policy over MANE, ACMG's supplementary material — so a machine that only pulled is missing four caches and the checks that read them skip themselves. This runs each lane by the route it has.

The route is a property of the lane, never a flag. A published lane pulls, because building it would spend an operator's bandwidth re-deriving bytes somebody already made; an unpublished one builds, because that is the only route there will ever be. Asking for the choice would be asking an operator to restate the licensing story.

A cache that is already present is left alone, exactly as cache pull leaves one alone, so this is idempotent and cheap to re-run. Re-cutting a snapshot that exists is cache rebuild, which writes somewhere else on purpose — a build straight into a live cache is visible half-done to anything reading it, and a short parquet still has a footer.

The Python counterpart is just_dna_enricher.caches.prepare_caches, which this calls.

Option Type Default Says
--only str [] Prepare just these caches (repeatable). Default: every one.
--use str unstated Declared use: one of ['commercial', 'non_commercial', 'unstated'].
--pin str [] lane=release, repeatable, for the lanes that are built rather than pulled.
--source str [] lane=path, repeatable: build from a file you already hold.

just-dna-enricher cache prune

Say what a published snapshot repo carries that its lane is not made of, and offer to delete it.

Deletion is never a side effect of publishing, and this is the command that makes that affordable (RM186). A published repo accumulates: the publisher adds and does not remove, so a layout change leaves the old spelling in place, and just-dna-seq/clinvar still carries the 159 MB single-file clinvar.parquet from before the per-chromosome split. Provisioning already refuses to download it — the glob is what defends this tier — but any consumer globbing data/*.parquet, the dataset viewer included, still gets two schemas under one relation.

Nothing here is a sweep. A file is a candidate only if the lane's own glob excludes it or a LayoutShift declares it retired; README.md, .gitattributes, release.json, LICENSE.txt and sidecar directories are never touched. Without --yes this reads and prints and does nothing else, which is the mode to run first.

Option Type Default Says
--only str [] Prune just these caches (repeatable). Default: every published one.
--yes flag Delete without asking. Without it this prints the plan and stops.

just-dna-enricher cache pull

Download the published parquet snapshots from HuggingFace into the local caches.

The provisioning step a hosted deployment runs once, so no pass ever reaches a source live per request. Already-complete caches are trusted without touching the network, so this is re-runnable and cheap; a truncated file is removed and refetched.

A lane with nothing published says so and names its reason, which is a field on the registry rather than a comment: PharmVar's and PubMind's are refusals, ACMG's and MANE's are permissions nobody has established. Build those with cache rebuild.

Option Type Default Says
--only str [] Pull just these caches (repeatable). Default: every publishable one.
--use str unstated Declared use for the licence-gated snapshots. They forbid sale, so they are SKIPPED when unstated and REFUSED when commercial — downloading is taking the data.

just-dna-enricher cache rebuild

Rebuild every cache this tier builds — acquire, convert, and optionally publish (RM176).

The one endpoint over eleven builders. Each per-lane X build command stays, and this calls the same download_*/build_* functions they do, so there is one conversion algorithm with two callers rather than two that have to agree. What differs is only flag plumbing: a per-lane command offers the local-file inputs an operator holds, and a rebuild pass by definition holds none.

Every lane is built into <base>/<lane>/, never in place over a resolved cache. A rebuild takes minutes and an enrich reading a half-written snapshot mid-flight would see a real but incomplete table — the failure a resolver cannot detect, because a short parquet is still a parquet. Point the caches at the new base when the run is done, or copy each directory across.

An outcome is three-valued. ACMG needs a workbook that is Elsevier supplementary material, PharmVar a personal key, CIViC a release date to pin — none of those is a failure, and a nightly rebuild reporting errors for them would be reporting the licences working as designed. They are printed as not run, with the reason, and the exit code counts only real failures.

Option Type Default Says
--out directory data/caches Base directory. Each lane is built into //, never in place. The default is under data/, which this workspace git-ignores wholesale.
--only str [] Rebuild just these caches (repeatable). Default: every one that can be.
--use str unstated Declared use: one of ['commercial', 'non_commercial', 'unstated'].
--pin str [] lane=release, repeatable. e.g. --pin mane=1.5 --pin civic=2026-08-01.
--source str [] lane=path, repeatable: build from a file you already hold instead of downloading. Required for acmg; the offline off-switch for clinvar, constraint, clinpgx, drug_labels, pubmind and strchive. mane and civic take three files each and refuse it.
--publish flag Also upload each rebuilt snapshot to its HuggingFace repo.
--dry-run flag With --publish: show what would be uploaded, send nothing.

just-dna-enricher cache status

Say which snapshots are present, where, and which release each holds.

Reads only: nothing is downloaded, so this is safe on a machine with no network and it is the first thing to run when a pass reports that a source was skipped.

just-dna-enricher check-acmg

Check each row's acmg_sf against the ACMG secondary-findings list (reports only).

Writes no authored cell, and records that the question was put — the same two halves as check-identifiers. acmg_sf is an authored cell this asks a registry about, not a fact this pass contributes, and filling it here would break the check (see hints.REDUNDANCY_BEARING). The verification.json record is an attestation, never a value: it says the list was consulted and over how many rows, which is the one thing a downstream reader cannot reconstruct from the artifact (RM45/RM72).

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Exit 1 if any acmg_sf disagrees.
--offline flag No network. Needs --sf-list, else nothing is checked.
--url str https://www.ncbi.nlm.nih.gov/clinvar/docs/acmg/ ACMG secondary-findings page URL (fallback).
--sf-list directory Built ACMG SF snapshot (see acmg build). Preferred: NCBI's page still serves v3.2. Omit it and a snapshot in $JUST_DNA_ACMG_CACHE (or the shared cache base) is used; the page is scraped only when neither is there.

just-dna-enricher check-identifiers

Report obsolete trait terms, retired gene symbols and unrecognised PGS accessions (online).

Writes no authored cell, and records that the question was put. Unlike the rsID check (whose verdict lands on resolution.csv), these are module-level identifiers with no sidecar column to record, and filling one from the registry being asked about it would make the comparison vacuous — see hints.REDUNDANCY_BEARING. What this does write is verification.json: an attestation that the five checks ran and over how many rows, never a value. A consumer holding the artifact has no other way to tell "asked and clean" from "never asked" (RM45/RM72).

The PGS leg also writes sources.csv (RM163), and that is not an exception to the sentence above: the Catalog's license is a field on each score record and it varies, so a module carrying an academic-research-use-only score must not compile claiming the generic terms. The rows are the terms, never a value in an authored cell.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Exit 1 if any identifier is stale.
--traits / --no-traits flag on Check trait_efo_id against OLS4.
--genes / --no-genes flag on Check gene symbols against HGNC.
--pgs / --no-pgs flag on Check pgs_id against the PGS Catalog, and the two authored cells beside it.
--use str unstated Declared use: unstated | non-commercial | commercial. A PGS score licensed for academic research only bars sale, so a module citing one compiles ONLY with a declaration — and this flag is the one the compile's own refusal tells you to re-run with.

just-dna-enricher check-repeat-bands

Compare a module's repeat_alleles.csv bands against STRchive's, and report the differences.

Writes no authored cell and never fails on a difference. Where a catalogue and an expert author draw a repeat threshold in different places, both are claims by an authority, and a compile that refused would make this format pick the winner — the rule the ClinVar clin_sig and PGx allele-function checks already follow. --strict is accepted so the flag means one thing across the tier, and it changes nothing here but the mode recorded in the report.

The catalogue's pathogenic_max is reported as its own finding and is never written: it is the longest allele the literature records, not a clinical ceiling, and a module that imported it would silently answer nothing at all for a longer one.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--catalogue path Built STRchive snapshot directory (see strchive build), or a STRchive-loci.json.
--strict / --best-effort flag Carried into the report. A band difference NEVER fails, in either mode.

just-dna-enricher civic

Build the CIViC snapshot from a dated bulk release. CC0, so unlike the PubMind snapshot this one may be published.

just-dna-enricher civic build

Reduce a dated CIViC release to the parquet snapshot the direction-axis drafter reads.

The bulk release, not the GraphQL API, and the two are not interchangeable. Every row of ClinicalEvidenceSummaries.tsv is status accepted; the API defaults to NON_REJECTED and serves roughly 2.35x as many evidence items. A snapshot has to be reproducible from a pinned input, so this reads the dated files and records the basis in release.json.

--submitted widens that basis without leaving the dated release (RM169). CIViC publishes <date>-civic_accepted_and_submitted.vcf beside the TSVs, so unreviewed evidence is pinnable too and no API read is needed. The TSVs stay primary — the VCF cannot carry a variant with no GRCh37 coordinate, which is exactly the class whose identity had to be read out of its name — and the VCF supplies the curation status, the submitted evidence, and the identity for the 112 variants VariantSummaries.tsv (itself accepted-only) does not describe. Those rows are stamped identity_derivation="vcf_csq", and nothing is placed from the VCF's own GRCh37 position.

There is no --use flag. CIViC is CC0 on every axis, so a declared-use gate would permit every build unconditionally, and a flag feeding a gate that never gates is a flag that does nothing (@acquisition-gate-is-not-a-read-gate).

Option Type Default Says
--release str Dated CIViC release to download, e.g. 01-Aug-2026. A DATED release, never the nightly: a snapshot that cannot name its input is one nothing can reproduce.
--evidence file Local ClinicalEvidenceSummaries.tsv. Use instead of --release to build offline.
--variants file Local VariantSummaries.tsv.
--profiles file Local MolecularProfileSummaries.tsv.
--out directory data/repro/civic Output snapshot directory (writes data/civic.parquet + release.json).
--submitted flag Also read the release's civic_accepted_and_submitted.vcf, so evidence a curator entered but no editor signed off joins the snapshot. Over 01-Aug-2026 that widens the direction corpus from 507 rows on 270 variants to 1149 on 397, adding 642 submitted rows and 127 variants, and every row carries the status CIViC gave it. Dated and pinnable like the TSVs, so the build stays reproducible.
--vcf file Local civic_accepted_and_submitted.vcf. Use with the local TSV flags to build offline.

just-dna-enricher civic citations

Append the citations a CIViC variant carries that the dated bulk release cannot reach (RM160).

Why this is not part of civic build. The builder reads a dated release and is byte- reproducible from it; civic reproduce proves it by building twice. The wider basis RM169 adopted comes from a VCF, and a VCF record needs a POS — so submitted evidence attached to a variant with no GRCh37 coordinate is published on exactly one surface, the GraphQL API, which has no release to pin. The read is also one request per variant by construction, because evidenceItems takes a single variantId. Batching it into the builder is the first repair anyone proposes and it is exactly the reproducibility bargain this shape refused.

A recovered citation lands in studies.csv; literature.csv is derived from those PMIDs by the literature command, and an article row nothing cites is dropped from the artifact. CIViC's own curation status rides in confidence/confidence_unit, unconverted, so an accepted row and a submitted row are not the same row. Appending only — a second run over an unchanged API adds nothing — and enrich re-asks later and reports what has moved since.

CIViC is CC0, so there is no --use flag: a declared-use gate would permit every call unconditionally, and a flag feeding a gate that never gates is a flag that does nothing.

Argument Type Default Says
spec directory required Module spec directory.
Option Type Default Says
--snapshot directory CIViC snapshot to map authored rows through. Default: the provisioned cache.
--variant-id int [] Ask about a CIViC variant id directly, repeatable. Its citations ground the MODULE rather than a variant, which is the only route to a record CIViC publishes no identity for — variant 1955 is the case this exists for.
--offline flag Do not fetch. Every subject is recorded as not-asked; no row is written.
--dry-run flag Report what would be appended, write nothing.

just-dna-enricher civic publish

Upload a built CIViC snapshot to a HuggingFace dataset repo (publisher/dev).

This one does not refuse, and the contrast with pubmind publish is the point. CIViC's content is CC0 1.0 — a public-domain dedication with no share-alike, no bar on sale and attribution requested rather than required — so there is no permission to establish before passing the bytes on. PubMind's command exists in order to say no because its terms are unstated; PharmVar's cache is unpublishable because its terms forbid it. Nothing here is in either position.

What the snapshot carries is a derivation of CIViC's release, not a copy of it: the germline direction rows, placed on GRCh38 through identifiers CIViC itself publishes. release.json records which dated release it came from and the accepted status basis, so a consumer can tell what they are looking at without re-deriving it.

The ClinGen Allele Registry's answers are NOT in here, and that is deliberate rather than an oversight: the registry states no terms, and a lookup performed at draft time is a read, while baking its responses into a published file would be redistribution of bytes nobody has established we may pass on. The snapshot carries the CAID; resolving it stays the consumer's own fetch.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/civic.parquet + release.json).
Option Type Default Says
--repo str Target HF dataset (owner/name). Default: just-dna-seq/civic.
--message / -m str Commit message.
--dry-run flag Show what would be uploaded without contacting HuggingFace.

just-dna-enricher civic reproduce

Build the CIViC snapshot from a dated release and check it, end to end.

Five checks, and the third is the one worth the network. The first two are about us; the third is about whether the coordinates we produced are real.

  1. The release downloads and its bytes are recorded — a sha256 per file, so a rerun that disagrees is a finding about the source rather than a mystery.
  2. Two independent builds are byte-identical (Principle 7). A parquet has no inherent row order, so this is the check that the sort is doing its job.
  3. Every placed coordinate is cross-checked against the GRCh38 reference sequence. This is the external validation: the snapshot's positions come from RefSeq accessions inside ClinVar HGVS, and this asks an unrelated service (refget/seqrepo) whether the reference base at each of those positions is what we wrote. A wrong-build or off-by-one placement fails here and nowhere else.
  4. The drop registry closes — every input row kept or counted, an equality over a walked set.
  5. The published file list is exactly what the publisher would upload.

Exits non-zero if any check fails, so it is usable in CI.

Option Type Default Says
--release str 01-Aug-2026 Dated CIViC release to reproduce, e.g. 01-Aug-2026.
--out directory data/repro/civic_reproduce Working directory. The release files and two independent builds land here. The default is under data/, which this workspace git-ignores wholesale.
--keep flag Leave the downloaded release files in place for inspection.
--offline flag Skip the reference cross-check. The build and determinism checks still run.
--submitted flag Reproduce the wider basis: also download the release's civic_accepted_and_submitted.vcf and build with it, so the submitted rows and their coordinates go through every check below rather than only the accepted ones.

just-dna-enricher clinpgx

Build the ClinPGx clinical-annotation and drug-label snapshots, and cross-check against them.

just-dna-enricher clinpgx build

Download + build the ClinPGx snapshot (dev surface; needs polars).

Option Type Default Says
--out directory data/repro/clinpgx Snapshot output directory.
--zip path An existing summaryAnnotations.zip (else downloaded).
--url str https://api.clinpgx.org/v1/download/file/data/summaryAnnotations.zip ClinPGx bulk download URL.
--use str unstated Declared use: unstated | non-commercial | commercial.

just-dna-enricher clinpgx build-labels

Download + build the regulator drug-label snapshot (dev surface; needs polars).

A second archive from a source this tier already adopted, with its own release.json: ClinPGx publishes at least twelve downloads on this endpoint and they do not refresh in lockstep, so the label snapshot is dated from its own CREATED_*.txt rather than from the annotation lane's.

There is no --offline: a builder's off-switch is passing --zip instead of downloading.

Option Type Default Says
--out directory data/repro/drug_labels Snapshot output directory.
--zip file A drugLabels.zip you already have. Without it the archive is downloaded.
--url str https://api.clinpgx.org/v1/download/file/data/drugLabels.zip ClinPGx bulk download URL.
--use str unstated Declared use: one of ['commercial', 'non_commercial', 'unstated'].

just-dna-enricher clinpgx check

Cross-check pharm_variants.csv against the ClinPGx snapshot.

The snapshot no longer has to be handed over by hand (RM38): explicit path → $JUST_DNA_CLINPGX_CACHE / the default cache → downloaded from HuggingFace. --offline stops at the second step.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--snapshot path Explicit ClinPGx snapshot dir. Omit it and the cache is used, or one is downloaded.
--offline flag Use a local snapshot only: never download one.
--strict / --best-effort flag Fail on a stale evidence level.
--use str unstated Declared use: unstated | non-commercial | commercial.

just-dna-enricher clinpgx check-labels

Compare a module's gene/allele/drug claims against the drug labels five regulators publish.

Two join tiers, reported apart, and the tier belongs to the question. What do the agencies say about this gene and this medicine is the gene-tier subject; …and this star allele or rsID is the allele-tier one. A label naming both answers both, because they are two questions rather than one asked twice, and a gene-level agreement is not an allele-level agreement.

Writes no authored cell and never fails on a difference. Five agencies genuinely disagree with each other — clopidogrel and CYP2C19 is Actionable PGx at four of them and Informative PGx at the EMA — and a compile that refused would make this format pick the winner. A blank Testing Level is a third of the file and is reported as unknown, never as No Clinical PGx.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--snapshot directory Built drug-label snapshot directory (see clinpgx build-labels). Omit it and $JUST_DNA_DRUG_LABELS_CACHE (or the shared cache base) is used.
--strict / --best-effort flag Carried into the report. A label difference NEVER fails, in either mode.
--use str unstated Declared use: one of ['commercial', 'non_commercial', 'unstated'].

just-dna-enricher clinpgx publish

Publish a built ClinPGx snapshot so clinpgx check can provision it (publisher/dev).

LICENSE.txt travels with the parquet — the terms ClinPGx ships inside its own archive are what license_sha256 pins, and a published snapshot without them pins nothing for whoever downloads it.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/*.parquet + release.json).
Option Type Default Says
--repo str just-dna-seq/clinpgx Target HuggingFace dataset repo (owner/name).
--dry-run flag Show what would be uploaded; send nothing.
--message / -m str Commit message.

just-dna-enricher clinpgx publish-labels

Publish a built drug-label snapshot so clinpgx check-labels can provision it (publisher/dev).

A second repo rather than a second table in just-dna-seq/clinpgx, for the reason the builder already gives its own release.json: the two ClinPGx archives do not refresh in lockstep, and one repo holding both would date the pair from whichever was published last.

Same grounds as clinpgx publish — CC BY-SA permits redistribution, forbids sale, and requires attribution, which sources.csv carries. LICENSE.txt travels with the parquet: a share-alike snapshot whose terms did not travel pins nothing for whoever downloads it.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/drug_labels.parquet + LICENSE.txt + release.json).
Option Type Default Says
--repo str just-dna-seq/clinpgx_drug_labels Target HuggingFace dataset repo (owner/name).
--dry-run flag Show what would be uploaded; send nothing.
--message / -m str Commit message.

just-dna-enricher clinvar

Build and publish the ClinVar reference snapshot (publisher/dev surface).

just-dna-enricher clinvar build

Convert a ClinVar VCF into the per-chromosome parquet snapshot the resolver reads.

Option Type Default Says
--vcf file Local ClinVar VCF (.vcf.gz). Omit and pass --download to fetch from NCBI.
--download flag Download the NCBI ClinVar GRCh38 VCF into --out first.
--out directory data/repro/clinvar Output snapshot directory (writes data/*.parquet + release.json).

just-dna-enricher clinvar citations

Add ClinVar's literature links to a snapshot: data/citations.parquet ([dev], needs polars).

Separate from clinvar build because ClinVar publishes citations separately from the VCF — which is precisely why a drafted gene panel could not compile without this: studies.csv is mandatory and the VCF carries no PMIDs. Written beside the snapshot, so an existing cache keeps its bytes.

Option Type Default Says
--out directory required Existing ClinVar snapshot dir.
--citations file Local var_citations.txt.
--download flag Fetch var_citations.txt first.
--url str https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt Source for --download.

just-dna-enricher clinvar publish

Create-or-update the dataset repo and upload the built ClinVar snapshot (publisher/dev).

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/*.parquet + release.json).
Option Type Default Says
--repo str Target HF dataset (owner/name). Default: just-dna-seq/clinvar.
--message / -m str Commit message.
--dry-run flag Show what would be uploaded. Reads the repo's file list; uploads nothing.

just-dna-enricher cpic

Build and publish the CPIC snapshot, so a hosted enricher never reaches CPIC per request.

just-dna-enricher cpic build

Fetch CPIC whole into data/*.parquet + release.json (dev surface; needs polars).

No gene filter, deliberately: the whole database is ~120k narrow rows, and a snapshot covering only the genes the operator thought of answers "CPIC has nothing" for the next one.

Option Type Default Says
--out directory data/repro/cpic Snapshot output directory.
--endpoint str https://api.cpicpgx.org/v1 CPIC PostgREST base URL.
--use str unstated Declared use: unstated | non-commercial | commercial.

just-dna-enricher cpic publish

Create-or-update the dataset repo and upload the built CPIC snapshot (publisher/dev).

Publishable because CPIC's recorded terms permit redistribution — CC BY-SA grants sharing under share-alike plus attribution, which sources.csv carries. PharmVar has no equivalent command, and that is the design rather than an omission.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/*.parquet + release.json).
Option Type Default Says
--repo str just-dna-seq/cpic Target HuggingFace dataset repo (owner/name).
--dry-run flag Show what would be uploaded; send nothing.
--message / -m str Commit message.

just-dna-enricher dosage

Add ClinGen dosage-sensitivity rows to gene_metrics.csv (haploinsufficiency/triplosensitivity).

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail unless every gene is ClinGen-curated.
--offline flag No-op with a warning: ClinGen's curation list is a live download with no snapshot.
--url str https://ftp.clinicalgenome.org/ClinGen_gene_curation_list_GRCh38.tsv ClinGen gene-curation list URL.
--use str unstated Declared use: unstated | non-commercial | commercial. ClinGen is CC0, so no declaration is refused here — it is recorded into sources.csv beside the rows it justifies.

just-dna-enricher draft

Draft PGx tables for one or more genes from CPIC — appends rows, never overwrites one.

Re-runnable and additive, so a multi-gene module is built up a gene at a time. A row whose key is already in the file is reported, never replaced: what CPIC now says about a row you already wrote is a finding for pgx, not an edit for this command to make.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--gene str required Gene to draft from CPIC (repeatable).
--drug str [] Also draft CPIC's prescribing recommendations for this drug (repeatable).
--allele str [] Draft only these star alleles, in all three tables (repeatable; *1 is always kept). A caller emits a bounded allele set, and n alleles is n(n+1)/2 pairs — CYP2D6 is 16,290 diplotypes unfiltered. Requires a single --gene, since a star name is gene-scoped.
--population str Draft only this CPIC clinical context (e.g. 'NVI'). Default: every context, as rows.
--use str unstated Declared use: unstated | non-commercial | commercial. CPIC forbids sale, so a draft is SKIPPED when unstated and REFUSED when commercial.
--offline flag Draft from a built CPIC snapshot only; never reach CPIC live.
--cpic-cache path Explicit CPIC snapshot dir.
--dry-run flag Report what would be added; write nothing.

just-dna-enricher draft-clinpgx

Draft pharm_variants.csv rows from the ClinPGx snapshot — appends, never overwrites a row.

Narrow with --drug and re-run as the module grows. A row already in the file is reported, never replaced: drift against ClinPGx is clinpgx check's finding, not this command's edit to make.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--snapshot directory required Built ClinPGx snapshot (see clinpgx build). Inject-only; nothing is downloaded.
--drug str [] Only annotations naming this drug (repeatable).
--gene str [] Only annotations naming this gene (repeatable).
--min-evidence-level str Keep annotations at least this strong: 1A|1B|2A|2B|3|4.
--use str unstated Declared use: unstated | non-commercial | commercial. ClinPGx forbids sale.
--dry-run flag Report what would be added; write nothing.

just-dna-enricher draft-panel

Draft a gene panel's variants.csv rows from an authority — appends, never overwrites a row.

The drafted rows carry a genotype placeholder, so the module will not compile until you decide what each finding is about. That is deliberate: ClinVar publishes alleles, and whether carrying one is a carrier state or an affected one follows from the condition's inheritance mode, which the source does not say. Rows land in their gene's block, and a re-run leaves anything already there — stub or filled — exactly as it is.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--gene str [] Gene to draft rows for (repeatable). Required for every source but mitomap-miss, whose increment is asked for as a whole and where --gene only filters.
--source str clinvar Which authority to draft the calls from: clinvar (the default); pubmind — an LLM's reading of the literature, which needs an operator-built snapshot and still reads the ClinVar one for its gene attribution; civic — curated cancer interpretations, which writes the DIRECTION axis rather than clin_sig and needs a civic build snapshot; or mitomap-miss — the curated mtDNA calls MITOMAP publishes and the ClinVar cache does not, which needs a mitomap miss snapshot.
--mitomap-miss-cache directory Built MITOMAP-miss snapshot (see mitomap miss). Only read under --source mitomap-miss; omit it and $JUST_DNA_MITOMAP_MISS_CACHE is used.
--civic-cache directory Built CIViC snapshot (see civic build). Only read under --source civic; omit it and $JUST_DNA_CIVIC_CACHE is used.
--snapshot directory Built ClinVar snapshot (see clinvar build). Omit it and the cache is used, or the published snapshot downloaded — the citations table comes with it, which is what a panel needs to compile. Read for its gene attribution under --source pubmind, which publishes no gene column of its own.
--pubmind-cache directory Built PubMind snapshot (see pubmind build), for --source pubmind. Omit it and $JUST_DNA_PUBMIND_CACHE is read; there is no published one to download.
--offline flag Use a local snapshot only: never download one.
--download / --no-download flag on Provision the published snapshot when no local one is found. Fetching it is this command's only network use, so --no-download coincides with --offline today; it is a separate switch because it says 'do not go and get one', not 'make no request'.
--clin-sig str Comma-separated calls to include. Default: pathogenic,likely_pathogenic.
--min-review-stars int range 2 Review-status floor, --source clinvar only. 2 = multiple submitters, no conflicts.
--max-citations int range 3 Study rows to draft per variant from ClinVar's literature links. 0 disables. --source clinvar only: PubMind's channel carries no PMID.
--min-confidence int range 1 Evidence-depth floor, --source pubmind only. PubMind's confidence counts how much of the literature spoke, 0-3; 1 means more than a single mention.
--use str unstated Declared use (ClinVar is public domain).
--dry-run flag Report what would be added; write nothing.

just-dna-enricher draft-repeats

Draft repeat_alleles.csv identity rows from STRchive — appends, never overwrites a row.

The bands are not drafted, and that is the design rather than a limitation. A drafted row carries the gene, the motif as the catalogue spells it, the trait CURIE where the locus names exactly one disease, and a conclusion placeholder — so the table cannot compile until a human has filled in what each band means. measure_min/measure_max stay empty: run check-repeat-bands once you have written them and it will report where the catalogue disagrees.

The catalogue's coordinates, ref_copies and locus_structure have no authored column to land in; the run counts them and says so rather than dropping them silently.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--gene / -g str [] Restrict to these genes. Repeatable; omit for every catalogue locus.
--catalogue path Built STRchive snapshot directory (see strchive build), or a STRchive-loci.json. Omit it and $JUST_DNA_STRCHIVE_CACHE (or the shared cache base) is used.
--use str unstated Declared use: one of ['commercial', 'non_commercial', 'unstated'].
--dry-run flag Report what would be added; write nothing.

just-dna-enricher enrich

Resolve a spec's variants into resolution.csv beside the spec. Exit 1 in strict mode if unresolved.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail unless every variant resolves.
--offline flag Cache-only: never touch the network.
--ensembl-cache path Explicit Ensembl cache dir/.duckdb.
--clinvar-cache path Explicit ClinVar snapshot dir.
--pubmind-cache path Built PubMind snapshot dir (from pubmind build) — the second authority in the clinical-significance concordance check. Omit it and $JUST_DNA_PUBMIND_CACHE is read; with neither, PubMind's leg reads unchecked rather than agreement.
--clinvar / --no-clinvar flag on Use the ClinVar link (after the Ensembl cache).
--gnomad / --no-gnomad flag on Use the gnomAD link (last, after live Ensembl).
--vrs / --no-vrs flag on Mint GA4GH VRS allele ids onto resolved rows.
--verify-ref / --no-verify-ref flag on Check each authored ref against the reference sequence and report disagreements.
--verify-clinsig / --no-verify-clinsig flag on Check each authored clin_sig against the ClinVar snapshot's own (warns, never fails).
--verify-rsids / --no-verify-rsids flag on Check each authored rsID against dbSNP for merges/withdrawals (online only).
--verify-datasets / --no-verify-datasets flag on Check each release recorded in sources.csv against the one that source publishes now, and report the gap. One request per source, and the cheap question to put before --rederive: it tells you whether re-asking every subject is worth the run.
--keep-par-twin flag Record both contigs of a pseudoautosomal locus. Default keeps only the X spelling, which is the one every annotation source uses and the only one a hard-masked GRCh38 analysis set can match.
--rederive flag Re-ask every source about every subject, including the ones already recorded, and report which of them changed value. An ordinary run gap-fills and never re-asks, so a source that quietly revised an answer moves nothing you could notice.
--keep-staging flag Leave the staged answers beside resolution.csv after a successful run. They are removed by default; a killed run leaves them either way, and the next run resumes from them.

just-dna-enricher enrich-and-compile

Enrich, then compile from the produced resolution.csv (offline, deterministic). Exit 1 on failure.

Argument Type Default Says
spec_dir directory required Module spec directory
output_dir directory required Output dir for parquet + manifest.json
Option Type Default Says
--strict / --best-effort flag Fail unless every variant resolves.
--offline flag Cache-only: never touch the network.
--ensembl-cache path Explicit Ensembl cache dir/.duckdb.
--clinvar-cache path Explicit ClinVar snapshot dir.
--clinvar / --no-clinvar flag on Use the ClinVar link (after the Ensembl cache).
--gnomad / --no-gnomad flag on Use the gnomAD link (last, after live Ensembl).
--frequencies flag Also run the frequency pass (writes frequencies.csv).
--gene-metrics flag Also run the gene-constraint pass (writes gene_metrics.csv).

just-dna-enricher frequencies

Fill frequencies.csv from the coordinates already in resolution.csv (pass 2, online only).

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail unless every resolved allele has a frequency.
--offline flag No-op with a warning: gnomAD frequency has no offline snapshot.
--populations str Comma-separated ancestry groups to keep (e.g. 'global' for one row per allele). Default: all.
--dataset str Override the dataset label recorded on each row.

just-dna-enricher gene-metrics

Fill gene_metrics.csv for the genes variants.csv mentions (pass 3, snapshot then live API).

With no local snapshot the v4.1 one is downloaded from HuggingFace first, exactly as enrich provisions the Ensembl and ClinVar snapshots — --offline is what turns that off, and then the pass is snapshot-only. Reaching the live API instead means v2.1.1 numbers, which the row's dataset records; provisioning is what keeps a plain install on v4.1.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail unless every gene has constraint metrics.
--offline flag Snapshot only: never touch the network.
--constraint-cache path Explicit gnomAD constraint snapshot dir.

just-dna-enricher gene-validity

Fill gene_validity.csv with curated gene-disease assertions for the genes variants.csv names.

One row per (gene, disease, mode of inheritance, submitter) — the source's own grain. Mode of inheritance is in the key because 59 ClinGen (gene, disease) pairs carry two curations that differ only there, and submitter is in it because GenCC publishes the disagreement between submitters, which is the thing it exists to publish.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--source str clingen Which submitter to read: clingen (expert panels) or gencc (an aggregate of nineteen).
--strict / --best-effort flag Fail unless every gene carries a curated assertion.
--offline flag No-op with a warning: neither ClinGen nor GenCC publishes an offline snapshot.
--url str Override the submitter's export URL.

just-dna-enricher gnomad

Build and publish the gnomAD gene-constraint snapshot (publisher/dev surface).

just-dna-enricher gnomad constraint

The gene-level constraint snapshot the gene-metrics pass reads offline.

just-dna-enricher gnomad constraint build

Reduce the per-transcript constraint TSV to the gene-level parquet the resolver reads.

Option Type Default Says
--tsv file Local gnomAD constraint metrics TSV. Omit and pass --download to fetch it.
--download flag Download the gnomAD v4.1 constraint TSV (95.5 MB) into --out first.
--out directory data/repro/gnomad_constraint Output snapshot directory (writes data/gnomad_constraint.parquet + release.json).

just-dna-enricher gnomad constraint publish

Create-or-update the dataset repo and upload the built constraint snapshot (publisher/dev).

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/*.parquet + release.json).
Option Type Default Says
--repo str Target HF dataset (owner/name). Default: just-dna-seq/gnomad_constraint.
--message / -m str Commit message.
--dry-run flag Show what would be uploaded without contacting HuggingFace.

just-dna-enricher gwas

Fill gwas_effects.csv with the GWAS Catalog's published effect sizes for this module's rsIDs.

One row per published association, not per variant — a well-studied variant carries dozens across different traits and papers. It does NOT fill weight: an authored weight is the author's model of the finding, and no tool writes one. The two sit side by side and a consumer picks.

Reads effect_unit verbatim, including the Catalog's uninformative "unit", because a beta whose scale is unknown must not look like one whose scale is shared. An association the Catalog published without establishing which allele carries the effect keeps a null effect_allele and is counted in the manifest, never dropped.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Severity ladder for findings; see the pass docstring.
--offline flag No-op with a warning: this pass reads the REST API, not a snapshot.
--use str unstated Declared use recorded on the licence row: unstated|non-commercial|commercial.
--study-facts / --no-study-facts flag on Follow each association's study and trait links. Costs 2 requests per association; measured at 382 requests for one real module. Off keeps effects, drops pmid/trait/ancestry PERMANENTLY for the rows it writes: the merge is keyed on association_id, so a later run with study facts on skips those rows rather than back-filling. Delete gwas_effects.csv to re-derive them.

just-dna-enricher hint

Look up what is known about a variant or citation. Writes nothing.

just-dna-enricher hint citation

Does this citation exist, and is it the paper you meant?

A paywall hides the fulltext, never the PubMed record, so existence is answerable for paywalled work; Crossref covers what PubMed does not index at all. Every answer is tri-state — unknown means the registry could not be asked, which is not the same as "no such paper".

Existence is not identity. PMIDs are densely allocated, so a recalled or invented number is very likely to be a real record for a different article, and pmid_exists alone cannot catch a fabricated citation. The title, journal, year and first author come back in the same response and are printed for exactly that comparison (S12).

--pmcid goes the other way. Every pmid in the schema keys on the PubMed id — studies.csv, a binning row's and a pharm_variants.csv row's alike — and a curator holding only a PMC… id had no route to it: the schema refused the cell and named no remedy. This resolves it and then asks PubMed which paper that is. The id is reported, never written: filling pmid from NCBI would make the existence check compare NCBI with itself.

Option Type Default Says
--pmid str PubMed id to check.
--doi str DOI to check (the one you authored).
--pmcid str PubMed Central id (PMC…) to resolve to the PubMed id tables key on.
--offline flag Skip the check and say so.
--json flag Emit the full machine answer.

just-dna-enricher hint gene

Is this gene symbol approved or retired?

Argument Type Default Says
symbol str required Gene symbol, e.g. MTHFR

just-dna-enricher hint recover

Which rs-number GRCh37 dbSNP records at an hg19/GRCh37 coordinate.

For a paper that predates GRCh38. Author the rs-number, not a converted position: an rs-number resolves into a coordinate the compiler can cross-examine, where a lifted-over position becomes the row's only witness to itself. Nothing is written — the rs-number is the row's identity, and a machine filling one migrates variant_key with no authored edit anywhere.

Option Type Default Says
--chrom str required Chromosome of the old coordinate.
--start int required 1-based GRCh37 position.
--ref str Reference allele, to narrow the answer.
--alts str Alt allele(s), comma-separated.
--offline flag Skip the lookup and say so.
--json flag Emit the full machine answer.

just-dna-enricher hint trait

Is this trait id current, obsolete, or unknown?

Argument Type Default Says
curie str required Trait CURIE, e.g. EFO_0004340

just-dna-enricher hint variant

Validity, coordinates, alleles, populations and clinical calls for one variant.

Nothing is decided for you: a one-to-many rsID returns every locus and a position matching several rsIDs returns every candidate. The coordinate is reported, never written into variants.csv — resolution puts it in resolution.csv, which is where it belongs.

Option Type Default Says
--rsid str dbSNP id to look up.
--chrom str Chromosome (with --start).
--start int 1-based position (with --chrom).
--ref str Reference allele, for an allele-exact lookup.
--alts str Alt allele(s), comma-separated.
--ambiguity flag Warn when the answer is not unique.
--frequencies flag Add gnomAD populations (paced: ~6s).
--offline flag Snapshots only; never touch the network.
--ensembl-cache path Explicit Ensembl cache.
--clinvar-cache path Explicit ClinVar snapshot.
--pubmind-cache path Explicit PubMind snapshot (see pubmind build); $JUST_DNA_PUBMIND_CACHE otherwise.
--json flag Emit the full machine answer.

just-dna-enricher literature

Fill literature.csv from a module's citations (pass 4, online only).

studies.csv is one citation site of several: a pmid on a binning row grounds the threshold it sits on, and one on a pharm_variants.csv row grounds that row's own drug and genotype claim. The pass reads every site, so a module citing only from those tables is enriched rather than refused.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail if a cited PMID does not resolve.
--offline flag No-op with a warning: there is no offline PubMed snapshot.
--fulltext / --no-fulltext flag on Also match provenance quotes against fulltext, falling back to the abstract.
--doi / --no-doi flag on Also confirm the authored DOI resolves in Crossref (covers preprints/books).

just-dna-enricher litvar

Which papers a variant-literature index holds for a module's alleles, and at which tier. Reports only; writes no authored cell and no table row.

just-dna-enricher litvar coverage

Report LitVar's literature coverage per locus, naming the tier that answered.

It answers which papers discuss an allele that is already identified. It does not answer which allele a name meant — those read as the same question and are not. Measured against the two hardest records in this repository (CIViC 1955 and 2131, four candidate alleles with registered CAIDs), the index returns no node for any of them, because PubTator3 mines titles and abstracts and those alleles live in a table inside a paywalled paper. Do not reach for this to recover an identity.

Writes no row and no sources.csv entry: nothing here reaches a module's tables, so the module does not use this source. What it does write is verification.json — an attestation that the question was put, over how many loci, and at which tier each was answered.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--offline flag No network; every locus is recorded as unchecked.
--quiet flag Only the tier summary, not a line per locus.

just-dna-enricher litvar gene

List every node LitVar holds under a gene symbol, grouped by tier. Writes nothing.

This is the endpoint that serves line-delimited Python repr() rather than JSON, and the tier split is the reason to look: of 588 HFE nodes on 2026-09-01, 220 are rsID-only, 69 carry a ClinGen allele id, exactly one is the gene node, and the remaining 298 are unnormalized protein strings — text a miner saw, not an identity anything should join on.

Argument Type Default Says
gene str required Gene symbol, e.g. HFE

just-dna-enricher mane

Build the MANE transcript snapshot — the numbering frame, cached and pinned instead of downloaded by hand and cited in prose.

just-dna-enricher mane build

Reduce one pinned MANE release to the three parquet tables the numbering frame needs.

All three files, in one pass. The summary is the frame; changed_select_accessions is the currency check and Update_Affects_CDS is the numbering-frame axis stated by the source; the negative roster is what makes "MANE has no answer for this gene" distinguishable from "nobody asked", with the reason attached. Shipping the cache without the thing that notices it going stale is the defect this command exists to close.

There is no --offline flag: the off-switch is passing the local files instead of --download. And there is no --use flag, because NCBI states a policy rather than a licence — every gating axis is unknown, and check_declared_use returns a skip for an unknown whatever the declaration says, so the gate would silently skip every build. A flag feeding a gate that never gates is a flag that does nothing (@acquisition-gate-is-not-a-read-gate).

MANE is the default, not the answer. A gene with two rows carries two CDS numbering frames and MANE_status says which; a gene with one row says nothing about the isoforms MANE does not carry, so a pass treating this table as an oracle would be wrong in a way the table cannot report.

Option Type Default Says
--download flag Fetch the release from NCBI. Without --release the newest version is discovered from current/README_versions.txt and then pinned to its versioned directory.
--release str MANE version to pin, e.g. 1.5. Resolves to release_/, never current/.
--summary file Local MANE.GRCh38.v.summary.txt.gz. Use instead of --download to build offline.
--changed file Local MANE.GRCh38.v.changed_select_accessions.txt.gz.
--not-in-mane file Local MANE.GRCh38.v.protein_coding_genes_not_in_mane.txt.gz.
--versions file Local README_versions.txt. Optional, and the only way an offline build can name its release: a filename is never parsed for one.
--out directory data/repro/mane Output snapshot directory (writes data/*.parquet + release.json).

just-dna-enricher mitomap

Build the MITOMAP snapshot and the derived miss lane. CC BY 3.0 with commercial use stated free, so a deployment may publish the snapshot.

just-dna-enricher mitomap build

Cut the two curated mtDNA variant tables, their citations and the references out of the dump.

602 mmutation rows and 494 rtmutation rows out of 6.76 million lines. The snapshot records every count this build computes — rows per table, the dump's own per-table edit dates, how much of reference.nlmid is a PMID, the alleles that cannot be spelled as VCF and the brackets that are not a documented VCEP class — because a number computed and dropped is one every reader has to recompute.

Option Type Default Says
--out directory data/repro/mitomap Output snapshot directory (writes data/mitomap-*.parquet + release.json).
--dump file A mitomap.dump.sql.gz you already have. Without it the dump is downloaded — the data surface answers plain curl, unlike the web surface. A local dump carries no Last-Modified, so that snapshot is honestly unlabelled.
--url str https://mitomap.org/downloads/mitomap.dump.sql.gz Source URL for the dump (used only when --dump is absent).

just-dna-enricher mitomap miss

Join MITOMAP against the ClinVar chrMT parquet and write the increment (RM171).

A derived lane, not a download. Its acquire stage is both parents being on disk, and a parent that is absent is reported as could-not-run rather than as an empty increment — a miss set computed without ClinVar would say MITOMAP publishes a thousand alleles nobody else has, from a comparison that never ran.

Exact (start, ref, alt) on chrMT, upper-cased both sides, no position-level fallback. Four buckets, and draft-panel --source mitomap-miss writes only one of them.

Option Type Default Says
--out directory data/repro/mitomap_miss Output snapshot directory (writes data/mitomap_miss.parquet + release.json).
--mitomap-cache directory Built MITOMAP snapshot (see mitomap build). Omit it and $JUST_DNA_MITOMAP_CACHE is used.
--clinvar-cache directory Built ClinVar snapshot (see clinvar build). Omit it and $JUST_DNA_CLINVAR_CACHE is used.

just-dna-enricher mitomap publish

Create-or-update the dataset repo and upload the built MITOMAP snapshot (publisher/dev).

Publishable on the source's own terms: CC BY 3.0, with commercial and clinical use stated free and attribution the one condition — which the snapshot's SourceRow carries. Only the parent lane publishes. The derived miss snapshot pins two parent digests, so a pulled copy would be an increment whose own currency check cannot be run by whoever pulled it; it is rebuilt locally from the parents instead, which is cheaper than the download and cannot be stale.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (data/ + release.json).
Option Type Default Says
--repo str just-dna-seq/mitomap Target HuggingFace dataset repo (owner/name).
--dry-run flag Show what would be uploaded; send nothing.
--message / -m str Commit message.

just-dna-enricher pgx

Cross-check star-allele tables against PharmVar/CPIC and record terms into sources.csv.

Snapshot first, live second (RM38). A built snapshot serves the check without egress and without spending a shared per-IP budget; --offline says snapshot-only, and a leg with neither is skipped with a reason rather than silently passing.

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--strict / --best-effort flag Fail on an allele-function discrepancy.
--offline flag Snapshots only: never reach PharmVar or CPIC live.
--use str unstated Declared use: unstated | non-commercial | commercial. Sources that forbid sale are SKIPPED when unstated and REFUSED when commercial.
--pharmvar / --no-pharmvar flag on Consult PharmVar (needs PHARMVAR_API_KEY).
--cpic / --no-cpic flag on Consult CPIC (open, no key).
--cpic-cache path Explicit CPIC snapshot dir.
--pharmvar-cache path Explicit PharmVar snapshot dir.

just-dna-enricher pharmvar

Build the PharmVar snapshot with your own key. Operator-built and inject-only; never published.

just-dna-enricher pharmvar build

Fetch PharmVar whole into data/*.parquet + release.json (dev surface; needs polars + a key).

There is no pharmvar publish, and there will not be. The data is pulled under a key PharmVar's terms §2 make personal and non-transferable, and no axis SourceTerms records covers passing that on — an unestablished permission is not a permission. Build your own; point at it with $JUST_DNA_PHARMVAR_CACHE or --pharmvar-cache.

Option Type Default Says
--out directory data/repro/pharmvar Snapshot output directory.
--use str unstated Declared use: unstated | non-commercial | commercial.

just-dna-enricher pubmind

Build the PubMind literature-derived snapshot. Operator-built and inject-only; never published.

just-dna-enricher pubmind build

Reduce the ANNOVAR-distributed PubMind table to the parquet snapshot the checks read.

There is no --use flag, and its absence is the design. PubMind's data terms could not be established, and unknown terms warn rather than gate (commercial_use=None never taints a module). A declared-use gate here would refuse every build unconditionally, and a flag feeding a gate that never gates is a flag that does nothing. What the unknown terms do gate is publishing a snapshot or a module carrying these bytes — see pubmind publish, which refuses.

Option Type Default Says
--table file Local hg38_pubmind_db.txt.gz. Omit and pass --download to fetch it from ANNOVAR.
--download flag Download the ANNOVAR-distributed PubMind table into --out first.
--out directory data/repro/pubmind Output snapshot directory (writes data/pubmind.parquet + release.json).

just-dna-enricher pubmind publish

Refuse to publish the PubMind snapshot, and say why.

The command exists in order to refuse. A missing one reads as an oversight somebody will helpfully add, and the reason belongs where a reader looks for it rather than only in a design document.

just-dna-enricher strchive

Build the STRchive repeat-locus snapshot. MIT-licensed, so a deployment may publish it.

just-dna-enricher strchive build

Fetch (or copy in) the STRchive catalogue and record its provenance beside it.

Pin a release: the default branch moves, so a comparison whose reference is "whatever was there that afternoon" cannot be re-run, and only a pinned build gets a dataset label the verification record can name.

Option Type Default Says
--out directory data/repro/strchive Output snapshot directory (writes STRchive-loci.json + release.json).
--catalogue file A STRchive-loci.json you already have. Without it the file is downloaded.
--release str Upstream release tag to pin, e.g. v2.26.0. Without it, the default branch, unlabelled.

just-dna-enricher strchive publish

Create-or-update the dataset repo and upload the built STRchive catalogue (publisher/dev).

Publishable on the source's own terms: STRchive is MIT, which grants redistribution outright. This lane had no publish command because it grew from a check rather than from a cache, not because anything withheld the permission — the same distinction the roster draws between CIViC's absent ensure_* (a gap) and PharmVar's (a refusal).

Publish a pinned build. An unlabelled snapshot carries no dataset, so whoever pulls it can run the comparison and cannot say which release they compared against — build with --release first, and this refuses nothing but says so.

Argument Type Default Says
snapshot_dir directory required Built snapshot directory (STRchive-loci.json + release.json).
Option Type Default Says
--repo str just-dna-seq/strchive Target HuggingFace dataset repo (owner/name).
--dry-run flag Show what would be uploaded; send nothing.
--message / -m str Commit message.

just-dna-enricher template

Print a header-only CSV for one authored table kind, generated from the live models.

Kept working here, but just-dna-compiler template is canonical: this needs no network, and an author who installed only the tier that owns the CSV shape should not have to add the network tier to get a header. See just-dna-compiler stub for a template with rows to replace.

Argument Type Default Says
kind str required Authored CSV to emit a header for, e.g. repeat_alleles.csv

just-dna-enricher upload

Upload a compiled module to a HuggingFace dataset collection (publisher/dev surface).

Writes data//, which keeps meaning "latest", and — when the manifest states a version — data//v/ under it, in that order, as two commits. With no version, the flat path alone, and the reason why.

Refuses when the versioned path already holds a different artifact (compare by artifact.digest), unless --force. The flat path means latest and is overwritten either way.

Argument Type Default Says
module_dir directory required Compiled module directory (at least one annotation parquet + manifest.json).
Option Type Default Says
--repo str Target HF dataset (owner/name). Default: just-dna-seq/annotators.
--name str Module name under data// (and data//v/) in the repo. Default: the directory basename.
--message / -m str Commit message. Default: 'Add module'.
--dry-run flag Show what would be uploaded without contacting HuggingFace.
--force flag Overwrite data//v/ even when it already holds a different artifact. Without this the publish refuses; the flat path is always overwritten.

just-dna-enricher vrs

GA4GH VRS allele identity for an already-resolved module.

just-dna-enricher vrs mint

Stamp ga4gh:VA.… allele ids onto resolution.csv (substitutions offline, indels online).

Argument Type Default Says
spec_dir directory required Module spec directory
Option Type Default Says
--offline flag Substitutions only: indels need the reference sequence, which means a network call.