just_dna_enricher.cli¶
just_dna_enricher.cli ¶
Command-line front door for the enricher (Typer) — the network tier's user-facing command.
just-dna-enricher enrich spec/ --strict --offline just-dna-enricher frequencies spec/ # pass 2: allele frequency (online only) just-dna-enricher gene-metrics spec/ # pass 3: gene constraint (offline capable) just-dna-enricher literature spec/ # pass 4: citations (online only) just-dna-enricher gene-validity spec/ --source gencc # curated gene-disease assertions (online) just-dna-enricher assertions spec/ # ClinVar call + review tier (offline capable) just-dna-enricher enrich-and-compile spec/ out/ --frequencies --gene-metrics just-dna-enricher gnomad constraint build --download --out gnomad_constraint/ # [dev] just-dna-enricher upload out/coronary --repo just-dna-seq/annotators # [dev]
enrich_ ¶
enrich_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every variant resolves.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Cache-only: never touch the network.",
),
ensembl_cache: Path | None = typer.Option(
None,
"--ensembl-cache",
help="Explicit Ensembl cache dir/.duckdb.",
),
clinvar_cache: Path | None = typer.Option(
None,
"--clinvar-cache",
help="Explicit ClinVar snapshot dir.",
),
pubmind_cache: Path | None = typer.Option(
None,
"--pubmind-cache",
help="Built PubMind snapshot dir (from `pubmind build`) — the second authority in the clinical-significance concordance check. Omit it and $JUST_DNA_PUBMIND_CACHE is read; with neither, PubMind's leg reads unchecked rather than agreement.",
),
use_clinvar: bool = typer.Option(
True,
"--clinvar/--no-clinvar",
help="Use the ClinVar link (after the Ensembl cache).",
),
use_gnomad: bool = typer.Option(
True,
"--gnomad/--no-gnomad",
help="Use the gnomAD link (last, after live Ensembl).",
),
mint_vrs: bool = typer.Option(
True,
"--vrs/--no-vrs",
help="Mint GA4GH VRS allele ids onto resolved rows.",
),
verify_ref: bool = typer.Option(
True,
"--verify-ref/--no-verify-ref",
help="Check each authored ref against the reference sequence and report disagreements.",
),
verify_clinsig: bool = typer.Option(
True,
"--verify-clinsig/--no-verify-clinsig",
help="Check each authored clin_sig against the ClinVar snapshot's own (warns, never fails).",
),
verify_rsids: bool = typer.Option(
True,
"--verify-rsids/--no-verify-rsids",
help="Check each authored rsID against dbSNP for merges/withdrawals (online only).",
),
verify_datasets: bool = typer.Option(
True,
"--verify-datasets/--no-verify-datasets",
help="Check each release recorded in sources.csv against the one that source publishes now, and report the gap. One request per source, and the cheap question to put before --rederive: it tells you whether re-asking every subject is worth the run.",
),
keep_par_twin: bool = typer.Option(
False,
"--keep-par-twin",
help="Record both contigs of a pseudoautosomal locus. Default keeps only the X spelling, which is the one every annotation source uses and the only one a hard-masked GRCh38 analysis set can match.",
),
rederive: bool = typer.Option(
False,
"--rederive",
help="Re-ask every source about every subject, including the ones already recorded, and report which of them changed value. An ordinary run gap-fills and never re-asks, so a source that quietly revised an answer moves nothing you could notice.",
),
keep_staging: bool = typer.Option(
False,
"--keep-staging",
help="Leave the staged answers beside resolution.csv after a successful run. They are removed by default; a killed run leaves them either way, and the next run resumes from them.",
),
) -> None
Resolve a spec's variants into resolution.csv beside the spec. Exit 1 in strict mode if unresolved.
Source code in enricher/src/just_dna_enricher/cli.py
238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 | |
frequencies_ ¶
frequencies_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every resolved allele has a frequency.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: gnomAD frequency has no offline snapshot.",
),
populations: str | None = typer.Option(
None,
"--populations",
help="Comma-separated ancestry groups to keep (e.g. 'global' for one row per allele). Default: all.",
),
dataset: str | None = typer.Option(
None,
"--dataset",
help="Override the dataset label recorded on each row.",
),
) -> None
Fill frequencies.csv from the coordinates already in resolution.csv (pass 2, online only).
Source code in enricher/src/just_dna_enricher/cli.py
gene_metrics_ ¶
gene_metrics_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every gene has constraint metrics.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Snapshot only: never touch the network.",
),
constraint_cache: Path | None = typer.Option(
None,
"--constraint-cache",
help="Explicit gnomAD constraint snapshot dir.",
),
) -> None
Fill gene_metrics.csv for the genes variants.csv mentions (pass 3, snapshot then live API).
With no local snapshot the v4.1 one is downloaded from HuggingFace first, exactly as enrich
provisions the Ensembl and ClinVar snapshots — --offline is what turns that off, and then the pass
is snapshot-only. Reaching the live API instead means v2.1.1 numbers, which the row's dataset
records; provisioning is what keeps a plain install on v4.1.
Source code in enricher/src/just_dna_enricher/cli.py
dosage_ ¶
dosage_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every gene is ClinGen-curated.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: ClinGen's curation list is a live download with no snapshot.",
),
url: str = typer.Option(
DEFAULT_CLINGEN_URL,
"--url",
help="ClinGen gene-curation list URL.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial. ClinGen is CC0, so no declaration is refused here — it is recorded into sources.csv beside the rows it justifies.",
),
) -> None
Add ClinGen dosage-sensitivity rows to gene_metrics.csv (haploinsufficiency/triplosensitivity).
Source code in enricher/src/just_dna_enricher/cli.py
gene_validity_ ¶
gene_validity_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
source: str = typer.Option(
CLINGEN_VALIDITY_SOURCE,
"--source",
help="Which submitter to read: clingen (expert panels) or gencc (an aggregate of nineteen).",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every gene carries a curated assertion.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: neither ClinGen nor GenCC publishes an offline snapshot.",
),
url: str | None = typer.Option(
None,
"--url",
help="Override the submitter's export URL.",
),
) -> None
Fill gene_validity.csv with curated gene-disease assertions for the genes variants.csv names.
One row per (gene, disease, mode of inheritance, submitter) — the source's own grain. Mode of
inheritance is in the key because 59 ClinGen (gene, disease) pairs carry two curations that differ
only there, and submitter is in it because GenCC publishes the disagreement between submitters,
which is the thing it exists to publish.
Source code in enricher/src/just_dna_enricher/cli.py
gwas_ ¶
gwas_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Severity ladder for findings; see the pass docstring.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: this pass reads the REST API, not a snapshot.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use recorded on the licence row: unstated|non-commercial|commercial.",
),
study_facts: bool = typer.Option(
True,
"--study-facts/--no-study-facts",
help="Follow each association's study and trait links. Costs 2 requests per association; measured at 382 requests for one real module. Off keeps effects, drops pmid/trait/ancestry PERMANENTLY for the rows it writes: the merge is keyed on association_id, so a later run with study facts on skips those rows rather than back-filling. Delete gwas_effects.csv to re-derive them.",
),
) -> None
Fill gwas_effects.csv with the GWAS Catalog's published effect sizes for this module's rsIDs.
One row per published association, not per variant — a well-studied variant carries dozens across
different traits and papers. It does NOT fill weight: an authored weight is the author's model
of the finding, and no tool writes one. The two sit side by side and a consumer picks.
Reads effect_unit verbatim, including the Catalog's uninformative "unit", because a beta whose
scale is unknown must not look like one whose scale is shared. An association the Catalog
published without establishing which allele carries the effect keeps a null effect_allele and is
counted in the manifest, never dropped.
Source code in enricher/src/just_dna_enricher/cli.py
assertions_ ¶
assertions_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every resolved allele has a ClinVar record.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Snapshot only: never touch the network.",
),
clinvar_cache: Path | None = typer.Option(
None,
"--clinvar-cache",
help="Explicit ClinVar snapshot directory.",
),
) -> None
Fill clinical_assertions.csv from the coordinates already in resolution.csv.
Records what ClinVar says about each allele and how much review sits behind it — the star rating a compiled module previously discarded, so a one-star single submission and a practice guideline stopped being the same claim. Offline-capable: with a snapshot provisioned this pass never touches the network, and with none reachable it is a no-op rather than a failure.
It records; it does not adjudicate. Whether the module's own clin_sig agrees with ClinVar's is the
enrich cross-check's question, and that one warns in both modes on purpose.
Source code in enricher/src/just_dna_enricher/cli.py
literature_ ¶
literature_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail if a cited PMID does not resolve.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: there is no offline PubMed snapshot.",
),
check_fulltext: bool = typer.Option(
True,
"--fulltext/--no-fulltext",
help="Also match provenance quotes against fulltext, falling back to the abstract.",
),
check_doi: bool = typer.Option(
True,
"--doi/--no-doi",
help="Also confirm the authored DOI resolves in Crossref (covers preprints/books).",
),
) -> None
Fill literature.csv from a module's citations (pass 4, online only).
studies.csv is one citation site of several: a pmid on a binning row grounds the threshold it
sits on, and one on a pharm_variants.csv row grounds that row's own drug and genotype claim. The
pass reads every site, so a module citing only from those tables is enriched rather than refused.
Source code in enricher/src/just_dna_enricher/cli.py
734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 | |
pgx_ ¶
pgx_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail on an allele-function discrepancy.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Snapshots only: never reach PharmVar or CPIC live.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial. Sources that forbid sale are SKIPPED when unstated and REFUSED when commercial.",
),
use_pharmvar: bool = typer.Option(
True,
"--pharmvar/--no-pharmvar",
help="Consult PharmVar (needs PHARMVAR_API_KEY).",
),
use_cpic: bool = typer.Option(
True,
"--cpic/--no-cpic",
help="Consult CPIC (open, no key).",
),
cpic_cache: Path | None = typer.Option(
None,
"--cpic-cache",
help="Explicit CPIC snapshot dir.",
),
pharmvar_cache: Path | None = typer.Option(
None,
"--pharmvar-cache",
help="Explicit PharmVar snapshot dir.",
),
) -> None
Cross-check star-allele tables against PharmVar/CPIC and record terms into sources.csv.
Snapshot first, live second (RM38). A built snapshot serves the check without egress and without
spending a shared per-IP budget; --offline says snapshot-only, and a leg with neither is skipped
with a reason rather than silently passing.
Source code in enricher/src/just_dna_enricher/cli.py
clinpgx_build_ ¶
clinpgx_build_(
out_dir: Path = typer.Option(
repro_out("clinpgx"),
"--out",
file_okay=False,
help="Snapshot output directory.",
),
zip_path: Path | None = typer.Option(
None,
"--zip",
help=f"An existing {CURRENT_ARCHIVE.archive} (else downloaded).",
),
url: str = typer.Option(
DEFAULT_CLINPGX_URL,
"--url",
help="ClinPGx bulk download URL.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial.",
),
) -> None
Download + build the ClinPGx snapshot (dev surface; needs polars).
Source code in enricher/src/just_dna_enricher/cli.py
clinpgx_check_ ¶
clinpgx_check_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
snapshot: Path | None = typer.Option(
None,
"--snapshot",
help="Explicit ClinPGx snapshot dir. Omit it and the cache is used, or one is downloaded.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Use a local snapshot only: never download one.",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail on a stale evidence level.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial.",
),
) -> None
Cross-check pharm_variants.csv against the ClinPGx snapshot.
The snapshot no longer has to be handed over by hand (RM38): explicit path → $JUST_DNA_CLINPGX_CACHE
/ the default cache → downloaded from HuggingFace. --offline stops at the second step.
Source code in enricher/src/just_dna_enricher/cli.py
draft_ ¶
draft_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
gene: list[str] = typer.Option(
...,
"--gene",
help="Gene to draft from CPIC (repeatable).",
),
drug: list[str] = typer.Option(
[],
"--drug",
help="Also draft CPIC's prescribing recommendations for this drug (repeatable).",
),
allele: list[str] = typer.Option(
[],
"--allele",
help="Draft only these star alleles, in all three tables (repeatable; `*1` is always kept). A caller emits a bounded allele set, and n alleles is n(n+1)/2 pairs — CYP2D6 is 16,290 diplotypes unfiltered. Requires a single --gene, since a star name is gene-scoped.",
),
population: str | None = typer.Option(
None,
"--population",
help="Draft only this CPIC clinical context (e.g. 'NVI'). Default: every context, as rows.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial. CPIC forbids sale, so a draft is SKIPPED when unstated and REFUSED when commercial.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Draft from a built CPIC snapshot only; never reach CPIC live.",
),
cpic_cache: Path | None = typer.Option(
None,
"--cpic-cache",
help="Explicit CPIC snapshot dir.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be added; write nothing.",
),
) -> None
Draft PGx tables for one or more genes from CPIC — appends rows, never overwrites one.
Re-runnable and additive, so a multi-gene module is built up a gene at a time. A row whose key is
already in the file is reported, never replaced: what CPIC now says about a row you already wrote
is a finding for pgx, not an edit for this command to make.
Source code in enricher/src/just_dna_enricher/cli.py
1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 | |
template_ ¶
template_(
kind: str = typer.Argument(
...,
help="Authored CSV to emit a header for, e.g. repeat_alleles.csv",
),
) -> None
Print a header-only CSV for one authored table kind, generated from the live models.
Kept working here, but just-dna-compiler template is canonical: this needs no network, and an
author who installed only the tier that owns the CSV shape should not have to add the network
tier to get a header. See just-dna-compiler stub for a template with rows to replace.
Source code in enricher/src/just_dna_enricher/cli.py
check_identifiers_ ¶
check_identifiers_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Exit 1 if any identifier is stale.",
),
traits: bool = typer.Option(
True,
"--traits/--no-traits",
help="Check trait_efo_id against OLS4.",
),
genes: bool = typer.Option(
True,
"--genes/--no-genes",
help="Check gene symbols against HGNC.",
),
pgs: bool = typer.Option(
True,
"--pgs/--no-pgs",
help="Check pgs_id against the PGS Catalog, and the two authored cells beside it.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial. A PGS score licensed for academic research only bars sale, so a module citing one compiles ONLY with a declaration — and this flag is the one the compile's own refusal tells you to re-run with.",
),
) -> None
Report obsolete trait terms, retired gene symbols and unrecognised PGS accessions (online).
Writes no authored cell, and records that the question was put. Unlike the rsID check (whose
verdict lands on resolution.csv), these are module-level identifiers with no sidecar column to
record, and filling one from the registry being asked about it would make the comparison vacuous
— see hints.REDUNDANCY_BEARING. What this does write is verification.json: an attestation that
the five checks ran and over how many rows, never a value. A consumer holding the artifact has no
other way to tell "asked and clean" from "never asked" (RM45/RM72).
The PGS leg also writes sources.csv (RM163), and that is not an exception to the sentence above:
the Catalog's license is a field on each score record and it varies, so a module carrying an
academic-research-use-only score must not compile claiming the generic terms. The rows are the
terms, never a value in an authored cell.
Source code in enricher/src/just_dna_enricher/cli.py
1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 | |
check_acmg_ ¶
check_acmg_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Exit 1 if any acmg_sf disagrees.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No network. Needs --sf-list, else nothing is checked.",
),
url: str = typer.Option(
DEFAULT_ACMG_URL,
"--url",
help="ACMG secondary-findings page URL (fallback).",
),
sf_list: Path | None = typer.Option(
None,
"--sf-list",
exists=True,
file_okay=False,
help="Built ACMG SF snapshot (see `acmg build`). Preferred: NCBI's page still serves v3.2. Omit it and a snapshot in $JUST_DNA_ACMG_CACHE (or the shared cache base) is used; the page is scraped only when neither is there.",
),
) -> None
Check each row's acmg_sf against the ACMG secondary-findings list (reports only).
Writes no authored cell, and records that the question was put — the same two halves as
check-identifiers. acmg_sf is an authored cell this asks a registry about, not a fact this
pass contributes, and filling it here would break the check (see hints.REDUNDANCY_BEARING). The
verification.json record is an attestation, never a value: it says the list was consulted and
over how many rows, which is the one thing a downstream reader cannot reconstruct from the
artifact (RM45/RM72).
Source code in enricher/src/just_dna_enricher/cli.py
1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 | |
enrich_and_compile ¶
enrich_and_compile(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
output_dir: Path = typer.Argument(
...,
file_okay=False,
help="Output dir for parquet + manifest.json",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Fail unless every variant resolves.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Cache-only: never touch the network.",
),
ensembl_cache: Path | None = typer.Option(
None,
"--ensembl-cache",
help="Explicit Ensembl cache dir/.duckdb.",
),
clinvar_cache: Path | None = typer.Option(
None,
"--clinvar-cache",
help="Explicit ClinVar snapshot dir.",
),
use_clinvar: bool = typer.Option(
True,
"--clinvar/--no-clinvar",
help="Use the ClinVar link (after the Ensembl cache).",
),
use_gnomad: bool = typer.Option(
True,
"--gnomad/--no-gnomad",
help="Use the gnomAD link (last, after live Ensembl).",
),
frequencies: bool = typer.Option(
False,
"--frequencies",
help="Also run the frequency pass (writes frequencies.csv).",
),
gene_metrics: bool = typer.Option(
False,
"--gene-metrics",
help="Also run the gene-constraint pass (writes gene_metrics.csv).",
),
) -> None
Enrich, then compile from the produced resolution.csv (offline, deterministic). Exit 1 on failure.
Source code in enricher/src/just_dna_enricher/cli.py
upload_ ¶
upload_(
module_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Compiled module directory (at least one annotation parquet + manifest.json).",
),
repo_id: str | None = typer.Option(
None,
"--repo",
help="Target HF dataset (owner/name). Default: just-dna-seq/annotators.",
),
name: str | None = typer.Option(
None,
"--name",
help="Module name under data/<name>/ (and data/<name>/v<version>/) in the repo. Default: the directory basename.",
),
commit_message: str | None = typer.Option(
None,
"--message",
"-m",
help="Commit message. Default: 'Add <name> module'.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded without contacting HuggingFace.",
),
force: bool = typer.Option(
False,
"--force",
help="Overwrite data/<name>/v<version>/ even when it already holds a different artifact. Without this the publish refuses; the flat path is always overwritten.",
),
) -> None
Upload a compiled module to a HuggingFace dataset collection (publisher/dev surface).
Writes data/
Refuses when the versioned path already holds a different artifact (compare by artifact.digest), unless --force. The flat path means latest and is overwritten either way.
Source code in enricher/src/just_dna_enricher/cli.py
1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 | |
acmg_build_ ¶
acmg_build_(
workbook: Path = typer.Argument(
...,
exists=True,
dir_okay=False,
help="ACMG SF supplementary workbook (.xlsx), downloaded by you.",
),
out: Path = typer.Option(
repro_out("acmg_sf"),
"--out",
file_okay=False,
help="Output snapshot directory (writes acmg_sf.csv + release.json).",
),
source_url: str | None = typer.Option(
None,
"--source-url",
help="Where the workbook came from, recorded in release.json.",
),
doi: str | None = typer.Option(
None,
"--doi",
help="DOI of the statement the workbook accompanies, recorded in release.json.",
),
) -> None
Convert ACMG's SF workbook into the snapshot check-acmg --sf-list reads.
Why this exists: NCBI's page serves v3.2 and ACMG published v3.3 in June 2025, so the live scrape reports correctly authored rows as wrong. Nothing is downloaded here — the workbook is ACMG/Elsevier supplementary material and the author supplies their own copy, which is the same inject-only shape every other reference in this repo uses.
Source code in enricher/src/just_dna_enricher/cli.py
clinvar_build_ ¶
clinvar_build_(
vcf: Path | None = typer.Option(
None,
"--vcf",
exists=True,
dir_okay=False,
help="Local ClinVar VCF (.vcf.gz). Omit and pass --download to fetch from NCBI.",
),
download: bool = typer.Option(
False,
"--download",
help="Download the NCBI ClinVar GRCh38 VCF into --out first.",
),
out: Path = typer.Option(
repro_out("clinvar"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/*.parquet + release.json).",
),
) -> None
Convert a ClinVar VCF into the per-chromosome parquet snapshot the resolver reads.
Source code in enricher/src/just_dna_enricher/cli.py
clinvar_publish_ ¶
clinvar_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/*.parquet + release.json).",
),
repo_id: str | None = typer.Option(
None,
"--repo",
help="Target HF dataset (owner/name). Default: just-dna-seq/clinvar.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded. Reads the repo's file list; uploads nothing.",
),
) -> None
Create-or-update the dataset repo and upload the built ClinVar snapshot (publisher/dev).
Source code in enricher/src/just_dna_enricher/cli.py
cache_status_ ¶
Say which snapshots are present, where, and which release each holds.
Reads only: nothing is downloaded, so this is safe on a machine with no network and it is the first thing to run when a pass reports that a source was skipped.
Source code in enricher/src/just_dna_enricher/cli.py
cache_pull_ ¶
cache_pull_(
only: list[str] = typer.Option(
[],
"--only",
help="Pull just these caches (repeatable). Default: every publishable one.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use for the licence-gated snapshots. They forbid sale, so they are SKIPPED when unstated and REFUSED when commercial — downloading is taking the data.",
),
) -> None
Download the published parquet snapshots from HuggingFace into the local caches.
The provisioning step a hosted deployment runs once, so no pass ever reaches a source live per request. Already-complete caches are trusted without touching the network, so this is re-runnable and cheap; a truncated file is removed and refetched.
A lane with nothing published says so and names its reason, which is a field on the registry
rather than a comment: PharmVar's and PubMind's are refusals, ACMG's and MANE's are permissions
nobody has established. Build those with cache rebuild.
Source code in enricher/src/just_dna_enricher/cli.py
1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 | |
cache_prepare_ ¶
cache_prepare_(
only: list[str] = typer.Option(
[],
"--only",
help="Prepare just these caches (repeatable). Default: every one.",
),
use: str = typer.Option(
"unstated",
"--use",
help=f"Declared use: one of {sorted(VALID_DECLARED_USE)}.",
),
pin: list[str] = typer.Option(
[],
"--pin",
help="lane=release, repeatable, for the lanes that are built rather than pulled.",
),
source: list[str] = typer.Option(
[],
"--source",
help="lane=path, repeatable: build from a file you already hold.",
),
) -> None
Leave this machine with every cache it can have — pull what is published, build what is not.
The complement of cache pull, and the one command a deployment actually wants. pull
fetches the published snapshots and stops; four lanes are not published for recorded reasons —
PharmVar's personal key, PubMind's absent terms, NCBI's policy over MANE, ACMG's supplementary
material — so a machine that only pulled is missing four caches and the checks that read them
skip themselves. This runs each lane by the route it has.
The route is a property of the lane, never a flag. A published lane pulls, because building it would spend an operator's bandwidth re-deriving bytes somebody already made; an unpublished one builds, because that is the only route there will ever be. Asking for the choice would be asking an operator to restate the licensing story.
A cache that is already present is left alone, exactly as cache pull leaves one alone, so
this is idempotent and cheap to re-run. Re-cutting a snapshot that exists is cache rebuild,
which writes somewhere else on purpose — a build straight into a live cache is visible half-done
to anything reading it, and a short parquet still has a footer.
The Python counterpart is just_dna_enricher.caches.prepare_caches, which this calls.
Source code in enricher/src/just_dna_enricher/cli.py
1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 | |
cache_prune_ ¶
cache_prune_(
only: list[str] = typer.Option(
[],
"--only",
help="Prune just these caches (repeatable). Default: every published one.",
),
yes: bool = typer.Option(
False,
"--yes",
help="Delete without asking. Without it this prints the plan and stops.",
),
) -> None
Say what a published snapshot repo carries that its lane is not made of, and offer to delete it.
Deletion is never a side effect of publishing, and this is the command that makes that
affordable (RM186). A published repo accumulates: the publisher adds and does not remove, so a
layout change leaves the old spelling in place, and just-dna-seq/clinvar still carries the
159 MB single-file clinvar.parquet from before the per-chromosome split. Provisioning already
refuses to download it — the glob is what defends this tier — but any consumer globbing
data/*.parquet, the dataset viewer included, still gets two schemas under one relation.
Nothing here is a sweep. A file is a candidate only if the lane's own glob excludes it or a
LayoutShift declares it retired; README.md, .gitattributes, release.json, LICENSE.txt
and sidecar directories are never touched. Without --yes this reads and prints and does nothing
else, which is the mode to run first.
Source code in enricher/src/just_dna_enricher/cli.py
2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 | |
cache_rebuild_ ¶
cache_rebuild_(
out: Path = typer.Option(
Path(CACHES_DIRNAME),
"--out",
file_okay=False,
help="Base directory. Each lane is built into <base>/<lane>/, never in place. The default is under data/, which this workspace git-ignores wholesale.",
),
only: list[str] = typer.Option(
[],
"--only",
help="Rebuild just these caches (repeatable). Default: every one that can be.",
),
use: str = typer.Option(
"unstated",
"--use",
help=f"Declared use: one of {sorted(VALID_DECLARED_USE)}.",
),
pin: list[str] = typer.Option(
[],
"--pin",
help="lane=release, repeatable. e.g. --pin mane=1.5 --pin civic=2026-08-01.",
),
source: list[str] = typer.Option(
[],
"--source",
help="lane=path, repeatable: build from a file you already hold instead of downloading. Required for acmg; the offline off-switch for clinvar, constraint, clinpgx, drug_labels, pubmind and strchive. mane and civic take three files each and refuse it.",
),
publish: bool = typer.Option(
False,
"--publish",
help="Also upload each rebuilt snapshot to its HuggingFace repo.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="With --publish: show what would be uploaded, send nothing.",
),
) -> None
Rebuild every cache this tier builds — acquire, convert, and optionally publish (RM176).
The one endpoint over eleven builders. Each per-lane X build command stays, and this calls
the same download_*/build_* functions they do, so there is one conversion algorithm with two
callers rather than two that have to agree. What differs is only flag plumbing: a per-lane command
offers the local-file inputs an operator holds, and a rebuild pass by definition holds none.
Every lane is built into <base>/<lane>/, never in place over a resolved cache. A rebuild
takes minutes and an enrich reading a half-written snapshot mid-flight would see a real but
incomplete table — the failure a resolver cannot detect, because a short parquet is still a
parquet. Point the caches at the new base when the run is done, or copy each directory across.
An outcome is three-valued. ACMG needs a workbook that is Elsevier supplementary material, PharmVar a personal key, CIViC a release date to pin — none of those is a failure, and a nightly rebuild reporting errors for them would be reporting the licences working as designed. They are printed as not run, with the reason, and the exit code counts only real failures.
Source code in enricher/src/just_dna_enricher/cli.py
2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 | |
cpic_build_ ¶
cpic_build_(
out_dir: Path = typer.Option(
repro_out("cpic"),
"--out",
file_okay=False,
help="Snapshot output directory.",
),
endpoint: str = typer.Option(
DEFAULT_CPIC_ENDPOINT,
"--endpoint",
help="CPIC PostgREST base URL.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial.",
),
) -> None
Fetch CPIC whole into data/*.parquet + release.json (dev surface; needs polars).
No gene filter, deliberately: the whole database is ~120k narrow rows, and a snapshot covering only the genes the operator thought of answers "CPIC has nothing" for the next one.
Source code in enricher/src/just_dna_enricher/cli.py
cpic_publish_ ¶
cpic_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/*.parquet + release.json).",
),
repo: str = typer.Option(
DEFAULT_CPIC_REPO_ID,
"--repo",
help="Target HuggingFace dataset repo (owner/name).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded; send nothing.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Create-or-update the dataset repo and upload the built CPIC snapshot (publisher/dev).
Publishable because CPIC's recorded terms permit redistribution — CC BY-SA grants sharing under
share-alike plus attribution, which sources.csv carries. PharmVar has no equivalent command, and
that is the design rather than an omission.
Source code in enricher/src/just_dna_enricher/cli.py
clinpgx_publish_ ¶
clinpgx_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/*.parquet + release.json).",
),
repo: str = typer.Option(
DEFAULT_CLINPGX_REPO_ID,
"--repo",
help="Target HuggingFace dataset repo (owner/name).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded; send nothing.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Publish a built ClinPGx snapshot so clinpgx check can provision it (publisher/dev).
LICENSE.txt travels with the parquet — the terms ClinPGx ships inside its own archive are what
license_sha256 pins, and a published snapshot without them pins nothing for whoever downloads it.
Source code in enricher/src/just_dna_enricher/cli.py
pharmvar_build_ ¶
pharmvar_build_(
out_dir: Path = typer.Option(
repro_out("pharmvar"),
"--out",
file_okay=False,
help="Snapshot output directory.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial.",
),
) -> None
Fetch PharmVar whole into data/*.parquet + release.json (dev surface; needs polars + a key).
There is no pharmvar publish, and there will not be. The data is pulled under a key
PharmVar's terms §2 make personal and non-transferable, and no axis SourceTerms records covers
passing that on — an unestablished permission is not a permission. Build your own; point at it with
$JUST_DNA_PHARMVAR_CACHE or --pharmvar-cache.
Source code in enricher/src/just_dna_enricher/cli.py
civic_build_ ¶
civic_build_(
release: str | None = typer.Option(
None,
"--release",
help="Dated CIViC release to download, e.g. 01-Aug-2026. A DATED release, never the nightly: a snapshot that cannot name its input is one nothing can reproduce.",
),
evidence: Path | None = typer.Option(
None,
"--evidence",
exists=True,
dir_okay=False,
help="Local ClinicalEvidenceSummaries.tsv. Use instead of --release to build offline.",
),
variants: Path | None = typer.Option(
None,
"--variants",
exists=True,
dir_okay=False,
help="Local VariantSummaries.tsv.",
),
profiles: Path | None = typer.Option(
None,
"--profiles",
exists=True,
dir_okay=False,
help="Local MolecularProfileSummaries.tsv.",
),
out: Path = typer.Option(
repro_out("civic"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/civic.parquet + release.json).",
),
submitted: bool = typer.Option(
False,
"--submitted",
help="Also read the release's civic_accepted_and_submitted.vcf, so evidence a curator entered but no editor signed off joins the snapshot. Over 01-Aug-2026 that widens the direction corpus from 507 rows on 270 variants to 1149 on 397, adding 642 submitted rows and 127 variants, and every row carries the status CIViC gave it. Dated and pinnable like the TSVs, so the build stays reproducible.",
),
vcf: Path | None = typer.Option(
None,
"--vcf",
exists=True,
dir_okay=False,
help="Local civic_accepted_and_submitted.vcf. Use with the local TSV flags to build offline.",
),
) -> None
Reduce a dated CIViC release to the parquet snapshot the direction-axis drafter reads.
The bulk release, not the GraphQL API, and the two are not interchangeable. Every row of
ClinicalEvidenceSummaries.tsv is status accepted; the API defaults to NON_REJECTED and
serves roughly 2.35x as many evidence items. A snapshot has to be reproducible from a pinned
input, so this reads the dated files and records the basis in release.json.
--submitted widens that basis without leaving the dated release (RM169). CIViC publishes
<date>-civic_accepted_and_submitted.vcf beside the TSVs, so unreviewed evidence is pinnable too
and no API read is needed. The TSVs stay primary — the VCF cannot carry a variant with no GRCh37
coordinate, which is exactly the class whose identity had to be read out of its name — and the VCF
supplies the curation status, the submitted evidence, and the identity for the 112 variants
VariantSummaries.tsv (itself accepted-only) does not describe. Those rows are stamped
identity_derivation="vcf_csq", and nothing is placed from the VCF's own GRCh37 position.
There is no --use flag. CIViC is CC0 on every axis, so a declared-use gate would permit
every build unconditionally, and a flag feeding a gate that never gates is a flag that does
nothing (@acquisition-gate-is-not-a-read-gate).
Source code in enricher/src/just_dna_enricher/cli.py
2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 | |
civic_citations_ ¶
civic_citations_(
spec: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory.",
),
snapshot: Path | None = typer.Option(
None,
"--snapshot",
exists=True,
file_okay=False,
help="CIViC snapshot to map authored rows through. Default: the provisioned cache.",
),
variant_id: list[int] = typer.Option(
[],
"--variant-id",
help="Ask about a CIViC variant id directly, repeatable. Its citations ground the MODULE rather than a variant, which is the only route to a record CIViC publishes no identity for — variant 1955 is the case this exists for.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Do not fetch. Every subject is recorded as not-asked; no row is written.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be appended, write nothing.",
),
) -> None
Append the citations a CIViC variant carries that the dated bulk release cannot reach (RM160).
Why this is not part of civic build. The builder reads a dated release and is byte-
reproducible from it; civic reproduce proves it by building twice. The wider basis RM169 adopted
comes from a VCF, and a VCF record needs a POS — so submitted evidence attached to a variant with
no GRCh37 coordinate is published on exactly one surface, the GraphQL API, which has no release to
pin. The read is also one request per variant by construction, because evidenceItems takes a
single variantId. Batching it into the builder is the first repair anyone proposes and it is
exactly the reproducibility bargain this shape refused.
A recovered citation lands in studies.csv; literature.csv is derived from those PMIDs by the
literature command, and an article row nothing cites is dropped from the artifact. CIViC's own
curation status rides in confidence/confidence_unit, unconverted, so an accepted row and a
submitted row are not the same row. Appending only — a second run over an unchanged API adds
nothing — and enrich re-asks later and reports what has moved since.
CIViC is CC0, so there is no --use flag: a declared-use gate would permit every call
unconditionally, and a flag feeding a gate that never gates is a flag that does nothing.
Source code in enricher/src/just_dna_enricher/cli.py
2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 2655 2656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 | |
civic_publish_ ¶
civic_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/civic.parquet + release.json).",
),
repo_id: str | None = typer.Option(
None,
"--repo",
help="Target HF dataset (owner/name). Default: just-dna-seq/civic.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded without contacting HuggingFace.",
),
) -> None
Upload a built CIViC snapshot to a HuggingFace dataset repo (publisher/dev).
This one does not refuse, and the contrast with pubmind publish is the point. CIViC's content
is CC0 1.0 — a public-domain dedication with no share-alike, no bar on sale and attribution
requested rather than required — so there is no permission to establish before passing the bytes
on. PubMind's command exists in order to say no because its terms are unstated; PharmVar's cache is
unpublishable because its terms forbid it. Nothing here is in either position.
What the snapshot carries is a derivation of CIViC's release, not a copy of it: the germline
direction rows, placed on GRCh38 through identifiers CIViC itself publishes. release.json records
which dated release it came from and the accepted status basis, so a consumer can tell what they
are looking at without re-deriving it.
The ClinGen Allele Registry's answers are NOT in here, and that is deliberate rather than an oversight: the registry states no terms, and a lookup performed at draft time is a read, while baking its responses into a published file would be redistribution of bytes nobody has established we may pass on. The snapshot carries the CAID; resolving it stays the consumer's own fetch.
Source code in enricher/src/just_dna_enricher/cli.py
civic_reproduce_ ¶
civic_reproduce_(
release: str = typer.Option(
"01-Aug-2026",
"--release",
help="Dated CIViC release to reproduce, e.g. 01-Aug-2026.",
),
out: Path = typer.Option(
repro_out("civic_reproduce"),
"--out",
file_okay=False,
help="Working directory. The release files and two independent builds land here. The default is under data/, which this workspace git-ignores wholesale.",
),
keep: bool = typer.Option(
False,
"--keep",
help="Leave the downloaded release files in place for inspection.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Skip the reference cross-check. The build and determinism checks still run.",
),
submitted: bool = typer.Option(
False,
"--submitted",
help="Reproduce the wider basis: also download the release's civic_accepted_and_submitted.vcf and build with it, so the submitted rows and their coordinates go through every check below rather than only the accepted ones.",
),
) -> None
Build the CIViC snapshot from a dated release and check it, end to end.
Five checks, and the third is the one worth the network. The first two are about us; the third is about whether the coordinates we produced are real.
- The release downloads and its bytes are recorded — a sha256 per file, so a rerun that disagrees is a finding about the source rather than a mystery.
- Two independent builds are byte-identical (Principle 7). A parquet has no inherent row order, so this is the check that the sort is doing its job.
- Every placed coordinate is cross-checked against the GRCh38 reference sequence. This is the external validation: the snapshot's positions come from RefSeq accessions inside ClinVar HGVS, and this asks an unrelated service (refget/seqrepo) whether the reference base at each of those positions is what we wrote. A wrong-build or off-by-one placement fails here and nowhere else.
- The drop registry closes — every input row kept or counted, an equality over a walked set.
- The published file list is exactly what the publisher would upload.
Exits non-zero if any check fails, so it is usable in CI.
Source code in enricher/src/just_dna_enricher/cli.py
2752 2753 2754 2755 2756 2757 2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 2772 2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 2794 2795 2796 2797 2798 2799 2800 2801 2802 2803 2804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 2879 2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 2898 2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 2918 2919 2920 2921 2922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 2942 2943 2944 2945 | |
pubmind_build_ ¶
pubmind_build_(
table: Path | None = typer.Option(
None,
"--table",
exists=True,
dir_okay=False,
help="Local hg38_pubmind_db.txt.gz. Omit and pass --download to fetch it from ANNOVAR.",
),
download: bool = typer.Option(
False,
"--download",
help="Download the ANNOVAR-distributed PubMind table into --out first.",
),
out: Path = typer.Option(
repro_out("pubmind"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/pubmind.parquet + release.json).",
),
) -> None
Reduce the ANNOVAR-distributed PubMind table to the parquet snapshot the checks read.
There is no --use flag, and its absence is the design. PubMind's data terms could not be
established, and unknown terms warn rather than gate (commercial_use=None never taints a
module). A declared-use gate here would refuse every build unconditionally, and a flag feeding a
gate that never gates is a flag that does nothing. What the unknown terms do gate is publishing
a snapshot or a module carrying these bytes — see pubmind publish, which refuses.
Source code in enricher/src/just_dna_enricher/cli.py
2978 2979 2980 2981 2982 2983 2984 2985 2986 2987 2988 2989 2990 2991 2992 2993 2994 2995 2996 2997 2998 2999 3000 3001 3002 3003 3004 3005 3006 3007 3008 3009 3010 3011 3012 3013 3014 3015 3016 3017 3018 3019 3020 3021 3022 3023 3024 3025 3026 3027 3028 3029 3030 3031 3032 3033 3034 3035 3036 3037 3038 3039 3040 3041 3042 3043 3044 3045 3046 3047 3048 | |
pubmind_publish_ ¶
Refuse to publish the PubMind snapshot, and say why.
The command exists in order to refuse. A missing one reads as an oversight somebody will helpfully add, and the reason belongs where a reader looks for it rather than only in a design document.
Source code in enricher/src/just_dna_enricher/cli.py
constraint_build_ ¶
constraint_build_(
tsv: Path | None = typer.Option(
None,
"--tsv",
exists=True,
dir_okay=False,
help="Local gnomAD constraint metrics TSV. Omit and pass --download to fetch it.",
),
download: bool = typer.Option(
False,
"--download",
help="Download the gnomAD v4.1 constraint TSV (95.5 MB) into --out first.",
),
out: Path = typer.Option(
repro_out("gnomad_constraint"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/gnomad_constraint.parquet + release.json).",
),
) -> None
Reduce the per-transcript constraint TSV to the gene-level parquet the resolver reads.
Source code in enricher/src/just_dna_enricher/cli.py
constraint_publish_ ¶
constraint_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/*.parquet + release.json).",
),
repo_id: str | None = typer.Option(
None,
"--repo",
help="Target HF dataset (owner/name). Default: just-dna-seq/gnomad_constraint.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded without contacting HuggingFace.",
),
) -> None
Create-or-update the dataset repo and upload the built constraint snapshot (publisher/dev).
Source code in enricher/src/just_dna_enricher/cli.py
vrs_mint_ ¶
vrs_mint_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
offline: bool = typer.Option(
False,
"--offline",
help="Substitutions only: indels need the reference sequence, which means a network call.",
),
) -> None
Stamp ga4gh:VA.… allele ids onto resolution.csv (substitutions offline, indels online).
Source code in enricher/src/just_dna_enricher/cli.py
hint_variant_ ¶
hint_variant_(
rsid: str | None = typer.Option(
None, "--rsid", help="dbSNP id to look up."
),
chrom: str | None = typer.Option(
None, "--chrom", help="Chromosome (with --start)."
),
start: int | None = typer.Option(
None,
"--start",
help="1-based position (with --chrom).",
),
ref: str | None = typer.Option(
None,
"--ref",
help="Reference allele, for an allele-exact lookup.",
),
alts: str | None = typer.Option(
None,
"--alts",
help="Alt allele(s), comma-separated.",
),
ambiguity: bool = typer.Option(
False,
"--ambiguity",
help="Warn when the answer is not unique.",
),
frequencies: bool = typer.Option(
False,
"--frequencies",
help="Add gnomAD populations (paced: ~6s).",
),
offline: bool = typer.Option(
False,
"--offline",
help="Snapshots only; never touch the network.",
),
ensembl_cache: Path | None = typer.Option(
None,
"--ensembl-cache",
help="Explicit Ensembl cache.",
),
clinvar_cache: Path | None = typer.Option(
None,
"--clinvar-cache",
help="Explicit ClinVar snapshot.",
),
pubmind_cache: Path | None = typer.Option(
None,
"--pubmind-cache",
help="Explicit PubMind snapshot (see `pubmind build`); $JUST_DNA_PUBMIND_CACHE otherwise.",
),
as_json: bool = typer.Option(
False,
"--json",
help="Emit the full machine answer.",
),
) -> None
Validity, coordinates, alleles, populations and clinical calls for one variant.
Nothing is decided for you: a one-to-many rsID returns every locus and a position matching
several rsIDs returns every candidate. The coordinate is reported, never written into
variants.csv — resolution puts it in resolution.csv, which is where it belongs.
Source code in enricher/src/just_dna_enricher/cli.py
3264 3265 3266 3267 3268 3269 3270 3271 3272 3273 3274 3275 3276 3277 3278 3279 3280 3281 3282 3283 3284 3285 3286 3287 3288 3289 3290 3291 3292 3293 3294 3295 3296 3297 3298 3299 3300 3301 3302 3303 3304 3305 3306 3307 3308 3309 3310 3311 3312 3313 3314 3315 3316 3317 3318 3319 3320 3321 3322 3323 3324 3325 3326 3327 3328 3329 3330 3331 3332 3333 3334 3335 3336 3337 3338 3339 3340 3341 3342 3343 3344 3345 3346 | |
hint_recover_ ¶
hint_recover_(
chrom: str = typer.Option(
...,
"--chrom",
help="Chromosome of the old coordinate.",
),
start: int = typer.Option(
...,
"--start",
help=f"1-based {GRCH37_BUILD} position.",
),
ref: str | None = typer.Option(
None,
"--ref",
help="Reference allele, to narrow the answer.",
),
alts: str | None = typer.Option(
None,
"--alts",
help="Alt allele(s), comma-separated.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Skip the lookup and say so.",
),
as_json: bool = typer.Option(
False,
"--json",
help="Emit the full machine answer.",
),
) -> None
Which rs-number GRCh37 dbSNP records at an hg19/GRCh37 coordinate.
For a paper that predates GRCh38. Author the rs-number, not a converted position: an
rs-number resolves into a coordinate the compiler can cross-examine, where a lifted-over position
becomes the row's only witness to itself. Nothing is written — the rs-number is the row's
identity, and a machine filling one migrates variant_key with no authored edit anywhere.
Source code in enricher/src/just_dna_enricher/cli.py
hint_citation_ ¶
hint_citation_(
pmid: str | None = typer.Option(
None, "--pmid", help="PubMed id to check."
),
doi: str | None = typer.Option(
None,
"--doi",
help="DOI to check (the one you authored).",
),
pmcid: str | None = typer.Option(
None,
"--pmcid",
help="PubMed Central id (PMC…) to resolve to the PubMed id tables key on.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Skip the check and say so.",
),
as_json: bool = typer.Option(
False,
"--json",
help="Emit the full machine answer.",
),
) -> None
Does this citation exist, and is it the paper you meant?
A paywall hides the fulltext, never the PubMed record, so existence is answerable for paywalled
work; Crossref covers what PubMed does not index at all. Every answer is tri-state — unknown
means the registry could not be asked, which is not the same as "no such paper".
Existence is not identity. PMIDs are densely allocated, so a recalled or invented number is
very likely to be a real record for a different article, and pmid_exists alone cannot catch a
fabricated citation. The title, journal, year and first author come back in the same response and
are printed for exactly that comparison (S12).
--pmcid goes the other way. Every pmid in the schema keys on the PubMed id — studies.csv,
a binning row's and a pharm_variants.csv row's alike — and a curator holding only a PMC… id
had no route to it: the schema refused the cell and named no remedy. This resolves it and then asks PubMed which paper that is. The id is reported,
never written: filling pmid from NCBI would make the existence check compare NCBI with itself.
Source code in enricher/src/just_dna_enricher/cli.py
3392 3393 3394 3395 3396 3397 3398 3399 3400 3401 3402 3403 3404 3405 3406 3407 3408 3409 3410 3411 3412 3413 3414 3415 3416 3417 3418 3419 3420 3421 3422 3423 3424 3425 3426 3427 3428 3429 3430 3431 3432 3433 3434 3435 3436 3437 3438 3439 3440 3441 3442 3443 3444 3445 3446 3447 3448 3449 3450 3451 3452 3453 3454 3455 3456 3457 3458 3459 3460 | |
hint_trait_ ¶
Is this trait id current, obsolete, or unknown?
hint_gene_ ¶
Is this gene symbol approved or retired?
draft_clinpgx_ ¶
draft_clinpgx_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
snapshot: Path = typer.Option(
...,
"--snapshot",
exists=True,
file_okay=False,
help="Built ClinPGx snapshot (see `clinpgx build`). Inject-only; nothing is downloaded.",
),
drug: list[str] = typer.Option(
[],
"--drug",
help="Only annotations naming this drug (repeatable).",
),
gene: list[str] = typer.Option(
[],
"--gene",
help="Only annotations naming this gene (repeatable).",
),
min_evidence_level: str | None = typer.Option(
None,
"--min-evidence-level",
help="Keep annotations at least this strong: 1A|1B|2A|2B|3|4.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use: unstated | non-commercial | commercial. ClinPGx forbids sale.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be added; write nothing.",
),
) -> None
Draft pharm_variants.csv rows from the ClinPGx snapshot — appends, never overwrites a row.
Narrow with --drug and re-run as the module grows. A row already in the file is reported, never
replaced: drift against ClinPGx is clinpgx check's finding, not this command's edit to make.
Source code in enricher/src/just_dna_enricher/cli.py
draft_panel_ ¶
draft_panel_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
gene: list[str] = typer.Option(
[],
"--gene",
help="Gene to draft rows for (repeatable). Required for every source but mitomap-miss, whose increment is asked for as a whole and where --gene only filters.",
),
source: str = typer.Option(
"clinvar",
"--source",
help="Which authority to draft the calls from: clinvar (the default); pubmind — an LLM's reading of the literature, which needs an operator-built snapshot and still reads the ClinVar one for its gene attribution; civic — curated cancer interpretations, which writes the DIRECTION axis rather than clin_sig and needs a `civic build` snapshot; or mitomap-miss — the curated mtDNA calls MITOMAP publishes and the ClinVar cache does not, which needs a `mitomap miss` snapshot.",
),
mitomap_miss_cache: Path | None = typer.Option(
None,
"--mitomap-miss-cache",
exists=True,
file_okay=False,
help="Built MITOMAP-miss snapshot (see `mitomap miss`). Only read under --source mitomap-miss; omit it and $JUST_DNA_MITOMAP_MISS_CACHE is used.",
),
civic_cache: Path | None = typer.Option(
None,
"--civic-cache",
exists=True,
file_okay=False,
help="Built CIViC snapshot (see `civic build`). Only read under --source civic; omit it and $JUST_DNA_CIVIC_CACHE is used.",
),
snapshot: Path | None = typer.Option(
None,
"--snapshot",
exists=True,
file_okay=False,
help="Built ClinVar snapshot (see `clinvar build`). Omit it and the cache is used, or the published snapshot downloaded — the citations table comes with it, which is what a panel needs to compile. Read for its gene attribution under --source pubmind, which publishes no gene column of its own.",
),
pubmind_cache: Path | None = typer.Option(
None,
"--pubmind-cache",
exists=True,
file_okay=False,
help="Built PubMind snapshot (see `pubmind build`), for --source pubmind. Omit it and $JUST_DNA_PUBMIND_CACHE is read; there is no published one to download.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Use a local snapshot only: never download one.",
),
download: bool = typer.Option(
True,
"--download/--no-download",
help="Provision the published snapshot when no local one is found. Fetching it is this command's only network use, so --no-download coincides with --offline today; it is a separate switch because it says 'do not go and get one', not 'make no request'.",
),
clin_sig: str | None = typer.Option(
None,
"--clin-sig",
help="Comma-separated calls to include. Default: pathogenic,likely_pathogenic.",
),
min_review_stars: int = typer.Option(
2,
"--min-review-stars",
min=0,
max=4,
help="Review-status floor, --source clinvar only. 2 = multiple submitters, no conflicts.",
),
max_citations: int = typer.Option(
3,
"--max-citations",
min=0,
help="Study rows to draft per variant from ClinVar's literature links. 0 disables. --source clinvar only: PubMind's channel carries no PMID.",
),
min_confidence: int = typer.Option(
DEFAULT_MIN_CONFIDENCE,
"--min-confidence",
min=0,
max=3,
help="Evidence-depth floor, --source pubmind only. PubMind's confidence counts how much of the literature spoke, 0-3; 1 means more than a single mention.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use (ClinVar is public domain).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be added; write nothing.",
),
) -> None
Draft a gene panel's variants.csv rows from an authority — appends, never overwrites a row.
The drafted rows carry a genotype placeholder, so the module will not compile until you decide what each finding is about. That is deliberate: ClinVar publishes alleles, and whether carrying one is a carrier state or an affected one follows from the condition's inheritance mode, which the source does not say. Rows land in their gene's block, and a re-run leaves anything already there — stub or filled — exactly as it is.
Source code in enricher/src/just_dna_enricher/cli.py
3552 3553 3554 3555 3556 3557 3558 3559 3560 3561 3562 3563 3564 3565 3566 3567 3568 3569 3570 3571 3572 3573 3574 3575 3576 3577 3578 3579 3580 3581 3582 3583 3584 3585 3586 3587 3588 3589 3590 3591 3592 3593 3594 3595 3596 3597 3598 3599 3600 3601 3602 3603 3604 3605 3606 3607 3608 3609 3610 3611 3612 3613 3614 3615 3616 3617 3618 3619 3620 3621 3622 3623 3624 3625 3626 3627 3628 3629 3630 3631 3632 3633 3634 3635 3636 3637 3638 3639 3640 3641 3642 3643 3644 3645 3646 3647 3648 3649 3650 3651 3652 3653 3654 3655 3656 3657 3658 3659 3660 3661 3662 3663 3664 3665 3666 3667 3668 3669 3670 3671 3672 3673 3674 3675 3676 3677 3678 3679 3680 3681 3682 3683 3684 3685 3686 3687 3688 3689 3690 3691 3692 3693 3694 3695 3696 3697 3698 3699 3700 3701 3702 3703 3704 3705 3706 3707 3708 3709 3710 3711 3712 3713 3714 3715 3716 3717 3718 3719 3720 3721 3722 3723 3724 3725 3726 3727 3728 3729 3730 3731 3732 3733 3734 3735 3736 3737 3738 3739 3740 3741 3742 3743 3744 3745 3746 3747 3748 3749 3750 3751 3752 3753 3754 3755 3756 3757 3758 3759 3760 3761 3762 3763 3764 3765 3766 3767 3768 3769 3770 3771 3772 3773 3774 3775 3776 3777 3778 3779 3780 3781 3782 3783 3784 3785 3786 3787 3788 3789 3790 3791 3792 3793 | |
clinvar_citations_ ¶
clinvar_citations_(
out: Path = typer.Option(
...,
"--out",
file_okay=False,
help="Existing ClinVar snapshot dir.",
),
citations_txt: Path | None = typer.Option(
None,
"--citations",
exists=True,
dir_okay=False,
help="Local var_citations.txt.",
),
download: bool = typer.Option(
False,
"--download",
help="Fetch var_citations.txt first.",
),
url: str = typer.Option(
DEFAULT_CITATIONS_URL,
"--url",
help="Source for --download.",
),
) -> None
Add ClinVar's literature links to a snapshot: data/citations.parquet ([dev], needs polars).
Separate from clinvar build because ClinVar publishes citations separately from the VCF — which
is precisely why a drafted gene panel could not compile without this: studies.csv is mandatory
and the VCF carries no PMIDs. Written beside the snapshot, so an existing cache keeps its bytes.
Source code in enricher/src/just_dna_enricher/cli.py
litvar_coverage_ ¶
litvar_coverage_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
offline: bool = typer.Option(
False,
"--offline",
help="No network; every locus is recorded as unchecked.",
),
quiet: bool = typer.Option(
False,
"--quiet",
help="Only the tier summary, not a line per locus.",
),
) -> None
Report LitVar's literature coverage per locus, naming the tier that answered.
It answers which papers discuss an allele that is already identified. It does not answer which allele a name meant — those read as the same question and are not. Measured against the two hardest records in this repository (CIViC 1955 and 2131, four candidate alleles with registered CAIDs), the index returns no node for any of them, because PubTator3 mines titles and abstracts and those alleles live in a table inside a paywalled paper. Do not reach for this to recover an identity.
Writes no row and no sources.csv entry: nothing here reaches a module's tables, so the module
does not use this source. What it does write is verification.json — an attestation that the
question was put, over how many loci, and at which tier each was answered.
Source code in enricher/src/just_dna_enricher/cli.py
3852 3853 3854 3855 3856 3857 3858 3859 3860 3861 3862 3863 3864 3865 3866 3867 3868 3869 3870 3871 3872 3873 3874 3875 3876 3877 3878 3879 3880 3881 3882 3883 3884 3885 3886 3887 3888 3889 3890 3891 3892 3893 3894 3895 3896 3897 3898 3899 3900 3901 3902 3903 3904 3905 3906 3907 3908 3909 3910 3911 3912 3913 3914 3915 3916 3917 3918 3919 3920 3921 3922 3923 3924 | |
litvar_gene_ ¶
List every node LitVar holds under a gene symbol, grouped by tier. Writes nothing.
This is the endpoint that serves line-delimited Python repr() rather than JSON, and the tier
split is the reason to look: of 588 HFE nodes on 2026-09-01, 220 are rsID-only, 69 carry a
ClinGen allele id, exactly one is the gene node, and the remaining 298 are unnormalized protein
strings — text a miner saw, not an identity anything should join on.
Source code in enricher/src/just_dna_enricher/cli.py
mane_build_ ¶
mane_build_(
download: bool = typer.Option(
False,
"--download",
help="Fetch the release from NCBI. Without --release the newest version is discovered from current/README_versions.txt and then pinned to its versioned directory.",
),
release: str | None = typer.Option(
None,
"--release",
help="MANE version to pin, e.g. 1.5. Resolves to release_<version>/, never current/.",
),
summary: Path | None = typer.Option(
None,
"--summary",
exists=True,
dir_okay=False,
help="Local MANE.GRCh38.v<ver>.summary.txt.gz. Use instead of --download to build offline.",
),
changed: Path | None = typer.Option(
None,
"--changed",
exists=True,
dir_okay=False,
help="Local MANE.GRCh38.v<ver>.changed_select_accessions.txt.gz.",
),
not_in_mane: Path | None = typer.Option(
None,
"--not-in-mane",
exists=True,
dir_okay=False,
help="Local MANE.GRCh38.v<ver>.protein_coding_genes_not_in_mane.txt.gz.",
),
versions: Path | None = typer.Option(
None,
"--versions",
exists=True,
dir_okay=False,
help="Local README_versions.txt. Optional, and the only way an offline build can name its release: a filename is never parsed for one.",
),
out: Path = typer.Option(
repro_out("mane"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/*.parquet + release.json).",
),
) -> None
Reduce one pinned MANE release to the three parquet tables the numbering frame needs.
All three files, in one pass. The summary is the frame; changed_select_accessions is the
currency check and Update_Affects_CDS is the numbering-frame axis stated by the source; the
negative roster is what makes "MANE has no answer for this gene" distinguishable from "nobody
asked", with the reason attached. Shipping the cache without the thing that notices it going
stale is the defect this command exists to close.
There is no --offline flag: the off-switch is passing the local files instead of
--download. And there is no --use flag, because NCBI states a policy rather than a licence —
every gating axis is unknown, and check_declared_use returns a skip for an unknown whatever
the declaration says, so the gate would silently skip every build. A flag feeding a gate that
never gates is a flag that does nothing (@acquisition-gate-is-not-a-read-gate).
MANE is the default, not the answer. A gene with two rows carries two CDS numbering frames
and MANE_status says which; a gene with one row says nothing about the isoforms MANE does not
carry, so a pass treating this table as an oracle would be wrong in a way the table cannot report.
Source code in enricher/src/just_dna_enricher/cli.py
3970 3971 3972 3973 3974 3975 3976 3977 3978 3979 3980 3981 3982 3983 3984 3985 3986 3987 3988 3989 3990 3991 3992 3993 3994 3995 3996 3997 3998 3999 4000 4001 4002 4003 4004 4005 4006 4007 4008 4009 4010 4011 4012 4013 4014 4015 4016 4017 4018 4019 4020 4021 4022 4023 4024 4025 4026 4027 4028 4029 4030 4031 4032 4033 4034 4035 4036 4037 4038 4039 4040 4041 4042 4043 4044 4045 4046 4047 4048 4049 4050 4051 4052 4053 4054 4055 4056 4057 4058 4059 4060 4061 4062 4063 4064 4065 4066 4067 4068 4069 4070 4071 4072 4073 4074 4075 4076 4077 4078 4079 4080 4081 4082 4083 4084 4085 4086 4087 4088 4089 4090 4091 4092 4093 4094 4095 4096 4097 4098 4099 4100 4101 4102 4103 4104 4105 4106 4107 4108 4109 4110 4111 4112 4113 4114 4115 4116 4117 4118 4119 4120 4121 4122 4123 4124 4125 4126 4127 4128 4129 4130 4131 4132 4133 4134 4135 4136 4137 4138 4139 4140 4141 4142 4143 4144 4145 4146 4147 4148 4149 4150 4151 4152 4153 4154 | |
strchive_build_ ¶
strchive_build_(
out: Path = typer.Option(
repro_out("strchive"),
"--out",
file_okay=False,
help="Output snapshot directory (writes STRchive-loci.json + release.json).",
),
catalogue: Path | None = typer.Option(
None,
"--catalogue",
exists=True,
dir_okay=False,
help="A STRchive-loci.json you already have. Without it the file is downloaded.",
),
release: str | None = typer.Option(
None,
"--release",
help="Upstream release tag to pin, e.g. v2.26.0. Without it, the default branch, unlabelled.",
),
) -> None
Fetch (or copy in) the STRchive catalogue and record its provenance beside it.
Pin a release: the default branch moves, so a comparison whose reference is "whatever was there
that afternoon" cannot be re-run, and only a pinned build gets a dataset label the verification
record can name.
Source code in enricher/src/just_dna_enricher/cli.py
strchive_publish_ ¶
strchive_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (STRchive-loci.json + release.json).",
),
repo: str = typer.Option(
DEFAULT_STRCHIVE_REPO_ID,
"--repo",
help="Target HuggingFace dataset repo (owner/name).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded; send nothing.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Create-or-update the dataset repo and upload the built STRchive catalogue (publisher/dev).
Publishable on the source's own terms: STRchive is MIT, which grants redistribution outright.
This lane had no publish command because it grew from a check rather than from a cache, not
because anything withheld the permission — the same distinction the roster draws between CIViC's
absent ensure_* (a gap) and PharmVar's (a refusal).
Publish a pinned build. An unlabelled snapshot carries no dataset, so whoever pulls it can
run the comparison and cannot say which release they compared against — build with --release
first, and this refuses nothing but says so.
Source code in enricher/src/just_dna_enricher/cli.py
mitomap_build_ ¶
mitomap_build_(
out: Path = typer.Option(
repro_out("mitomap"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/mitomap-*.parquet + release.json).",
),
dump: Path | None = typer.Option(
None,
"--dump",
exists=True,
dir_okay=False,
help="A mitomap.dump.sql.gz you already have. Without it the dump is downloaded — the data surface answers plain curl, unlike the web surface. A local dump carries no Last-Modified, so that snapshot is honestly unlabelled.",
),
url: str = typer.Option(
DEFAULT_MITOMAP_URL,
"--url",
help="Source URL for the dump (used only when --dump is absent).",
),
) -> None
Cut the two curated mtDNA variant tables, their citations and the references out of the dump.
602 mmutation rows and 494 rtmutation rows out of 6.76 million lines. The snapshot records
every count this build computes — rows per table, the dump's own per-table edit dates, how much of
reference.nlmid is a PMID, the alleles that cannot be spelled as VCF and the brackets that are
not a documented VCEP class — because a number computed and dropped is one every reader has to
recompute.
Source code in enricher/src/just_dna_enricher/cli.py
4279 4280 4281 4282 4283 4284 4285 4286 4287 4288 4289 4290 4291 4292 4293 4294 4295 4296 4297 4298 4299 4300 4301 4302 4303 4304 4305 4306 4307 4308 4309 4310 4311 4312 4313 4314 4315 4316 4317 4318 4319 4320 4321 4322 4323 4324 4325 4326 4327 4328 4329 4330 4331 4332 4333 4334 4335 4336 4337 4338 4339 4340 4341 4342 4343 4344 4345 4346 4347 4348 4349 4350 4351 4352 4353 4354 4355 4356 4357 4358 4359 4360 4361 4362 | |
mitomap_miss_ ¶
mitomap_miss_(
out: Path = typer.Option(
repro_out("mitomap_miss"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/mitomap_miss.parquet + release.json).",
),
mitomap_cache: Path | None = typer.Option(
None,
"--mitomap-cache",
exists=True,
file_okay=False,
help="Built MITOMAP snapshot (see `mitomap build`). Omit it and $JUST_DNA_MITOMAP_CACHE is used.",
),
clinvar_cache: Path | None = typer.Option(
None,
"--clinvar-cache",
exists=True,
file_okay=False,
help="Built ClinVar snapshot (see `clinvar build`). Omit it and $JUST_DNA_CLINVAR_CACHE is used.",
),
) -> None
Join MITOMAP against the ClinVar chrMT parquet and write the increment (RM171).
A derived lane, not a download. Its acquire stage is both parents being on disk, and a parent that is absent is reported as could-not-run rather than as an empty increment — a miss set computed without ClinVar would say MITOMAP publishes a thousand alleles nobody else has, from a comparison that never ran.
Exact (start, ref, alt) on chrMT, upper-cased both sides, no position-level fallback. Four
buckets, and draft-panel --source mitomap-miss writes only one of them.
Source code in enricher/src/just_dna_enricher/cli.py
4365 4366 4367 4368 4369 4370 4371 4372 4373 4374 4375 4376 4377 4378 4379 4380 4381 4382 4383 4384 4385 4386 4387 4388 4389 4390 4391 4392 4393 4394 4395 4396 4397 4398 4399 4400 4401 4402 4403 4404 4405 4406 4407 4408 4409 4410 4411 4412 4413 4414 4415 4416 4417 4418 4419 4420 4421 4422 4423 4424 4425 4426 4427 4428 4429 4430 4431 4432 4433 4434 4435 4436 4437 4438 4439 4440 4441 4442 4443 4444 4445 4446 4447 4448 4449 4450 4451 4452 4453 | |
mitomap_publish_ ¶
mitomap_publish_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/ + release.json).",
),
repo: str = typer.Option(
DEFAULT_MITOMAP_REPO_ID,
"--repo",
help="Target HuggingFace dataset repo (owner/name).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded; send nothing.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Create-or-update the dataset repo and upload the built MITOMAP snapshot (publisher/dev).
Publishable on the source's own terms: CC BY 3.0, with commercial and clinical use stated free and
attribution the one condition — which the snapshot's SourceRow carries. Only the parent lane
publishes. The derived miss snapshot pins two parent digests, so a pulled copy would be an
increment whose own currency check cannot be run by whoever pulled it; it is rebuilt locally from
the parents instead, which is cheaper than the download and cannot be stale.
Source code in enricher/src/just_dna_enricher/cli.py
draft_repeats_ ¶
draft_repeats_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
genes: list[str] = typer.Option(
[],
"--gene",
"-g",
help="Restrict to these genes. Repeatable; omit for every catalogue locus.",
),
catalogue: Path | None = typer.Option(
None,
"--catalogue",
exists=True,
help="Built STRchive snapshot directory (see `strchive build`), or a STRchive-loci.json. Omit it and $JUST_DNA_STRCHIVE_CACHE (or the shared cache base) is used.",
),
use: str = typer.Option(
"unstated",
"--use",
help=f"Declared use: one of {sorted(VALID_DECLARED_USE)}.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be added; write nothing.",
),
) -> None
Draft repeat_alleles.csv identity rows from STRchive — appends, never overwrites a row.
The bands are not drafted, and that is the design rather than a limitation. A drafted row
carries the gene, the motif as the catalogue spells it, the trait CURIE where the locus names
exactly one disease, and a conclusion placeholder — so the table cannot compile until a human
has filled in what each band means. measure_min/measure_max stay empty: run
check-repeat-bands once you have written them and it will report where the catalogue disagrees.
The catalogue's coordinates, ref_copies and locus_structure have no authored column to land
in; the run counts them and says so rather than dropping them silently.
Source code in enricher/src/just_dna_enricher/cli.py
4504 4505 4506 4507 4508 4509 4510 4511 4512 4513 4514 4515 4516 4517 4518 4519 4520 4521 4522 4523 4524 4525 4526 4527 4528 4529 4530 4531 4532 4533 4534 4535 4536 4537 4538 4539 4540 4541 4542 4543 4544 4545 4546 4547 4548 4549 4550 4551 4552 4553 4554 4555 4556 4557 4558 4559 4560 4561 4562 4563 4564 4565 4566 | |
check_repeat_bands_ ¶
check_repeat_bands_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
catalogue: Path | None = typer.Option(
None,
"--catalogue",
exists=True,
help="Built STRchive snapshot directory (see `strchive build`), or a STRchive-loci.json.",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Carried into the report. A band difference NEVER fails, in either mode.",
),
) -> None
Compare a module's repeat_alleles.csv bands against STRchive's, and report the differences.
Writes no authored cell and never fails on a difference. Where a catalogue and an expert
author draw a repeat threshold in different places, both are claims by an authority, and a compile
that refused would make this format pick the winner — the rule the ClinVar clin_sig and PGx
allele-function checks already follow. --strict is accepted so the flag means one thing across
the tier, and it changes nothing here but the mode recorded in the report.
The catalogue's pathogenic_max is reported as its own finding and is never written: it is the
longest allele the literature records, not a clinical ceiling, and a module that imported it would
silently answer nothing at all for a longer one.
Source code in enricher/src/just_dna_enricher/cli.py
clinpgx_build_labels_ ¶
clinpgx_build_labels_(
out_dir: Path = typer.Option(
repro_out("drug_labels"),
"--out",
file_okay=False,
help="Snapshot output directory.",
),
zip_path: Path | None = typer.Option(
None,
"--zip",
exists=True,
dir_okay=False,
help="A drugLabels.zip you already have. Without it the archive is downloaded.",
),
url: str = typer.Option(
DEFAULT_DRUG_LABELS_URL,
"--url",
help="ClinPGx bulk download URL.",
),
use: str = typer.Option(
"unstated",
"--use",
help=f"Declared use: one of {sorted(VALID_DECLARED_USE)}.",
),
) -> None
Download + build the regulator drug-label snapshot (dev surface; needs polars).
A second archive from a source this tier already adopted, with its own release.json: ClinPGx
publishes at least twelve downloads on this endpoint and they do not refresh in lockstep, so the
label snapshot is dated from its own CREATED_*.txt rather than from the annotation lane's.
There is no --offline: a builder's off-switch is passing --zip instead of downloading.
Source code in enricher/src/just_dna_enricher/cli.py
4622 4623 4624 4625 4626 4627 4628 4629 4630 4631 4632 4633 4634 4635 4636 4637 4638 4639 4640 4641 4642 4643 4644 4645 4646 4647 4648 4649 4650 4651 4652 4653 4654 4655 4656 4657 4658 4659 4660 4661 4662 4663 4664 4665 4666 4667 4668 4669 4670 4671 4672 4673 4674 4675 4676 4677 4678 4679 4680 4681 4682 4683 4684 4685 4686 4687 | |
clinpgx_check_labels_ ¶
clinpgx_check_labels_(
spec_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
snapshot: Path | None = typer.Option(
None,
"--snapshot",
exists=True,
file_okay=False,
help="Built drug-label snapshot directory (see `clinpgx build-labels`). Omit it and $JUST_DNA_DRUG_LABELS_CACHE (or the shared cache base) is used.",
),
strict: bool = typer.Option(
False,
"--strict/--best-effort",
help="Carried into the report. A label difference NEVER fails, in either mode.",
),
use: str = typer.Option(
"unstated",
"--use",
help=f"Declared use: one of {sorted(VALID_DECLARED_USE)}.",
),
) -> None
Compare a module's gene/allele/drug claims against the drug labels five regulators publish.
Two join tiers, reported apart, and the tier belongs to the question. What do the agencies say about this gene and this medicine is the gene-tier subject; …and this star allele or rsID is the allele-tier one. A label naming both answers both, because they are two questions rather than one asked twice, and a gene-level agreement is not an allele-level agreement.
Writes no authored cell and never fails on a difference. Five agencies genuinely disagree with
each other — clopidogrel and CYP2C19 is Actionable PGx at four of them and Informative PGx at
the EMA — and a compile that refused would make this format pick the winner. A blank Testing
Level is a third of the file and is reported as unknown, never as No Clinical PGx.
Source code in enricher/src/just_dna_enricher/cli.py
4690 4691 4692 4693 4694 4695 4696 4697 4698 4699 4700 4701 4702 4703 4704 4705 4706 4707 4708 4709 4710 4711 4712 4713 4714 4715 4716 4717 4718 4719 4720 4721 4722 4723 4724 4725 4726 4727 4728 4729 4730 4731 4732 4733 4734 4735 4736 4737 4738 4739 4740 4741 4742 4743 4744 4745 4746 4747 4748 4749 4750 4751 4752 4753 4754 4755 4756 4757 4758 4759 4760 4761 4762 4763 | |
clinpgx_publish_labels_ ¶
clinpgx_publish_labels_(
snapshot_dir: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Built snapshot directory (data/drug_labels.parquet + LICENSE.txt + release.json).",
),
repo: str = typer.Option(
DEFAULT_DRUG_LABELS_REPO_ID,
"--repo",
help="Target HuggingFace dataset repo (owner/name).",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded; send nothing.",
),
commit_message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Publish a built drug-label snapshot so clinpgx check-labels can provision it (publisher/dev).
A second repo rather than a second table in just-dna-seq/clinpgx, for the reason the builder
already gives its own release.json: the two ClinPGx archives do not refresh in lockstep, and one
repo holding both would date the pair from whichever was published last.
Same grounds as clinpgx publish — CC BY-SA permits redistribution, forbids sale, and requires
attribution, which sources.csv carries. LICENSE.txt travels with the parquet: a share-alike
snapshot whose terms did not travel pins nothing for whoever downloads it.
Source code in enricher/src/just_dna_enricher/cli.py
atlas_generate_ ¶
atlas_generate_(
refetch: bool = typer.Option(
False,
"--refetch",
help="Re-download the pinned sources even if they are already on disk and match.",
),
) -> None
Fetch the pinned Atlas .proto sources and generate the gRPC bindings from them.
The sources are not vendored: the repository carries a commit id and a sha256 per file, and a
file that does not match its pin is refused rather than used (RM196). Needs grpcio-tools,
which is in \[dev] and deliberately not in \[atlas] — the runtime imports the bindings without
it. A released wheel carries both the sources and the bindings already, so this is a checkout
command.
Source code in enricher/src/just_dna_enricher/cli.py
alphagenome_avi_build_ ¶
alphagenome_avi_build_(
input_: Path = typer.Option(
...,
"--input",
exists=True,
dir_okay=False,
help="The extracted alphagenome_variant_impact_score_snvs.tsv.gz (its .tbi must be beside it). Required, and there is no default URL: acquisition is yours, under your own acceptance of the AlphaGenome Services Additional Terms.",
),
out: Path = typer.Option(
repro_out("alphagenome_avi"),
"--out",
file_okay=False,
help="Output snapshot directory (writes data/alphagenome_avi-*.parquet, avi_knots.parquet, release.json, LICENSE.txt).",
),
contig: list[str] = typer.Option(
None,
"--contig",
help="Build only these contigs, repeatable. Omit for every contig the .tbi index knows.",
),
workers: int = typer.Option(
12,
"--workers",
min=1,
help="How many contigs to read at once. Twelve ran 24 contigs in 41-46 minutes; one takes about four times as long.",
),
no_hash: bool = typer.Option(
False,
"--no-hash",
help="Skip the source sha256. It is a few minutes over 88.5 GB; release.json then records null, which is unknown rather than unpinned.",
),
) -> None
Build the AVI snapshot from a local copy of the artifact.
raw_score is stored as Int32 at a scale of 105 — exactly lossless, since the artifact
prints at most five decimals — and PHRED is not** stored: it is a rank, a function of
raw_score, and the 466 KB knot table beside the data reconstructs it while also carrying the
per-value ambiguity interval a threshold has to be checked against.
Source code in enricher/src/just_dna_enricher/cli.py
4889 4890 4891 4892 4893 4894 4895 4896 4897 4898 4899 4900 4901 4902 4903 4904 4905 4906 4907 4908 4909 4910 4911 4912 4913 4914 4915 4916 4917 4918 4919 4920 4921 4922 4923 4924 4925 4926 4927 4928 4929 4930 4931 4932 4933 4934 4935 4936 4937 4938 4939 4940 4941 4942 4943 4944 4945 4946 4947 4948 4949 4950 4951 4952 4953 4954 4955 4956 4957 4958 4959 4960 4961 4962 4963 4964 | |
alphagenome_check_ ¶
alphagenome_check_(
spec: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
reference: Path | None = typer.Option(
None,
"--reference",
exists=True,
file_okay=False,
help="An AVI snapshot directory. Omit to use $JUST_DNA_ALPHAGENOME_AVI_CACHE.",
),
threshold: float | None = typer.Option(
None,
"--threshold",
help="A PHRED cut to check the module's variants against. Without one the pass is entirely offline: there is no question the local artifact cannot answer.",
),
offline: bool = typer.Option(
False,
"--offline",
help="Never reach the Atlas. Straddling variants are recorded as nobody-asked, not as decided.",
),
refinement_cap: int = typer.Option(
DEFAULT_REFINEMENT_CAP,
"--refinement-cap",
min=1,
help="Refuse rather than refine more than this many variants over the network in one run.",
),
strict: bool = typer.Option(
False,
"--strict",
help="Carried for the report; see the docstring.",
),
) -> None
Cross-check a module's variants against AlphaGenome's AVI scores. Reports, never repairs.
The local snapshot answers most of it. The Atlas is asked only where the knot table says the local data genuinely cannot decide — a threshold falling inside a printed score's PHRED interval — and that set is computed offline, before any request is spent.
Source code in enricher/src/just_dna_enricher/cli.py
4967 4968 4969 4970 4971 4972 4973 4974 4975 4976 4977 4978 4979 4980 4981 4982 4983 4984 4985 4986 4987 4988 4989 4990 4991 4992 4993 4994 4995 4996 4997 4998 4999 5000 5001 5002 5003 5004 5005 5006 5007 5008 5009 5010 5011 5012 5013 5014 5015 5016 5017 5018 5019 5020 5021 5022 5023 5024 5025 5026 5027 5028 5029 5030 5031 5032 5033 5034 5035 5036 5037 5038 5039 5040 5041 | |
alphagenome_expression_ ¶
alphagenome_expression_(
spec: Path = typer.Argument(
...,
exists=True,
file_okay=False,
help="Module spec directory",
),
gene: str = typer.Option(
...,
"--gene",
help="HGNC symbol. REQUIRED even with an explicit interval: the server-side gene filter is not an optimisation, and an unfiltered interval query is refused before it is sent.",
),
chrom: str | None = typer.Option(
None,
"--chrom",
help="Contig of an explicit interval. Wins over the gene's MANE span.",
),
start: int | None = typer.Option(
None,
"--start",
min=0,
help="1-based start of that interval.",
),
end: int | None = typer.Option(
None,
"--end",
min=0,
help="1-based end of that interval.",
),
min_score: float | None = typer.Option(
None,
"--min-score",
help="Keep only pairs whose magnitude reaches this. Distal scores run ~10x lower than scores at the gene, so a flat bar keeps the proximal rows and looks like it filtered on effect.",
),
max_rows: int = typer.Option(
DEFAULT_MAX_ROWS,
"--max-rows",
min=1,
help="Refuse rather than write more rows than this. Raising it is a deliberate act.",
),
offline: bool = typer.Option(
False,
"--offline",
help="No-op with a warning: this pass reads the Atlas, not a snapshot.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Report what would be written without writing it.",
),
use: str = typer.Option(
"unstated",
"--use",
help="Declared use recorded on the licence row: unstated|non-commercial|commercial. AlphaGenome Output is NON-COMMERCIAL ONLY, so an undeclared run writes nothing and says so — pass --use non-commercial.",
),
) -> None
Fill expression_effects.csv with AlphaGenome's per-gene expression effects for one gene.
Example — and the --use is not decoration, an undeclared run is a no-op:
just-dna-enricher alphagenome expression ./my_module --gene TBX1 --use non-commercial
One row per (variant, gene): which way the variant moves that gene's predicted expression, how many of the 371 tissue tracks agree, and how far it sits from the gene. The interval is the gene's MANE span widened by the model's measured +/-512 kb attribution horizon, unless --chrom/--start/--end supply one; either way the gene names the server-side filter, and the MANE lane is still consulted for the distance, which an explicit interval cannot supply.
A whole gene is ~3.3 M SNVs at the measured 1,091 SNVs/s — about 50 minutes — and the cost is printed before the query runs rather than discovered during it.
Source code in enricher/src/just_dna_enricher/cli.py
5134 5135 5136 5137 5138 5139 5140 5141 5142 5143 5144 5145 5146 5147 5148 5149 5150 5151 5152 5153 5154 5155 5156 5157 5158 5159 5160 5161 5162 5163 5164 5165 5166 5167 5168 5169 5170 5171 5172 5173 5174 5175 5176 5177 5178 5179 5180 5181 5182 5183 5184 5185 5186 5187 5188 5189 5190 5191 5192 5193 5194 5195 5196 5197 5198 5199 5200 5201 5202 5203 5204 5205 5206 5207 5208 5209 5210 5211 5212 5213 5214 5215 5216 5217 5218 5219 5220 5221 5222 5223 5224 5225 5226 5227 5228 5229 5230 5231 5232 5233 5234 5235 5236 5237 5238 | |
alphagenome_publish_ ¶
alphagenome_publish_(
snapshot: Path | None = typer.Argument(
None,
exists=True,
file_okay=False,
help="The built snapshot directory. Omit to use the resolved cache ($JUST_DNA_ALPHAGENOME_AVI_CACHE, then the cache base), falling back to where `alphagenome build` writes.",
),
repo: str | None = typer.Option(
None,
"--repo",
help=f"Target HF dataset. Default: {DEFAULT_ALPHAGENOME_AVI_REPO_ID}.",
),
dry_run: bool = typer.Option(
False,
"--dry-run",
help="Show what would be uploaded. Reads the repo's file list; sends nothing.",
),
message: str | None = typer.Option(
None, "--message", "-m", help="Commit message."
),
) -> None
Publish the AVI snapshot to HuggingFace.
Its own command because no other one can reach this lane (RM202). cache rebuild --publish
walks lanes that have a rebuild adapter, and this lane cannot have one — its source is behind an
eligibility gate, so there is nothing for an unattended rebuild to fetch. RM198 gave the lane a
publish_repo and left it unreachable.
The upload is two commits and, above 5 GB, goes through the resumable uploader (RM199): the
payload first, then release.json — the description must never arrive before the bytes it
describes.
Source code in enricher/src/just_dna_enricher/cli.py
5241 5242 5243 5244 5245 5246 5247 5248 5249 5250 5251 5252 5253 5254 5255 5256 5257 5258 5259 5260 5261 5262 5263 5264 5265 5266 5267 5268 5269 5270 5271 5272 5273 5274 5275 5276 5277 5278 5279 5280 5281 5282 5283 5284 5285 5286 5287 5288 5289 5290 5291 5292 5293 5294 5295 5296 5297 5298 5299 5300 5301 5302 5303 5304 5305 5306 5307 5308 5309 5310 5311 5312 5313 5314 5315 5316 5317 5318 | |