Skip to content

just_dna_enricher.upload

just_dna_enricher.upload

HuggingFace upload — publisher surface of the network tier.

Two publish shapes share one create-or-update pathway (ensure_repo):

  • module — a compiled module's artifacts (every parquet the artifact carries + manifest.json, and a logo and readme if present) to datasets/<repo>/data/<name>/, matching the layout just-dna-lite's discovery scans, and to data/<name>/v<version>/ when the manifest states a version, so the discovery path can address a particular release rather than only "whatever is there now" (upload_module; RM84).
  • reference snapshot — a built ClinVar (or Ensembl) parquet snapshot + release.json to the root of a dataset repo, matching the download.ensure_*_snapshot layout (publish_reference_snapshot).

Both require a HuggingFace token with write access (hf auth login or HF_TOKEN). This is the dev/publisher half of the enricher's HuggingFace use: snapshot download is a runtime enrich path (download.ensure_snapshot); upload is for authors republishing modules (e.g. the Gen-I v1-port recreation) or publishing a rebuilt reference. Install as just-dna-enricher[dev].

Extracted from just_dna_pipelines.v1_port.publish (just-dna-lite); create_repo was added here (the origin assumed the repo pre-existed) so a brand-new dataset repo can be created on first push.

LayoutShift dataclass

LayoutShift(
    repo_id: str, retires: str, introduces: str, reason: str
)

A retirement declared by the commit that causes it, so the publish carries out the migration.

The rule (maintainer, 2026-09-03). A change that retires one published file and introduces another must also carry a conditioned update to the publication procedure: if the new file is absent from the repo and the old one is present, upload the new and delete the old. The condition is a predicate over the remote, so it fires exactly once per repo and is a no-op forever after — nothing is deleted on a repo that has already moved, and re-running a publish cannot re-delete.

This is the only deletion a publish performs. General cleanup is cache prune, which prints what it would remove and asks; a deletion nobody declared is not a side effect a publish may have. What makes this one safe is not that it is small but that it is named: the pair is in the commit that changed the layout, where a reviewer sees both halves at once.

Deleting is recoverable in a way this repository's @snapshot-accumulates note did not credit — a HuggingFace dataset repo is git-backed, and a superseded revision still resolves (three of them were read off the hub while auditing this on 2026-09-03). The reason to declare rather than sweep is not that bytes are lost; it is that a retired file goes on answering 200 to whoever still asks for it, which is how a lane's default archive stayed frozen for a year (CLINPGX_ARCHIVES).

OrphanedSidecarError

Bases: RuntimeError

The publish would overwrite a release.json describing sidecars it is not carrying (RM185).

Its own type for the reason PublishCollisionError is one: the CLI has to tell it from the refusals plan_* raises. Those say the local snapshot is not publishable; this one says the local snapshot is fine and the remote holds bytes this publish would silently stop describing.

The incident. A ClinVar snapshot is two halves on two cadences — the per-chromosome parquet and citations/citations.parquet — and release.json describes both because build_citations merges a block into it. A rebuild that built only the VCF half published a citations-free release.json over a repo whose sidecar it had not replaced, and the publisher adds without deleting: the sidecar outlived its own description, and the published artifact carried records from one ClinVar release beside citations from another while saying nothing. RM179 fixed the lane; this is the general shape, refused at the boundary where any lane could repeat it.

PublishCollisionError

Bases: RuntimeError

The versioned path already holds a different release, and no --force was given (RM88).

Its own type rather than a FileNotFoundError, because the CLI has to tell it apart from the three refusals plan_upload raises: those say the local module is not publishable, this one says the local module is fine and the remote already has that version.

UploadPlan

Bases: BaseModel

What an upload would send (also the dry-run result).

SnapshotPlan

Bases: BaseModel

What a reference-snapshot publish would send (also the dry-run result).

PruneCandidate

Bases: BaseModel

One remote file cache prune would remove, and the reason it is nameable as removable.

PrunePlan

Bases: BaseModel

What a prune would delete from one repo. The dry run and the deletion read the same object.

ensure_repo

ensure_repo(repo_id: str, token: str | None = None)

Create-or-update: ensure the dataset repo exists, returning the authenticated HfApi.

create_repo(..., exist_ok=True) is a no-op when the repo already exists, so create and update are one pathway. The returned api is reused by the caller's upload_folder so only one HfApi is constructed per publish.

Source code in enricher/src/just_dna_enricher/upload.py
def ensure_repo(repo_id: str, token: str | None = None):
    """Create-or-update: ensure the dataset repo exists, returning the authenticated ``HfApi``.

    ``create_repo(..., exist_ok=True)`` is a no-op when the repo already exists, so create and update
    are one pathway. The returned api is reused by the caller's ``upload_folder`` so only one
    ``HfApi`` is constructed per publish.
    """
    api = _hf_api(repo_id, token)
    api.create_repo(repo_id=repo_id, repo_type="dataset", exist_ok=True)
    return api

layout_shifts_to_apply

layout_shifts_to_apply(
    repo_files: Iterable[str], repo_id: str
) -> list[LayoutShift]

The declared shifts whose condition the remote currently satisfies. Reads; decides nothing else.

Separate from the publish so it can be asserted directly: the whole safety of this mechanism is that the predicate is false the moment it has run once, and that is a property of this function rather than of the caller that acts on it.

Source code in enricher/src/just_dna_enricher/upload.py
def layout_shifts_to_apply(repo_files: Iterable[str], repo_id: str) -> list[LayoutShift]:
    """The declared shifts whose condition the remote currently satisfies. Reads; decides nothing else.

    Separate from the publish so it can be asserted directly: the whole safety of this mechanism is
    that the predicate is false the moment it has run once, and that is a property of this function
    rather than of the caller that acts on it.
    """
    remote = list(repo_files)
    due: list[LayoutShift] = []
    for shift in LAYOUT_SHIFTS:
        if shift.repo_id != repo_id:
            continue
        if shift.retires not in remote:
            continue  # already retired, or never there
        if any(fnmatch(name, shift.introduces) for name in remote):
            continue  # the repo has already moved: not this one's job
        due.append(shift)
    return due

plan_upload

plan_upload(
    module_dir: Path, name: str, repo_id: str | None = None
) -> UploadPlan

Resolve the upload plan and validate the compiled artifacts are present.

Two destinations, both planned here (RM84): the flat data/<name>/, which is what the reference consumer's discovery scans and therefore has to keep working and keep meaning latest, and data/<name>/v<version>/ inside it — nested rather than a sibling, because the flat path is the one deployed readers are pointed at — which is the only thing on that path that can name a particular release. A module with no readable version gets the flat path alone, never a vNone directory, and version_unknown_reason says which of the four reasons applies. The reason is a field rather than a log line for the reason every other withheld answer here is (clin_sig_not_checked, skipped_offline): a caller that has to parse prose to learn what happened has not been told.

What it refuses is three positive rules, not one required list (RM89). The old rule demanded the three SNP-core parquets, which RM2 made optional four releases ago, so seven of the sixteen reference examples could not be published at all. In their place, ordered most specific first so a refusal names the actual fault:

  1. The plan carries everything the artifact attests. manifest.artifact.files states a name, a sha256 and a size per parquet, and artifact.digest is a Merkle root over exactly those — so a file attested and not sent makes the published manifest a false claim about bytes that are not there. This is the general guard, and it is a self-check as much as a module check: it fires if this publisher's allowlist ever falls behind the compiler's output list again.
  2. weights.parquet never travels alone. A SNP core compiles to all three, so a missing annotations/studies beside it is an interrupted compile.
  3. At least one lead table. A directory with a manifest, a readme and no annotation rows is not a module; publishing one is the silent failure a bare widening of the old rule would have produced, and it is what discovery would then be unable to see.
Source code in enricher/src/just_dna_enricher/upload.py
def plan_upload(
    module_dir: Path,
    name: str,
    repo_id: str | None = None,
) -> UploadPlan:
    """Resolve the upload plan and validate the compiled artifacts are present.

    Two destinations, both planned here (RM84): the flat `data/<name>/`, which is what the reference
    consumer's discovery scans and therefore has to keep working and keep meaning *latest*, and
    `data/<name>/v<version>/` **inside it** — nested rather than a sibling, because the flat path is
    the one deployed readers are pointed at — which is the only thing on that path that can name a
    particular release. A module with no readable version gets the flat path alone, never a `vNone`
    directory, and `version_unknown_reason` says which of the four reasons applies. The reason is a
    field rather than a log line for the reason every other withheld answer here is
    (`clin_sig_not_checked`, `skipped_offline`): a caller that has to parse prose to learn what
    happened has not been told.

    **What it refuses is three positive rules, not one required list (RM89).** The old rule demanded
    the three SNP-core parquets, which RM2 made optional four releases ago, so seven of the sixteen
    reference examples could not be published at all. In their place, ordered most specific first so a
    refusal names the actual fault:

    1. **The plan carries everything the artifact attests.** `manifest.artifact.files` states a name,
       a sha256 and a size per parquet, and `artifact.digest` is a Merkle root over exactly those — so
       a file attested and not sent makes the published manifest a false claim about bytes that are not
       there. This is the general guard, and it is a self-check as much as a module check: it fires if
       this publisher's allowlist ever falls behind the compiler's output list again.
    2. **`weights.parquet` never travels alone.** A SNP core compiles to all three, so a missing
       `annotations`/`studies` beside it is an interrupted compile.
    3. **At least one lead table.** A directory with a manifest, a readme and no annotation rows is not
       a module; publishing one is the silent failure a bare widening of the old rule would have
       produced, and it is what discovery would then be unable to see.
    """
    resolved_repo = repo_id or DEFAULT_REPO_ID
    present = [f for f in _ALLOW_PATTERNS if (module_dir / f).exists()]
    attested = _attested_parquets(module_dir)
    unsent = [] if attested is None else [f for f in attested if f not in present]
    if unsent:
        raise FileNotFoundError(
            f"{name}: {_MANIFEST_FILENAME} attests {unsent} but the upload would not carry "
            f"{'it' if len(unsent) == 1 else 'them'} — `artifact.digest` is computed over the "
            f"attested files, so what arrives could not be verified against the manifest that "
            f"describes it. Re-compile {module_dir} if the file is missing."
        )
    if "weights.parquet" in present:
        half = [f for f in _EXPECTED_WITH_WEIGHTS if f not in present]
        if half:
            raise FileNotFoundError(
                f"{name}: missing compiled artifact(s) {half} in {module_dir} beside "
                f"weights.parquet — a SNP core compiles to all three, so this directory is a "
                f"half-finished compile. Re-run `just-dna-enricher enrich-and-compile` "
                f"or `just-dna-compiler compile`."
            )
    if not any(f in LEAD_PARQUETS for f in present):
        raise FileNotFoundError(
            f"{name}: no annotation table in {module_dir} — a module carries at least one of "
            f"{list(LEAD_PARQUETS)}. Compile the module first (e.g. "
            f"`just-dna-enricher enrich-and-compile` or `just-dna-compiler compile`)."
        )
    flat = f"data/{name}"
    version, reason = _module_version(module_dir)
    return UploadPlan(
        module=name,
        repo_id=resolved_repo,
        path_in_repo=flat,
        versioned_path_in_repo=None if version is None else f"{flat}/v{version}",
        version_unknown_reason=reason,
        files=present,
    )

upload_module

upload_module(
    module_dir: Path,
    name: str,
    repo_id: str | None = None,
    token: str | None = None,
    commit_message: str | None = None,
    force: bool = False,
) -> UploadPlan

Upload the compiled module to a HuggingFace dataset collection.

Ensures the dataset repo exists (create-or-update), then writes the same files to the flat path and — when the manifest states a version — to the versioned subdirectory under it.

The two writes are two commits, not one. upload_folder commits per call, so a reader can briefly see the flat path refreshed while the versioned copy is not there yet; the flat path goes first because it is the one anything reads today. Making it atomic means create_commit over an explicit operation list, which is a different shape from the allow-pattern plumbing every publish here uses, and it was not worth holding the fix for. If the second call fails the first has already landed — data/<name>/ is the new release and the versioned copy is absent until a re-run, which is idempotent.

The versioned path refuses to be overwritten with different bytes (RM88). Before either write, the published manifest.json at data/<name>/v<version>/ is read and its artifact.digest compared with this module's. A different digest is a PublishCollisionError unless force=True. The flat path is deliberately not guarded: it means latest, so overwriting it is what it is for, and the whole point of the versioned copy is that it does not.

The policy is refuse-unless---force rather than warn-and-proceed or refuse-outright, decided 2026-08-18. The flag's existence is itself the claim that overwriting is sometimes right — a curator re-cutting a draft release is a real workflow, and a gate with no override becomes a gate people route around.

Recompiling under a newer compiler also moves the digest, and will trip this. That is correct rather than a false positive: P4 scopes byte-reproducibility to a fixed compiler_version, so the versioned path really would come to hold different bytes than the ones it was published with. The refusal says so, because "I changed nothing" is the first thing its first user will think.

Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.

Source code in enricher/src/just_dna_enricher/upload.py
def upload_module(
    module_dir: Path,
    name: str,
    repo_id: str | None = None,
    token: str | None = None,
    commit_message: str | None = None,
    force: bool = False,
) -> UploadPlan:
    """Upload the compiled module to a HuggingFace dataset collection.

    Ensures the dataset repo exists (create-or-update), then writes the same files to the flat path
    and — when the manifest states a version — to the versioned subdirectory under it.

    **The two writes are two commits, not one.** `upload_folder` commits per call, so a reader can
    briefly see the flat path refreshed while the versioned copy is not there yet; the flat path goes
    first because it is the one anything reads today. Making it atomic means `create_commit` over an
    explicit operation list, which is a different shape from the allow-pattern plumbing every publish
    here uses, and it was not worth holding the fix for. If the second call fails the first has already
    landed — `data/<name>/` is the new release and the versioned copy is absent until a re-run, which
    is idempotent.

    **The versioned path refuses to be overwritten with different bytes (RM88).** Before either
    write, the published `manifest.json` at `data/<name>/v<version>/` is read and its
    `artifact.digest` compared with this module's. A different digest is a `PublishCollisionError`
    unless `force=True`. The flat path is deliberately **not** guarded: it means *latest*, so
    overwriting it is what it is for, and the whole point of the versioned copy is that it does not.

    The policy is refuse-unless-`--force` rather than warn-and-proceed or refuse-outright, decided
    2026-08-18. The flag's existence is itself the claim that overwriting is sometimes right — a
    curator re-cutting a draft release is a real workflow, and a gate with no override becomes a gate
    people route around.

    **Recompiling under a newer compiler also moves the digest**, and will trip this. That is correct
    rather than a false positive: P4 scopes byte-reproducibility to a fixed `compiler_version`, so the
    versioned path really would come to hold different bytes than the ones it was published with. The
    refusal says so, because "I changed nothing" is the first thing its first user will think.

    Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.
    """
    plan = plan_upload(module_dir, name, repo_id)
    api = ensure_repo(plan.repo_id, token)
    published_digest = None if force else _versioned_digest_conflict(api, plan, _artifact_digest(module_dir))
    if published_digest is not None:
        raise PublishCollisionError(
            f"{name}: {plan.versioned_path_in_repo}/ is already published with a different artifact "
            f"({published_digest[:19]}… there, {_artifact_digest(module_dir)[:19]}… here), so this "
            f"upload would replace the bytes that version names. Bump `version:` in "
            f"module_spec.yaml and re-compile to publish this as a new release, or pass --force to "
            f"overwrite it deliberately. Note that recompiling an unchanged spec under a newer "
            f"compiler moves the digest too — if that is what happened, this refusal is still "
            f"protecting the bytes the published version was cut from."
        )
    api.upload_folder(
        folder_path=str(module_dir),
        path_in_repo=plan.path_in_repo,
        repo_id=plan.repo_id,
        repo_type="dataset",
        allow_patterns=_ALLOW_PATTERNS,
        commit_message=commit_message or f"Add {name} module",
    )
    if plan.versioned_path_in_repo is not None:
        segment = plan.versioned_path_in_repo.rsplit("/", 1)[-1]
        api.upload_folder(
            folder_path=str(module_dir),
            path_in_repo=plan.versioned_path_in_repo,
            repo_id=plan.repo_id,
            repo_type="dataset",
            allow_patterns=_ALLOW_PATTERNS,
            commit_message=commit_message or f"Add {name} module {segment}",
        )
    return plan

plan_reference_snapshot

plan_reference_snapshot(
    snapshot_dir: Path,
    repo_id: str | None = None,
    *,
    payload: str | None = None,
) -> SnapshotPlan

Resolve a snapshot publish plan and validate the built artifacts are present.

payload names a snapshot whose content is one file at the root rather than data/*.parquet — STRchive's STRchive-loci.json is the case, and ACMG's acmg_sf.csv is the shape's second member. The parameter is the caller's, deliberately: a lane knows the name of the file it builds, and putting a roster of lane filenames in the publisher would make this function the fourth place that has to learn about a new snapshot kind.

A payload snapshot is refused when the file is missing, exactly as a parquet one is when data/ is empty — the refusal is the same claim either way, that there is nothing built here to publish.

Source code in enricher/src/just_dna_enricher/upload.py
def plan_reference_snapshot(
    snapshot_dir: Path,
    repo_id: str | None = None,
    *,
    payload: str | None = None,
) -> SnapshotPlan:
    """Resolve a snapshot publish plan and validate the built artifacts are present.

    `payload` names a snapshot whose content is **one file at the root** rather than
    `data/*.parquet` — STRchive's `STRchive-loci.json` is the case, and ACMG's `acmg_sf.csv` is the
    shape's second member. The parameter is the caller's, deliberately: a lane knows the name of the
    file it builds, and putting a roster of lane filenames in the publisher would make this function
    the fourth place that has to learn about a new snapshot kind.

    A `payload` snapshot is refused when the file is missing, exactly as a parquet one is when
    `data/` is empty — the refusal is the same claim either way, that there is nothing built here to
    publish.
    """
    resolved_repo = repo_id or DEFAULT_CLINVAR_REPO_ID
    if payload is not None:
        if not (snapshot_dir / payload).is_file():
            raise FileNotFoundError(
                f"no {payload} in {snapshot_dir} — build the snapshot first "
                f"(e.g. `just-dna-enricher strchive build`)"
            )
        # Same registry as the parquet branch below, and for the same reason: this branch had its own
        # copy of the pair, so a payload-only lane that gained a root file would have dropped it
        # exactly the way the parquet branch dropped `LICENSE.txt`.
        files = [payload]
        files.extend(name for name in SNAPSHOT_ROOT_FILENAMES if (snapshot_dir / name).is_file())
        return SnapshotPlan(repo_id=resolved_repo, files=files)
    data_dir = snapshot_dir / SNAPSHOT_DATA_DIRNAME
    parquet = sorted(p.name for p in data_dir.glob("*.parquet")) if data_dir.is_dir() else []
    if not parquet:
        raise FileNotFoundError(
            f"no {SNAPSHOT_DATA_DIRNAME}/*.parquet in {snapshot_dir} — build the snapshot first "
            f"(e.g. `just-dna-enricher clinvar build`)"
        )
    files = [f"{SNAPSHOT_DATA_DIRNAME}/{name}" for name in parquet]
    # Sidecars, when the snapshot has them. Only ClinVar does, and only when `clinvar citations` was
    # run, so absence is normal rather than a reason to refuse.
    for sidecar in SNAPSHOT_SIDECAR_DIRNAMES:
        directory = snapshot_dir / sidecar
        if directory.is_dir():
            files.extend(f"{sidecar}/{p.name}" for p in sorted(directory.glob("*.parquet")))
    # Walked from `locations.SNAPSHOT_ROOT_FILENAMES`, never listed here. A file this function does
    # not name is a file the publish silently drops, and the snapshot still looks complete on the
    # other side — which is how a share-alike snapshot went out without its terms
    # (`@publisher-allowlist-derived`). Absence stays normal: only ClinPGx carries a licence and only
    # the AVI lane carries a knot table.
    for name in SNAPSHOT_ROOT_FILENAMES:
        if (snapshot_dir / name).is_file():
            files.append(name)
    return SnapshotPlan(repo_id=resolved_repo, files=files)

check_publish_orphans_no_sidecar

check_publish_orphans_no_sidecar(
    plan: SnapshotPlan, api=None, token: str | None = None
) -> None

Refuse a publish that would leave a remote sidecar undescribed. Reads the repo; writes nothing.

The remote tree, not the remote release.json. The block is a description of the bytes and the bytes are what a puller gets, so the description is the half that can already be missing — which is exactly the state just-dna-seq/clinvar was found in on 2026-09-03: the citations parquet present, the block gone. A guard reading the block would have passed the second bad publish as happily as the first.

Scoped to publishes that carry release.json, because a publish that carries none overwrites no provenance. A repo that does not exist yet lists nothing and passes — that is a first publish, not an orphan.

api is the caller's when it has one (a publish has just built an authenticated client and there is no reason to build a second); a dry run passes none and gets an anonymous reader, since listing a public repo needs no token and a dry run must not require write access to say what a publish would do.

Source code in enricher/src/just_dna_enricher/upload.py
def check_publish_orphans_no_sidecar(plan: SnapshotPlan, api=None, token: str | None = None) -> None:
    """Refuse a publish that would leave a remote sidecar undescribed. Reads the repo; writes nothing.

    **The remote tree, not the remote `release.json`.** The block is a *description* of the bytes and
    the bytes are what a puller gets, so the description is the half that can already be missing —
    which is exactly the state `just-dna-seq/clinvar` was found in on 2026-09-03: the citations parquet
    present, the block gone. A guard reading the block would have passed the second bad publish as
    happily as the first.

    Scoped to publishes that carry `release.json`, because a publish that carries none overwrites no
    provenance. A repo that does not exist yet lists nothing and passes — that is a first publish, not
    an orphan.

    `api` is the caller's when it has one (a publish has just built an authenticated client and there
    is no reason to build a second); a dry run passes none and gets an anonymous reader, since listing
    a public repo needs no token and a dry run must not require write access to say what a publish
    would do.
    """
    if RELEASE_FILENAME not in plan.files:
        return
    if api is None:
        try:
            from huggingface_hub import HfApi, get_token
        except ImportError as exc:
            raise ImportError(
                "huggingface_hub is required to check a publish against the published repo"
            ) from exc
        api = HfApi(token=env_value("HF_TOKEN") or get_token())
    try:
        remote = list(api.list_repo_files(repo_id=plan.repo_id, repo_type="dataset"))
    except Exception as exc:
        # Nobody has published here yet, or the listing failed. Neither is an orphan, and a publish
        # that cannot read the repo will fail on its own terms a moment later with a better message.
        logger.info(
            "Could not list %s (%s); publishing without the sidecar check.", plan.repo_id, type(exc).__name__
        )
        return
    carried = {path.split("/", 1)[0] for path in plan.files if "/" in path}
    for sidecar in SNAPSHOT_SIDECAR_DIRNAMES:
        if sidecar in carried:
            continue
        orphaned = sorted(f for f in remote if f.startswith(f"{sidecar}/") and f.endswith(".parquet"))
        if not orphaned:
            continue
        raise OrphanedSidecarError(
            f"{plan.repo_id} carries {', '.join(orphaned)}, and this publish would replace "
            f"{RELEASE_FILENAME} without carrying {sidecar}/ — so the published artifact would hold "
            f"two vintages and describe one. Build the missing half into the snapshot before "
            f"publishing (for ClinVar: `just-dna-enricher clinvar citations --out <snapshot>`), or "
            f"publish from a directory that has it."
        )

plan_prune

plan_prune(
    repo_id: str, filename_glob: str, api=None
) -> PrunePlan

What the published repo carries that this lane's snapshot is not made of. Reads; deletes nothing.

Two sources of a name, and neither is a guess. A file under data/ that the lane's own glob excludes is not part of the snapshot by the same definition provisioning uses — _provision_snapshot already refuses to download it and warns when it finds one locally. And a file a LayoutShift declares retired is nameable even where the glob cannot see it: for the four lanes whose glob is *.parquet the exclusion set is empty by construction, so declaration is the only way a retired file there is ever identified.

Everything else is left alone, including files this tier never wrote — README.md, .gitattributes, release.json, LICENSE.txt, and any sidecar directory. A pruner that removed what it did not recognise would be the sweep this design exists to avoid.

Source code in enricher/src/just_dna_enricher/upload.py
def plan_prune(repo_id: str, filename_glob: str, api=None) -> PrunePlan:
    """What the published repo carries that this lane's snapshot is not made of. Reads; deletes nothing.

    **Two sources of a name, and neither is a guess.** A file under `data/` that the lane's own glob
    excludes is not part of the snapshot by the same definition provisioning uses — `_provision_snapshot`
    already refuses to download it and warns when it finds one locally. And a file a `LayoutShift`
    declares retired is nameable even where the glob cannot see it: for the four lanes whose glob is
    `*.parquet` the exclusion set is empty by construction, so declaration is the only way a retired
    file there is ever identified.

    Everything else is left alone, including files this tier never wrote — `README.md`,
    `.gitattributes`, `release.json`, `LICENSE.txt`, and any sidecar directory. A pruner that removed
    what it did not recognise would be the sweep this design exists to avoid.
    """
    if api is None:
        try:
            from huggingface_hub import HfApi, get_token
        except ImportError as exc:
            raise ImportError("huggingface_hub is required to inspect a published repo") from exc
        api = HfApi(token=env_value("HF_TOKEN") or get_token())
    entries = list(api.list_repo_tree(repo_id=repo_id, repo_type="dataset", recursive=True, expand=True))
    sizes = {getattr(e, "path", ""): getattr(e, "size", None) for e in entries}
    remote = [path for path in sizes if path]

    candidates: dict[str, PruneCandidate] = {}
    prefix = f"{SNAPSHOT_DATA_DIRNAME}/"
    for path in sorted(remote):
        if not (path.startswith(prefix) and path.endswith(".parquet")):
            continue
        if fnmatch(path[len(prefix) :], filename_glob):
            continue
        candidates[path] = PruneCandidate(
            path=path,
            size=sizes.get(path),
            reason=f"under {SNAPSHOT_DATA_DIRNAME}/ but outside this snapshot ({filename_glob})",
        )
    for shift in LAYOUT_SHIFTS:
        if shift.repo_id != repo_id or shift.retires not in sizes:
            continue
        # A declared retirement wins the reason slot: it says *when* and *why* the file stopped being
        # part of the snapshot, which "outside the glob" does not.
        candidates[shift.retires] = PruneCandidate(
            path=shift.retires,
            size=sizes.get(shift.retires),
            reason=f"declared retired — {shift.reason}",
        )
    return PrunePlan(repo_id=repo_id, candidates=[candidates[k] for k in sorted(candidates)])

prune_repo

prune_repo(
    plan: PrunePlan,
    token: str | None = None,
    commit_message: str | None = None,
) -> int

Delete a prune plan's files. Called only after the caller has shown the plan and been told yes.

Returns the number of paths deleted. Deleting is a commit on a git-backed repo, so a superseded revision still resolves — but that is a reason to be able to undo a mistake, never a reason to remove something nobody looked at, which is why this takes a plan rather than a repo id.

Source code in enricher/src/just_dna_enricher/upload.py
def prune_repo(plan: PrunePlan, token: str | None = None, commit_message: str | None = None) -> int:
    """Delete a prune plan's files. Called only after the caller has shown the plan and been told yes.

    Returns the number of paths deleted. Deleting is a commit on a git-backed repo, so a superseded
    revision still resolves — but that is a reason to be able to undo a mistake, never a reason to
    remove something nobody looked at, which is why this takes a plan rather than a repo id.
    """
    if not plan.candidates:
        return 0
    api = _hf_api(plan.repo_id, token)
    api.delete_files(
        repo_id=plan.repo_id,
        delete_patterns=[c.path for c in plan.candidates],
        repo_type="dataset",
        commit_message=commit_message
        or (f"Prune {len(plan.candidates)} file(s) that are not part of this snapshot"),
    )
    return len(plan.candidates)

publish_reference_snapshot

publish_reference_snapshot(
    snapshot_dir: Path,
    repo_id: str | None = None,
    token: str | None = None,
    commit_message: str | None = None,
    *,
    payload: str | None = None,
) -> SnapshotPlan

Create-or-update a dataset repo and upload a built reference snapshot to its root.

Uploads data/*.parquet + any parquet sidecars (ClinVar's citations/) + release.json, so the tree matches download.ensure_*_snapshot and a provisioned snapshot is the same artifact a built one is — PMIDs included, which is what a drafted gene panel needs to compile. Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.

Source code in enricher/src/just_dna_enricher/upload.py
def publish_reference_snapshot(
    snapshot_dir: Path,
    repo_id: str | None = None,
    token: str | None = None,
    commit_message: str | None = None,
    *,
    payload: str | None = None,
) -> SnapshotPlan:
    """Create-or-update a dataset repo and upload a built reference snapshot to its root.

    Uploads ``data/*.parquet`` + any parquet sidecars (ClinVar's ``citations/``) + ``release.json``, so
    the tree matches ``download.ensure_*_snapshot`` and a provisioned snapshot is the same artifact a
    built one is — PMIDs included, which is what a drafted gene panel needs to compile.
    Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.
    """
    plan = plan_reference_snapshot(snapshot_dir, repo_id, payload=payload)
    api = ensure_repo(plan.repo_id, token)
    check_publish_orphans_no_sidecar(plan, api)
    # A declared retirement rides in the same commit as the upload that replaces it (RM186). Listed
    # once and reused: `check_publish_orphans_no_sidecar` has already read the repo, but its answer is
    # about sidecars, and a second listing is cheaper to reason about than a shared cache of a remote
    # state two checks disagree about the meaning of.
    try:
        remote = list(api.list_repo_files(repo_id=plan.repo_id, repo_type="dataset"))
    except Exception as exc:
        logger.info(
            "Could not list %s (%s); no declared retirement can apply.", plan.repo_id, type(exc).__name__
        )
        remote = []
    due = layout_shifts_to_apply(remote, plan.repo_id)
    if due:
        logger.info(
            "Retiring %s from %s in this commit: %s",
            ", ".join(shift.retires for shift in due),
            plan.repo_id,
            "; ".join(shift.reason for shift in due),
        )

    # **One `upload_folder`, and the size branch is gone** (RM199, revised 2026-09-11).
    #
    # This used to pick `upload_large_folder` above 5 GB, because `upload_folder` was a single
    # non-resumable commit and 32 GB in one of those starts over on any failure. `huggingface_hub`
    # 1.x settled it the other way: **`upload_folder` is itself multi-commit now**, and
    # `upload_large_folder` emits a `FutureWarning` telling callers to stop using it. Found the way
    # deprecations should be — the warning appeared in a real publish.
    #
    # Two things fall out. The threshold and its two constants are gone, and so is the refusal this
    # function used to raise when a declared retirement met the large path: `upload_folder` takes
    # `delete_patterns`, so RM186's arrival-and-departure stays on one call and there is no longer a
    # collision to refuse. What remains upstream's rather than ours is that a *large* upload is
    # several commits either way — so the one-commit guarantee holds for the payloads that fit in one
    # and is the Hub's business for the payloads that do not.
    #
    # **The description still goes last** (RM199's real content). `release.json` is what a puller reads to learn
    # which release it holds, so it must never arrive before the bytes it describes: a publish that
    # lands the description and then fails leaves a snapshot that *reads as* provisioned and is not.
    # That is `@a-publish-may-not-orphan-the-bytes-it-stops-describing` from the other direction, and
    # it costs one extra commit on a path that is already several.
    payload_files = [f for f in plan.files if f != RELEASE_FILENAME]
    description = [f for f in plan.files if f == RELEASE_FILENAME]

    api.upload_folder(
        folder_path=str(snapshot_dir),
        path_in_repo="",
        repo_id=plan.repo_id,
        repo_type="dataset",
        # The retirement is a `delete_patterns` on the same call, so the new file arriving and the old
        # one leaving are one operation rather than a window in which the repo has neither or both.
        delete_patterns=[shift.retires for shift in due] or None,
        # Derived from the plan rather than restated as a pattern list. The two had to agree and did
        # not: `--dry-run` printed a file the patterns then dropped, which is the failure mode that
        # lost `citations/` and `LICENSE.txt` in the first place. One list, computed once, so what a
        # dry run promises is exactly what an upload sends (`@publisher-allowlist-derived`).
        allow_patterns=payload_files,
        commit_message=commit_message
        or (
            f"Publish reference snapshot ({len(payload_files)} files)"
            + (f", retiring {', '.join(s.retires for s in due)}" if due else "")
        ),
    )

    if description:
        api.upload_folder(
            folder_path=str(snapshot_dir),
            path_in_repo="",
            repo_id=plan.repo_id,
            repo_type="dataset",
            allow_patterns=description,
            commit_message=f"Describe the snapshot ({RELEASE_FILENAME})",
        )
    return plan