just_dna_enricher.upload¶
just_dna_enricher.upload ¶
HuggingFace upload — publisher surface of the network tier.
Two publish shapes share one create-or-update pathway (ensure_repo):
- module — a compiled module's artifacts (every parquet the artifact carries + manifest.json,
and a logo and readme if present) to
datasets/<repo>/data/<name>/, matching the layout just-dna-lite's discovery scans, and todata/<name>/v<version>/when the manifest states a version, so the discovery path can address a particular release rather than only "whatever is there now" (upload_module; RM84). - reference snapshot — a built ClinVar (or Ensembl) parquet snapshot +
release.jsonto the root of a dataset repo, matching thedownload.ensure_*_snapshotlayout (publish_reference_snapshot).
Both require a HuggingFace token with write access (hf auth login or HF_TOKEN). This is the
dev/publisher half of the enricher's HuggingFace use: snapshot download is a runtime enrich
path (download.ensure_snapshot); upload is for authors republishing modules (e.g. the Gen-I
v1-port recreation) or publishing a rebuilt reference. Install as just-dna-enricher[dev].
Extracted from just_dna_pipelines.v1_port.publish (just-dna-lite); create_repo was added here
(the origin assumed the repo pre-existed) so a brand-new dataset repo can be created on first push.
LayoutShift
dataclass
¶
A retirement declared by the commit that causes it, so the publish carries out the migration.
The rule (maintainer, 2026-09-03). A change that retires one published file and introduces another must also carry a conditioned update to the publication procedure: if the new file is absent from the repo and the old one is present, upload the new and delete the old. The condition is a predicate over the remote, so it fires exactly once per repo and is a no-op forever after — nothing is deleted on a repo that has already moved, and re-running a publish cannot re-delete.
This is the only deletion a publish performs. General cleanup is cache prune, which prints
what it would remove and asks; a deletion nobody declared is not a side effect a publish may have.
What makes this one safe is not that it is small but that it is named: the pair is in the commit
that changed the layout, where a reviewer sees both halves at once.
Deleting is recoverable in a way this repository's @snapshot-accumulates note did not credit —
a HuggingFace dataset repo is git-backed, and a superseded revision still resolves (three of them
were read off the hub while auditing this on 2026-09-03). The reason to declare rather than sweep
is not that bytes are lost; it is that a retired file goes on answering 200 to whoever still asks
for it, which is how a lane's default archive stayed frozen for a year (CLINPGX_ARCHIVES).
OrphanedSidecarError ¶
Bases: RuntimeError
The publish would overwrite a release.json describing sidecars it is not carrying (RM185).
Its own type for the reason PublishCollisionError is one: the CLI has to tell it from the
refusals plan_* raises. Those say the local snapshot is not publishable; this one says the local
snapshot is fine and the remote holds bytes this publish would silently stop describing.
The incident. A ClinVar snapshot is two halves on two cadences — the per-chromosome parquet and
citations/citations.parquet — and release.json describes both because build_citations merges
a block into it. A rebuild that built only the VCF half published a citations-free release.json
over a repo whose sidecar it had not replaced, and the publisher adds without deleting: the
sidecar outlived its own description, and the published artifact carried records from one ClinVar
release beside citations from another while saying nothing. RM179 fixed the lane; this is the
general shape, refused at the boundary where any lane could repeat it.
PublishCollisionError ¶
Bases: RuntimeError
The versioned path already holds a different release, and no --force was given (RM88).
Its own type rather than a FileNotFoundError, because the CLI has to tell it apart from the
three refusals plan_upload raises: those say the local module is not publishable, this one says
the local module is fine and the remote already has that version.
UploadPlan ¶
Bases: BaseModel
What an upload would send (also the dry-run result).
SnapshotPlan ¶
Bases: BaseModel
What a reference-snapshot publish would send (also the dry-run result).
PruneCandidate ¶
Bases: BaseModel
One remote file cache prune would remove, and the reason it is nameable as removable.
PrunePlan ¶
Bases: BaseModel
What a prune would delete from one repo. The dry run and the deletion read the same object.
ensure_repo ¶
Create-or-update: ensure the dataset repo exists, returning the authenticated HfApi.
create_repo(..., exist_ok=True) is a no-op when the repo already exists, so create and update
are one pathway. The returned api is reused by the caller's upload_folder so only one
HfApi is constructed per publish.
Source code in enricher/src/just_dna_enricher/upload.py
layout_shifts_to_apply ¶
The declared shifts whose condition the remote currently satisfies. Reads; decides nothing else.
Separate from the publish so it can be asserted directly: the whole safety of this mechanism is that the predicate is false the moment it has run once, and that is a property of this function rather than of the caller that acts on it.
Source code in enricher/src/just_dna_enricher/upload.py
plan_upload ¶
Resolve the upload plan and validate the compiled artifacts are present.
Two destinations, both planned here (RM84): the flat data/<name>/, which is what the reference
consumer's discovery scans and therefore has to keep working and keep meaning latest, and
data/<name>/v<version>/ inside it — nested rather than a sibling, because the flat path is
the one deployed readers are pointed at — which is the only thing on that path that can name a
particular release. A module with no readable version gets the flat path alone, never a vNone
directory, and version_unknown_reason says which of the four reasons applies. The reason is a
field rather than a log line for the reason every other withheld answer here is
(clin_sig_not_checked, skipped_offline): a caller that has to parse prose to learn what
happened has not been told.
What it refuses is three positive rules, not one required list (RM89). The old rule demanded the three SNP-core parquets, which RM2 made optional four releases ago, so seven of the sixteen reference examples could not be published at all. In their place, ordered most specific first so a refusal names the actual fault:
- The plan carries everything the artifact attests.
manifest.artifact.filesstates a name, a sha256 and a size per parquet, andartifact.digestis a Merkle root over exactly those — so a file attested and not sent makes the published manifest a false claim about bytes that are not there. This is the general guard, and it is a self-check as much as a module check: it fires if this publisher's allowlist ever falls behind the compiler's output list again. weights.parquetnever travels alone. A SNP core compiles to all three, so a missingannotations/studiesbeside it is an interrupted compile.- At least one lead table. A directory with a manifest, a readme and no annotation rows is not a module; publishing one is the silent failure a bare widening of the old rule would have produced, and it is what discovery would then be unable to see.
Source code in enricher/src/just_dna_enricher/upload.py
upload_module ¶
upload_module(
module_dir: Path,
name: str,
repo_id: str | None = None,
token: str | None = None,
commit_message: str | None = None,
force: bool = False,
) -> UploadPlan
Upload the compiled module to a HuggingFace dataset collection.
Ensures the dataset repo exists (create-or-update), then writes the same files to the flat path and — when the manifest states a version — to the versioned subdirectory under it.
The two writes are two commits, not one. upload_folder commits per call, so a reader can
briefly see the flat path refreshed while the versioned copy is not there yet; the flat path goes
first because it is the one anything reads today. Making it atomic means create_commit over an
explicit operation list, which is a different shape from the allow-pattern plumbing every publish
here uses, and it was not worth holding the fix for. If the second call fails the first has already
landed — data/<name>/ is the new release and the versioned copy is absent until a re-run, which
is idempotent.
The versioned path refuses to be overwritten with different bytes (RM88). Before either
write, the published manifest.json at data/<name>/v<version>/ is read and its
artifact.digest compared with this module's. A different digest is a PublishCollisionError
unless force=True. The flat path is deliberately not guarded: it means latest, so
overwriting it is what it is for, and the whole point of the versioned copy is that it does not.
The policy is refuse-unless---force rather than warn-and-proceed or refuse-outright, decided
2026-08-18. The flag's existence is itself the claim that overwriting is sometimes right — a
curator re-cutting a draft release is a real workflow, and a gate with no override becomes a gate
people route around.
Recompiling under a newer compiler also moves the digest, and will trip this. That is correct
rather than a false positive: P4 scopes byte-reproducibility to a fixed compiler_version, so the
versioned path really would come to hold different bytes than the ones it was published with. The
refusal says so, because "I changed nothing" is the first thing its first user will think.
Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.
Source code in enricher/src/just_dna_enricher/upload.py
plan_reference_snapshot ¶
plan_reference_snapshot(
snapshot_dir: Path,
repo_id: str | None = None,
*,
payload: str | None = None,
) -> SnapshotPlan
Resolve a snapshot publish plan and validate the built artifacts are present.
payload names a snapshot whose content is one file at the root rather than
data/*.parquet — STRchive's STRchive-loci.json is the case, and ACMG's acmg_sf.csv is the
shape's second member. The parameter is the caller's, deliberately: a lane knows the name of the
file it builds, and putting a roster of lane filenames in the publisher would make this function
the fourth place that has to learn about a new snapshot kind.
A payload snapshot is refused when the file is missing, exactly as a parquet one is when
data/ is empty — the refusal is the same claim either way, that there is nothing built here to
publish.
Source code in enricher/src/just_dna_enricher/upload.py
check_publish_orphans_no_sidecar ¶
Refuse a publish that would leave a remote sidecar undescribed. Reads the repo; writes nothing.
The remote tree, not the remote release.json. The block is a description of the bytes and
the bytes are what a puller gets, so the description is the half that can already be missing —
which is exactly the state just-dna-seq/clinvar was found in on 2026-09-03: the citations parquet
present, the block gone. A guard reading the block would have passed the second bad publish as
happily as the first.
Scoped to publishes that carry release.json, because a publish that carries none overwrites no
provenance. A repo that does not exist yet lists nothing and passes — that is a first publish, not
an orphan.
api is the caller's when it has one (a publish has just built an authenticated client and there
is no reason to build a second); a dry run passes none and gets an anonymous reader, since listing
a public repo needs no token and a dry run must not require write access to say what a publish
would do.
Source code in enricher/src/just_dna_enricher/upload.py
plan_prune ¶
What the published repo carries that this lane's snapshot is not made of. Reads; deletes nothing.
Two sources of a name, and neither is a guess. A file under data/ that the lane's own glob
excludes is not part of the snapshot by the same definition provisioning uses — _provision_snapshot
already refuses to download it and warns when it finds one locally. And a file a LayoutShift
declares retired is nameable even where the glob cannot see it: for the four lanes whose glob is
*.parquet the exclusion set is empty by construction, so declaration is the only way a retired
file there is ever identified.
Everything else is left alone, including files this tier never wrote — README.md,
.gitattributes, release.json, LICENSE.txt, and any sidecar directory. A pruner that removed
what it did not recognise would be the sweep this design exists to avoid.
Source code in enricher/src/just_dna_enricher/upload.py
prune_repo ¶
Delete a prune plan's files. Called only after the caller has shown the plan and been told yes.
Returns the number of paths deleted. Deleting is a commit on a git-backed repo, so a superseded revision still resolves — but that is a reason to be able to undo a mistake, never a reason to remove something nobody looked at, which is why this takes a plan rather than a repo id.
Source code in enricher/src/just_dna_enricher/upload.py
publish_reference_snapshot ¶
publish_reference_snapshot(
snapshot_dir: Path,
repo_id: str | None = None,
token: str | None = None,
commit_message: str | None = None,
*,
payload: str | None = None,
) -> SnapshotPlan
Create-or-update a dataset repo and upload a built reference snapshot to its root.
Uploads data/*.parquet + any parquet sidecars (ClinVar's citations/) + release.json, so
the tree matches download.ensure_*_snapshot and a provisioned snapshot is the same artifact a
built one is — PMIDs included, which is what a drafted gene panel needs to compile.
Raises PermissionError if no token is available and ImportError if huggingface_hub is absent.
Source code in enricher/src/just_dna_enricher/upload.py
784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 | |