Skip to content

just_dna_enricher.locations

just_dna_enricher.locations

Ensembl reference cache resolution — mirrors just-dna-lite's on-disk layout so a marketplace or standalone compile can reuse an existing just-dna-lite deployment's cache (disk economy, no re-download), pointed via .env.

Layout (identical to just-dna-pipelines)::

<base>/ensembl_variations/data/*.parquet
<base>/ensembl_variations/ensembl_variations.duckdb    # optional prebuilt view

where <base> is $JUST_DNA_PIPELINES_CACHE_DIR (the same var just-dna-lite uses), or the platformdirs user cache for "just-dna-pipelines". $JUST_DNA_ENSEMBL_CACHE (a .duckdb file or a directory) overrides everything for explicit pointing.

This module never downloads: if no cache is present, resolution returns None and the resolver skips with a warning. Provisioning the reference is the deployment's job.

repro_out

repro_out(name: str) -> Path

The default --out for a builder, under data/repro/<name>/.

Derived rather than restated, because a rule spelled out once per command is a rule that holds until somebody adds the tenth command. An AST guard over the CLI asserts every --out default resolves under data/, so a new builder that writes its own literal fails the suite rather than the operator's working tree.

Source code in enricher/src/just_dna_enricher/locations.py
def repro_out(name: str) -> Path:
    """The default `--out` for a builder, under `data/repro/<name>/`.

    Derived rather than restated, because a rule spelled out once per command is a rule that holds
    until somebody adds the tenth command. An AST guard over the CLI asserts every `--out` default
    resolves under `data/`, so a new builder that writes its own literal fails the suite rather than
    the operator's working tree.
    """
    return Path(REPRO_DIRNAME) / name

read_release

read_release(reference: Path) -> dict | None

A snapshot's release.json as a dict, or None when it is absent or unreadable.

Sixth party to the layout agreement above, and the first reader of it outside cache status: the file was written by every builder and consulted by nothing, so a caller who needed to know which release a local snapshot is could only guess. None is the honest answer for both absence and corruption — a caller must not be able to mistake "this snapshot does not say" for a release id, so callers branch on None rather than on a default (the tri-state rule).

Source code in enricher/src/just_dna_enricher/locations.py
def read_release(reference: Path) -> dict | None:
    """A snapshot's `release.json` as a dict, or `None` when it is absent or unreadable.

    Sixth party to the layout agreement above, and the first *reader* of it outside `cache status`:
    the file was written by every builder and consulted by nothing, so a caller who needed to know
    which release a local snapshot is could only guess. `None` is the honest answer for both absence
    and corruption — a caller must not be able to mistake "this snapshot does not say" for a release
    id, so callers branch on `None` rather than on a default (the tri-state rule).
    """
    path = Path(reference) / RELEASE_FILENAME
    if not path.is_file():
        return None
    try:
        payload = json.loads(path.read_text(encoding="utf-8"))
    except (OSError, json.JSONDecodeError) as exc:
        logger.warning("Could not read %s (%s); treating the release as unstated.", path, exc)
        return None
    return payload if isinstance(payload, dict) else None

load_env

load_env(override: bool = False) -> str | None

Load the nearest .env (walking up from CWD) into os.environ, so cache paths can be set there. Returns the loaded path, or None.

Every variable in the file is exported, so only the cache resolvers call this. A credential is read with env_value, which exports nothing (RM301, S124).

Source code in enricher/src/just_dna_enricher/locations.py
def load_env(override: bool = False) -> str | None:
    """Load the nearest `.env` (walking up from CWD) into `os.environ`, so cache paths can be set there.
    Returns the loaded path, or None.

    **Every variable in the file is exported, so only the cache resolvers call this.** A credential is
    read with `env_value`, which exports nothing (RM301, S124).
    """
    env_path = find_dotenv(usecwd=True)
    if env_path:
        load_dotenv(env_path, override=override)
        return env_path
    return None

env_value

env_value(var: str) -> str | None

One variable, from the process environment or else the nearest .env, without exporting the file.

This is how a credential is read where it is used (@credential-where-read). The precedence is load_env's: override=False keeps a variable that is present, so an exported value wins over the file, and an exported empty string stays empty. The difference is that nothing is written into os.environ. RM301 (S124): every client used to call load_env(), which copied the whole .env into the host's environment. A host that reports which layer each of its own settings came from then saw every file value as an exported shell variable.

The file is found per call, walking up from the CWD, like load_env. None means neither source names the variable. An empty string means one of them sets it empty.

Source code in enricher/src/just_dna_enricher/locations.py
def env_value(var: str) -> str | None:
    """One variable, from the process environment or else the nearest `.env`, **without exporting the file**.

    This is how a credential is read where it is used (`@credential-where-read`). The precedence is
    `load_env`'s: `override=False` keeps a variable that is present, so an exported value wins over the
    file, and an exported empty string stays empty. The difference is that nothing is written into
    `os.environ`. RM301 (S124): every client used to call `load_env()`, which copied the *whole* `.env`
    into the host's environment. A host that reports which layer each of its own settings came from
    then saw every file value as an exported shell variable.

    The file is found per call, walking up from the CWD, like `load_env`. `None` means neither source
    names the variable. An empty string means one of them sets it empty.
    """
    if var in os.environ:
        return os.environ[var]
    path = find_dotenv(usecwd=True)
    if not path:
        return None
    return dotenv_values(path).get(var)

missing_credential_reason

missing_credential_reason(var: str) -> str

Why $var is unusable — absent and exported empty are two states, not one.

load_env uses override=False, so a variable that is present is kept whatever the .env says — and an empty string is present. That makes export FOO= strictly stronger than deleting the variable, which is the opposite of what anyone expects and is the same edge the tier's own tests exploit deliberately (a test neutralizes a credential with "" precisely because delenv would let the developer's .env refill it).

It bites for real: a shell that ran a snippet whose placeholder was edited out — export PHARMVAR_API_KEY= — reports no key for the rest of the session on a machine whose .env holds a working one, and nothing in the message says why. So the two readings are named separately and the empty one carries its own remedy, because unset and "go and get a key" are different actions (@rsid-absent-two-readings is the same rule about a different absence).

Source code in enricher/src/just_dna_enricher/locations.py
def missing_credential_reason(var: str) -> str:
    """Why `$var` is unusable — **absent** and **exported empty** are two states, not one.

    `load_env` uses ``override=False``, so a variable that is *present* is kept whatever the `.env`
    says — and an empty string is present. That makes `export FOO=` strictly stronger than deleting
    the variable, which is the opposite of what anyone expects and is the same edge the tier's own
    tests exploit deliberately (a test neutralizes a credential with `""` precisely because `delenv`
    would let the developer's `.env` refill it).

    It bites for real: a shell that ran a snippet whose placeholder was edited out — `export
    PHARMVAR_API_KEY=` — reports *no key* for the rest of the session on a machine whose `.env` holds
    a working one, and nothing in the message says why. So the two readings are named separately and
    the empty one carries its own remedy, because `unset` and "go and get a key" are different
    actions (`@rsid-absent-two-readings` is the same rule about a different absence).
    """
    value = os.getenv(var)
    if value is None and env_value(var) == "":
        return (
            f"${var} is set EMPTY in the `.env` beside the working directory. Give it a value there, "
            f"or export it"
        )
    if value is None:
        return (
            f"no ${var} is set. A `.env` beside the working directory is read automatically, so "
            f"either add it there or export it"
        )
    return (
        f"${var} is set but EMPTY, and an empty exported variable outranks a `.env`: the loader uses "
        f"override=False, so it keeps a variable that is present. Run `unset {var}` — deleting it is "
        f"what lets the file supply the real one"
    )

default_ensembl_cache_dir

default_ensembl_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/ensembl_variations directory, matching just-dna-lite's convention.

Source code in enricher/src/just_dna_enricher/locations.py
def default_ensembl_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/ensembl_variations` directory, matching just-dna-lite's convention."""
    return _cache_dir(ENSEMBL_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_ensembl_reference

resolve_ensembl_reference(
    ensembl_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a usable Ensembl reference without downloading.

Precedence: explicit ensembl_cache → $JUST_DNA_ENSEMBL_CACHE → the just-dna-lite layout under $JUST_DNA_PIPELINES_CACHE_DIR / platformdirs. Prefers a prebuilt ensembl_variations.duckdb; otherwise the directory of parquet files. Returns the resolved path (a .duckdb file or a directory), or None if nothing is present.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_ensembl_reference(
    ensembl_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a usable Ensembl reference without downloading.

    Precedence: explicit `ensembl_cache` → ``$JUST_DNA_ENSEMBL_CACHE`` → the just-dna-lite layout
    under ``$JUST_DNA_PIPELINES_CACHE_DIR`` / platformdirs. Prefers a prebuilt
    ``ensembl_variations.duckdb``; otherwise the directory of parquet files. Returns the resolved
    path (a ``.duckdb`` file or a directory), or ``None`` if nothing is present.
    """
    if load_dotenv_file:
        load_env()

    candidate = ensembl_cache or os.getenv(ENSEMBL_CACHE_VAR)
    search_dir = (
        Path(candidate) if candidate else default_ensembl_cache_dir(load_dotenv_file=load_dotenv_file)
    )

    # Explicit pointing at a specific DuckDB file.
    if search_dir.is_file() and search_dir.suffix == ".duckdb":
        return search_dir

    # Otherwise return the cache directory if it holds a prebuilt db or parquet data; the
    # connection layer decides whether the db is usable and falls back to parquet if not.
    if search_dir.is_dir():
        data_dir = search_dir / "data"
        has_db = (search_dir / DUCKDB_NAME).is_file()
        has_parquet = (data_dir.is_dir() and any(data_dir.glob("*.parquet"))) or any(
            search_dir.glob("*.parquet")
        )
        if has_db or has_parquet:
            return search_dir
    return None

default_clinvar_cache_dir

default_clinvar_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/clinvar directory (same base as the Ensembl cache).

Source code in enricher/src/just_dna_enricher/locations.py
def default_clinvar_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/clinvar` directory (same base as the Ensembl cache)."""
    return _cache_dir(CLINVAR_SUBDIR, load_dotenv_file=load_dotenv_file)

default_constraint_cache_dir

default_constraint_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/gnomad_constraint directory (same base as the other two caches).

Source code in enricher/src/just_dna_enricher/locations.py
def default_constraint_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/gnomad_constraint` directory (same base as the other two caches)."""
    return _cache_dir(CONSTRAINT_SUBDIR, load_dotenv_file=load_dotenv_file)

default_clinpgx_cache_dir

default_clinpgx_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/clinpgx directory — the ClinPGx clinical-annotation snapshot.

Source code in enricher/src/just_dna_enricher/locations.py
def default_clinpgx_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/clinpgx` directory — the ClinPGx clinical-annotation snapshot."""
    return _cache_dir(CLINPGX_SUBDIR, load_dotenv_file=load_dotenv_file)

default_cpic_cache_dir

default_cpic_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/cpic directory — the CPIC allele/diplotype/recommendation snapshot.

Source code in enricher/src/just_dna_enricher/locations.py
def default_cpic_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/cpic` directory — the CPIC allele/diplotype/recommendation snapshot."""
    return _cache_dir(CPIC_SUBDIR, load_dotenv_file=load_dotenv_file)

default_pharmvar_cache_dir

default_pharmvar_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/pharmvar directory — operator-built only (see PHARMVAR_SUBDIR).

Source code in enricher/src/just_dna_enricher/locations.py
def default_pharmvar_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/pharmvar` directory — operator-built only (see `PHARMVAR_SUBDIR`)."""
    return _cache_dir(PHARMVAR_SUBDIR, load_dotenv_file=load_dotenv_file)

default_pubmind_cache_dir

default_pubmind_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/pubmind directory — operator-built only (see PUBMIND_SUBDIR).

Source code in enricher/src/just_dna_enricher/locations.py
def default_pubmind_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/pubmind` directory — operator-built only (see `PUBMIND_SUBDIR`)."""
    return _cache_dir(PUBMIND_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_constraint_reference

resolve_constraint_reference(
    constraint_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a usable gnomAD constraint snapshot without downloading.

Explicit argument → $JUST_DNA_GNOMAD_CONSTRAINT_CACHE → $JUST_DNA_PIPELINES_CACHE_DIR /platformdirs gnomad_constraint/. Parquet only (like ClinVar, there is no prebuilt .duckdb), and a bare .parquet may be pointed at directly since the snapshot is one file.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_constraint_reference(
    constraint_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a usable gnomAD constraint snapshot without downloading.

    Explicit argument → ``$JUST_DNA_GNOMAD_CONSTRAINT_CACHE`` → ``$JUST_DNA_PIPELINES_CACHE_DIR``
    /platformdirs ``gnomad_constraint/``. Parquet only (like ClinVar, there is no prebuilt
    ``.duckdb``), and a bare ``.parquet`` may be pointed at directly since the snapshot is one file.
    """
    return _resolve_parquet_cache(
        constraint_cache,
        CONSTRAINT_CACHE_VAR,
        default_constraint_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
        accept_bare_file=True,
    )

resolve_clinvar_reference

resolve_clinvar_reference(
    clinvar_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a usable ClinVar reference without downloading.

Mirrors resolve_ensembl_reference's precedence: explicit clinvar_cache → $JUST_DNA_CLINVAR_CACHE → $JUST_DNA_PIPELINES_CACHE_DIR/platformdirs clinvar/. The ClinVar snapshot ships as parquet only (no prebuilt .duckdb); a directory is returned when it holds data/*.parquet (or bare *.parquet), else None. Never downloads — provisioning is the enricher's download.ensure_clinvar_snapshot or the deployment's job.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_clinvar_reference(
    clinvar_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a usable ClinVar reference without downloading.

    Mirrors `resolve_ensembl_reference`'s precedence: explicit `clinvar_cache` →
    ``$JUST_DNA_CLINVAR_CACHE`` → ``$JUST_DNA_PIPELINES_CACHE_DIR``/platformdirs `clinvar/`. The
    ClinVar snapshot ships as parquet only (no prebuilt ``.duckdb``); a directory is returned when it
    holds ``data/*.parquet`` (or bare ``*.parquet``), else ``None``. Never downloads — provisioning is
    the enricher's `download.ensure_clinvar_snapshot` or the deployment's job.
    """
    return _resolve_parquet_cache(
        clinvar_cache,
        CLINVAR_CACHE_VAR,
        default_clinvar_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

resolve_clinpgx_reference

resolve_clinpgx_reference(
    clinpgx_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built ClinPGx snapshot without downloading ($JUST_DNA_CLINPGX_CACHE).

The builder shipped a release before this existed, so the snapshot was reachable only by handing enrich_clinpgx an explicit --snapshot; with no path the pass skipped itself and said so, which on a hosted deployment is the check simply not running. Provisioning is download.ensure_clinpgx_snapshot.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_clinpgx_reference(
    clinpgx_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built ClinPGx snapshot without downloading (`$JUST_DNA_CLINPGX_CACHE`).

    The builder shipped a release before this existed, so the snapshot was reachable only by handing
    `enrich_clinpgx` an explicit `--snapshot`; with no path the pass skipped itself and said so, which
    on a hosted deployment is the check simply not running. Provisioning is
    `download.ensure_clinpgx_snapshot`.
    """
    return _resolve_parquet_cache(
        clinpgx_cache,
        CLINPGX_CACHE_VAR,
        default_clinpgx_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

resolve_cpic_reference

resolve_cpic_reference(
    cpic_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built CPIC snapshot without downloading ($JUST_DNA_CPIC_CACHE).

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_cpic_reference(cpic_cache: Path | None = None, *, load_dotenv_file: bool = True) -> Path | None:
    """Locate a built CPIC snapshot without downloading (`$JUST_DNA_CPIC_CACHE`)."""
    return _resolve_parquet_cache(
        cpic_cache,
        CPIC_CACHE_VAR,
        default_cpic_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

resolve_pharmvar_reference

resolve_pharmvar_reference(
    pharmvar_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate an operator-built PharmVar snapshot ($JUST_DNA_PHARMVAR_CACHE).

There is deliberately no download.ensure_pharmvar_snapshot to pair with this — see PHARMVAR_SUBDIR. A deployment builds its own with its own key, or this returns None and the PharmVar leg degrades exactly as it does when no key is configured.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_pharmvar_reference(
    pharmvar_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate an **operator-built** PharmVar snapshot (`$JUST_DNA_PHARMVAR_CACHE`).

    There is deliberately no `download.ensure_pharmvar_snapshot` to pair with this — see
    `PHARMVAR_SUBDIR`. A deployment builds its own with its own key, or this returns `None` and the
    PharmVar leg degrades exactly as it does when no key is configured.
    """
    return _resolve_parquet_cache(
        pharmvar_cache,
        PHARMVAR_CACHE_VAR,
        default_pharmvar_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_civic_cache_dir

default_civic_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/civic directory — operator-built for now (see CIVIC_SUBDIR).

Source code in enricher/src/just_dna_enricher/locations.py
def default_civic_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/civic` directory — operator-built for now (see `CIVIC_SUBDIR`)."""
    return _cache_dir(CIVIC_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_civic_reference

resolve_civic_reference(
    civic_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate an operator-built CIViC snapshot ($JUST_DNA_CIVIC_CACHE).

None when there is none, and the drafter reads that as nobody-asked rather than as an empty source (@unreachable-not-absent). Build one with civic build --release <date>.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_civic_reference(civic_cache: Path | None = None, *, load_dotenv_file: bool = True) -> Path | None:
    """Locate an operator-built CIViC snapshot (`$JUST_DNA_CIVIC_CACHE`).

    `None` when there is none, and the drafter reads that as nobody-asked rather than as an empty
    source (`@unreachable-not-absent`). Build one with `civic build --release <date>`.
    """
    return _resolve_parquet_cache(
        civic_cache,
        CIVIC_CACHE_VAR,
        default_civic_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

resolve_pubmind_reference

resolve_pubmind_reference(
    pubmind_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate an operator-built PubMind snapshot ($JUST_DNA_PUBMIND_CACHE).

There is deliberately no download.ensure_pubmind_snapshot to pair with this — see PUBMIND_SUBDIR. A deployment builds its own with pubmind build, or this returns None and the PubMind leg reads unchecked rather than absent: nobody asked is a third state beside asked-and-failed and asked-and-absent (@unreachable-not-absent).

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_pubmind_reference(
    pubmind_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate an **operator-built** PubMind snapshot (`$JUST_DNA_PUBMIND_CACHE`).

    There is deliberately no `download.ensure_pubmind_snapshot` to pair with this — see
    `PUBMIND_SUBDIR`. A deployment builds its own with `pubmind build`, or this returns `None` and the
    PubMind leg reads `unchecked` rather than absent: nobody asked is a third state beside
    asked-and-failed and asked-and-absent (`@unreachable-not-absent`).
    """
    return _resolve_parquet_cache(
        pubmind_cache,
        PUBMIND_CACHE_VAR,
        default_pubmind_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_mane_cache_dir

default_mane_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/mane directory — operator-built only (see MANE_SUBDIR).

Source code in enricher/src/just_dna_enricher/locations.py
def default_mane_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/mane` directory — operator-built only (see `MANE_SUBDIR`)."""
    return _cache_dir(MANE_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_mane_reference

resolve_mane_reference(
    mane_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate an operator-built MANE snapshot ($JUST_DNA_MANE_CACHE).

There is deliberately no download.ensure_mane_snapshot to pair with this: nothing publishes a MANE snapshot, and NCBI's policy neither grants nor withholds permission to (see MANE_SUBDIR). None when there is none, and a reader must treat that as nobody-asked rather than as a gene MANE has no transcript for — the snapshot's own negative roster is what answers the second question, and it can only answer it once the snapshot exists (@unreachable-not-absent).

A bare .parquet is not accepted here, unlike the constraint snapshot: this cache is three tables read by filename, so a single file pointed at directly would be a snapshot missing the currency check and the negative roster.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_mane_reference(mane_cache: Path | None = None, *, load_dotenv_file: bool = True) -> Path | None:
    """Locate an **operator-built** MANE snapshot (`$JUST_DNA_MANE_CACHE`).

    There is deliberately no `download.ensure_mane_snapshot` to pair with this: nothing publishes a
    MANE snapshot, and NCBI's policy neither grants nor withholds permission to (see `MANE_SUBDIR`).
    `None` when there is none, and a reader must treat that as nobody-asked rather than as a gene
    MANE has no transcript for — the snapshot's own negative roster is what answers the second
    question, and it can only answer it once the snapshot exists (`@unreachable-not-absent`).

    A bare `.parquet` is **not** accepted here, unlike the constraint snapshot: this cache is three
    tables read by filename, so a single file pointed at directly would be a snapshot missing the
    currency check and the negative roster.
    """
    return _resolve_parquet_cache(
        mane_cache,
        MANE_CACHE_VAR,
        default_mane_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_acmg_cache_dir

default_acmg_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/acmg_sf directory — operator-built only (see ACMG_SUBDIR).

Source code in enricher/src/just_dna_enricher/locations.py
def default_acmg_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/acmg_sf` directory — operator-built only (see `ACMG_SUBDIR`)."""
    return _cache_dir(ACMG_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_acmg_reference

resolve_acmg_reference(
    acmg_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built ACMG SF snapshot ($JUST_DNA_ACMG_CACHE), without downloading.

None means nobody provisioned one, and check-acmg reads that as leave to fall back to scraping NCBI's page — which serves v3.2 while ACMG published v3.3 in June 2025, so a correctly authored v3.3 row is reported as wrong by the fallback. That is the whole reason this resolver exists: the snapshot was buildable and unreachable unless a caller passed --sf-list by hand, so the accurate list was the one path a deployment never took.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_acmg_reference(acmg_cache: Path | None = None, *, load_dotenv_file: bool = True) -> Path | None:
    """Locate a built ACMG SF snapshot (`$JUST_DNA_ACMG_CACHE`), without downloading.

    `None` means nobody provisioned one, and `check-acmg` reads that as leave to fall back to
    scraping NCBI's page — which serves **v3.2** while ACMG published v3.3 in June 2025, so a
    correctly authored v3.3 row is reported as wrong by the fallback. That is the whole reason this
    resolver exists: the snapshot was buildable and unreachable unless a caller passed `--sf-list` by
    hand, so the accurate list was the one path a deployment never took.
    """
    return _resolve_named_cache(
        acmg_cache,
        ACMG_CACHE_VAR,
        default_acmg_cache_dir(load_dotenv_file=load_dotenv_file),
        ACMG_SNAPSHOT_FILENAME,
        load_dotenv_file=load_dotenv_file,
    )

default_strchive_cache_dir

default_strchive_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/strchive directory — see STRCHIVE_SUBDIR.

Source code in enricher/src/just_dna_enricher/locations.py
def default_strchive_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/strchive` directory — see `STRCHIVE_SUBDIR`."""
    return _cache_dir(STRCHIVE_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_strchive_reference

resolve_strchive_reference(
    strchive_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built STRchive catalogue ($JUST_DNA_STRCHIVE_CACHE), without downloading.

None is nobody-asked rather than "STRchive lists no such locus", and check-repeat-bands says so in those words (@unreachable-not-absent). A directory is preferred over a bare STRchive-loci.json because only the directory carries release.json, and without it the check cannot name the release it compared against.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_strchive_reference(
    strchive_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built STRchive catalogue (`$JUST_DNA_STRCHIVE_CACHE`), without downloading.

    `None` is nobody-asked rather than "STRchive lists no such locus", and `check-repeat-bands` says
    so in those words (`@unreachable-not-absent`). A directory is preferred over a bare
    `STRchive-loci.json` because only the directory carries `release.json`, and without it the check
    cannot name the release it compared against.
    """
    return _resolve_named_cache(
        strchive_cache,
        STRCHIVE_CACHE_VAR,
        default_strchive_cache_dir(load_dotenv_file=load_dotenv_file),
        STRCHIVE_CATALOGUE_FILENAME,
        load_dotenv_file=load_dotenv_file,
    )

default_mitomap_cache_dir

default_mitomap_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/mitomap directory — see MITOMAP_SUBDIR.

Source code in enricher/src/just_dna_enricher/locations.py
def default_mitomap_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/mitomap` directory — see `MITOMAP_SUBDIR`."""
    return _cache_dir(MITOMAP_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_mitomap_reference

resolve_mitomap_reference(
    mitomap_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built MITOMAP snapshot ($JUST_DNA_MITOMAP_CACHE), without downloading.

None is nobody-asked, not "MITOMAP publishes no such variant". It is also the parent absence the derived miss lane reports as could not run rather than as an empty increment: a miss set derived with one parent missing would be a claim about a comparison that never happened.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_mitomap_reference(
    mitomap_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built MITOMAP snapshot (`$JUST_DNA_MITOMAP_CACHE`), without downloading.

    `None` is nobody-asked, not "MITOMAP publishes no such variant". It is also the parent absence the
    derived miss lane reports as *could not run* rather than as an empty increment: a miss set derived
    with one parent missing would be a claim about a comparison that never happened.
    """
    return _resolve_parquet_cache(
        mitomap_cache,
        MITOMAP_CACHE_VAR,
        default_mitomap_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_alphagenome_avi_cache_dir

default_alphagenome_avi_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/alphagenome_avi directory — see ALPHAGENOME_AVI_SUBDIR.

Source code in enricher/src/just_dna_enricher/locations.py
def default_alphagenome_avi_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/alphagenome_avi` directory — see `ALPHAGENOME_AVI_SUBDIR`."""
    return _cache_dir(ALPHAGENOME_AVI_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_alphagenome_avi_reference

resolve_alphagenome_avi_reference(
    alphagenome_avi_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built AVI snapshot ($JUST_DNA_ALPHAGENOME_AVI_CACHE), without downloading.

None here is nobody-asked in its strongest form: this lane has no acquire stage at all, so an absent snapshot means the operator has not run alphagenome build against a file they hold — not that a fetch failed, and not that AlphaGenome scores nothing at that position (@unreachable-not-absent).

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_alphagenome_avi_reference(
    alphagenome_avi_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built AVI snapshot (`$JUST_DNA_ALPHAGENOME_AVI_CACHE`), without downloading.

    `None` here is *nobody-asked* in its strongest form: this lane has no acquire stage at all, so an
    absent snapshot means the operator has not run `alphagenome build` against a file they hold — not
    that a fetch failed, and not that AlphaGenome scores nothing at that position
    (`@unreachable-not-absent`).
    """
    return _resolve_parquet_cache(
        alphagenome_avi_cache,
        ALPHAGENOME_AVI_CACHE_VAR,
        default_alphagenome_avi_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_mitomap_miss_cache_dir

default_mitomap_miss_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/mitomap_miss directory — see MITOMAP_MISS_SUBDIR.

Source code in enricher/src/just_dna_enricher/locations.py
def default_mitomap_miss_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/mitomap_miss` directory — see `MITOMAP_MISS_SUBDIR`."""
    return _cache_dir(MITOMAP_MISS_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_mitomap_miss_reference

resolve_mitomap_miss_reference(
    mitomap_miss_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built MITOMAP-miss snapshot ($JUST_DNA_MITOMAP_MISS_CACHE), without downloading.

A resolver for a derived snapshot answers a narrower question than the others: it says the join was run and its result is on disk, never that the result is still current. Currency is mitomap_miss_build.stale_parents, which re-reads both parents' release.json and names the one that moved — a snapshot present and stale is a different answer from absent, and the drafter says which (@currency-asks-the-source-not-the-cache).

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_mitomap_miss_reference(
    mitomap_miss_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built MITOMAP-miss snapshot (`$JUST_DNA_MITOMAP_MISS_CACHE`), without downloading.

    A resolver for a derived snapshot answers a narrower question than the others: it says the join
    was run and its result is on disk, never that the result is still *current*. Currency is
    `mitomap_miss_build.stale_parents`, which re-reads both parents' `release.json` and names the one
    that moved — a snapshot present and stale is a different answer from absent, and the drafter says
    which (`@currency-asks-the-source-not-the-cache`).
    """
    return _resolve_parquet_cache(
        mitomap_miss_cache,
        MITOMAP_MISS_CACHE_VAR,
        default_mitomap_miss_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )

default_drug_labels_cache_dir

default_drug_labels_cache_dir(
    *, load_dotenv_file: bool = True
) -> Path

The <base>/drug_labels directory — see DRUG_LABELS_SUBDIR.

Source code in enricher/src/just_dna_enricher/locations.py
def default_drug_labels_cache_dir(*, load_dotenv_file: bool = True) -> Path:
    """The `<base>/drug_labels` directory — see `DRUG_LABELS_SUBDIR`."""
    return _cache_dir(DRUG_LABELS_SUBDIR, load_dotenv_file=load_dotenv_file)

resolve_drug_labels_reference

resolve_drug_labels_reference(
    drug_labels_cache: Path | None = None,
    *,
    load_dotenv_file: bool = True,
) -> Path | None

Locate a built regulator drug-label snapshot ($JUST_DNA_DRUG_LABELS_CACHE).

A parquet snapshot like ClinPGx's, and a separate cache from it on purpose: the two archives are dated from their own CREATED_*.txt and do not refresh together (DRUG_LABELS_SUBDIR). Provisioning is download.ensure_drug_labels_snapshot.

Source code in enricher/src/just_dna_enricher/locations.py
def resolve_drug_labels_reference(
    drug_labels_cache: Path | None = None, *, load_dotenv_file: bool = True
) -> Path | None:
    """Locate a built regulator drug-label snapshot (`$JUST_DNA_DRUG_LABELS_CACHE`).

    A parquet snapshot like ClinPGx's, and a **separate** cache from it on purpose: the two archives
    are dated from their own `CREATED_*.txt` and do not refresh together (`DRUG_LABELS_SUBDIR`).
    Provisioning is `download.ensure_drug_labels_snapshot`.
    """
    return _resolve_parquet_cache(
        drug_labels_cache,
        DRUG_LABELS_CACHE_VAR,
        default_drug_labels_cache_dir(load_dotenv_file=load_dotenv_file),
        load_dotenv_file=load_dotenv_file,
    )