Skip to content

just_dna_enricher.clin_sig

just_dna_enricher.clin_sig

One clinical-significance normalizer, shared by every source that reports one (RM134 section A).

Why this is a module and not a helper inside clinvar_build. The three-way concordance check this release builds toward reports agreement between two authorities by comparing their normalized calls. A second hand-written map would make any drift between the two maps read as a disagreement between ClinVar and PubMind, which is a finding about our own code wearing the costume of a finding about the field. So there is one map, one severity order and one function, and every snapshot builder calls it.

It is deliberately dependency-free. Only VALID_CLIN_SIG is imported, so a runtime pass can read it without the [dev] polars extra the builders need — which is the whole argument for a shared home rather than one builder importing another.

Two defects were fixed on the way out of clinvar_build, and both were invisible from ClinVar's side alone. The map's keys are underscored because that is how ClinVar spells CLNSIG; PubMind spells the same concepts with spaces. So Uncertain significance and Conflicting both fell through to other, against uncertain_significance and conflicting for ClinVar's own wording of the same two concepts — a manufactured disagreement, on the largest disagreeing class in the measured corpus join. The repair is a whitespace→underscore step in the tokenizer, which is an identity on every key below, plus the bare conflicting key PubMind's single-word token needs.

A composite token is still resolved by severity, not by a special case. PubMind's Benign/Likely benign folds to likely_benign, exactly as ClinVar's Benign/Likely_benign does, because that is what one normalizer means: the same concept gets the same answer whoever spelled it.

normalize_clin_sig

normalize_clin_sig(raw: str | None) -> str

Fold a raw clinical-significance value into a single VALID_CLIN_SIG member.

Absent, empty or whitespace-only is not_provided — the source states no classification, which is a member of the vocabulary rather than an unknown to withhold. A token with no module axis becomes other, and the two are different answers: a source that said nothing has not disagreed with anybody, while one that said something we do not model has.

A value that splits into no tokens at all ("|||") reaches not_provided for the same reason. The extracted original answered other there, which read a blank cell as an unmodelled wording; no ClinVar CLNSIG takes that shape, so nothing built to date moves.

Source code in enricher/src/just_dna_enricher/clin_sig.py
def normalize_clin_sig(raw: str | None) -> str:
    """Fold a raw clinical-significance value into a single `VALID_CLIN_SIG` member.

    Absent, empty **or whitespace-only** is `not_provided` — the source states no classification,
    which is a member of the vocabulary rather than an unknown to withhold. A token with no module
    axis becomes `other`, and the two are different answers: a source that said nothing has not
    disagreed with anybody, while one that said something we do not model has.

    A value that splits into no tokens at all (`"|||"`) reaches `not_provided` for the same reason.
    The extracted original answered `other` there, which read a blank cell as an unmodelled wording;
    no ClinVar `CLNSIG` takes that shape, so nothing built to date moves.
    """
    if not (raw or "").strip():
        return "not_provided"
    tokens = {_tokenize(tok) for tok in CLIN_SIG_SPLIT.split(raw) if tok.strip()}
    if not tokens:
        return "not_provided"
    mapped = {CLIN_SIG_MAP.get(tok, "other") for tok in tokens}
    for sig in CLIN_SIG_SEVERITY:
        if sig in mapped:
            return sig
    return "other"