Skip to content

just_dna_format.pgx

just_dna_format.pgx

PGx star-allele model (0.4 — see docs/CHANGELOG.md). Three definition/lookup tables that, with the per-gene ActivityPhenotypeRow binning table (binning.py), form the four-table model validated against the Aldy / Cyrius / PharmCAT stack:

  1. HaplotypeRow — junction: variant ↔ allele is many-to-many (one allele = many variants; one variant recurs across many alleles). One row per (haplotype × variant).
  2. AlleleFunctionRow — allele-unit → activity value + function category. The star-string is the canonical identity, stored verbatim (*4, *1x2, *36+*10); copy number / SV are attributes of the cis allele-unit, optional parsed conveniences (the string is truth — PharmVar has no structured SV field).
  3. DiplotypeRow — the safe canonical fallback for structural/duplication/unphased cases, keyed on a canonicalized haplotype pair.

Data-agnostic (design north star — see CLAUDE.md): the format supplies these tables; a consumer star-allele caller supplies the phased diplotype + CN/SV calls and computes the phenotype. Copy number attaches to a specific cis allele-unit, so *2x2/*4 (AS 2 → NM) ≠ *2/*4x2 (AS 1 → IM) — a consumer that multiplies by total CN gets it wrong.

HaplotypeRow

Bases: AuthoredModel

Junction row: one defining variant of a named haplotype/allele. Many rows per haplotype; a variant recurs across many haplotypes (CYP2D6 rs1065852 is core-defining in 22 alleles).

Inherits AuthoredModel (reserved-namespace guard + shared rsid validator).

AlleleFunctionRow

Bases: AuthoredModel

Allele-unit → activity value + function category. The star-string allele is the required canonical key. suballele is optional-extra (Aldy's Minor, e.g. 1.001); the core star is the identity. copy_number/sv_type/hybrid_orientation are optional parsed conveniences of the cis allele-unit — the star-string remains truth.

Inherits AuthoredModel (reserved-namespace guard).

DiplotypeRow

Bases: AuthoredModel

Canonical fallback: a diplotype (haplotype pair) → phenotype. The pair is canonicalized (haplotype_a <= haplotype_b) so a lookup is order-independent; multiple rows per pair are allowed (a pleiotropic diplotype affecting several traits).

No requires_callable here, and the absence is the decision (RM70, 0.7). The column landed on HaplotypeRow and PharmVariantRow because each of those rows names a locus, so a callability claim on one is about a position the row states — exactly what the column means on VariantRow. A diplotype names a star-allele pair, not a locus, so the same column here could only mean "the variants defining these two haplotypes were callable", which is a fact about haplotypes.csv rows restated one table over and free to drift the moment a definition is edited. One concept, one home. extra="forbid" on AuthoredModel is what enforces it, and a test pins the refusal.

Inherits AuthoredModel (reserved-namespace guard + shared direction/clin_sig/ evidence_level/trait_efo_id validators).

PharmVariantRow

Bases: AuthoredModel

Single-variant PharmGKB drug-response annotation (item 9) — pharm_variants.csv.

A distinct rowtype rather than columns on VariantRow, so the SNP core stays free of the drug-response domain: a module includes this table only when it carries drug annotations (one CSV = one concern; no empty variants.csv). Diplotype-keyed drug response instead rides on DiplotypeRow's optional drug columns. A row maps a variant → a drug → a response + a PharmGKB evidence level (1A…4) — a different axis from a risk weight (why it is not a VariantRow).

genotype is part of the identity, not decoration (0.5). A PharmGKB clinical annotation is published per genotype: the summary row names the variant and drug, and a child table gives one annotation per call, and the large majority carry exactly three. They are not variations on one finding but distinct, sometimes opposed, ones: for rs4149056/simvastatin, CC and CT read "decreased response" while TT reads "increased". Modelling only (variant, drug) collapsed them, and the compiler's duplicate-row check rejected the real data outright — the axis is therefore in the dedup key (variant_key, drug, genotype), mirroring the SNP core's (variant, genotype) rule. It is not derivable: nothing else on this row distinguishes the calls but free text.

The grammar is the shared one on AuthoredModel, so a genotype means here exactly what it means on a VariantRow. Two shapes upstream deliberately do not land in this column: a haplotype-keyed annotation (*1, *1xN) belongs on DiplotypeRow, which already models a haplotype pair, and a symbolic allele (C/del, del/del) is RM5 and is skipped rather than coerced. PharmGKB writes a diploid call concatenated (CC); the canonical form here is sorted and slash-separated (C/C), since CC would otherwise read as a single two-base allele.

Inherits AuthoredModel (reserved-namespace guard + shared rsid/evidence_level/ trait_efo_id/genotype validators).

validate_haplotype_name

validate_haplotype_name(value: str, field_name: str) -> str

One rule for a haplotype name, shared by all three PGx tables so they cannot disagree.

Source code in schema/src/just_dna_format/pgx.py
def validate_haplotype_name(value: str, field_name: str) -> str:
    """One rule for a haplotype name, shared by all three PGx tables so they cannot disagree."""
    if not HAPLOTYPE_NAME_PATTERN.match(value or ""):
        raise ValueError(
            f"{field_name} must be a non-empty haplotype name without whitespace (e.g. *4, e4, "
            f"\u03b54), got: {value!r}"
        )
    return value