Skip to content

diplotypes.csv

One row is one haplotype pair → one phenotype, and the pair is canonicalised (haplotype_a <= haplotype_b) so a lookup is order-independent. Several rows per pair are legal — a pleiotropic diplotype affecting several traits is several rows, not one row with a list.

There is deliberately no requires_callable column, and the absence is the decision. A diplotype names a pair, not a locus, so the column could only mean "the variants defining these two haplotypes were callable" — a fact about haplotypes.csv rows, restated one table over, free to drift the moment a definition is edited. One concept, one home. extra="forbid" is what enforces it, so adding the column by hand is a compile error rather than a silent second source of truth.

Identity

Row model DiplotypeRow (just_dna_format.pgx)
Becomes diplotypes.parquet — lead parquet: diplotypes.parquet
Authored or derived authored — a person writes it (a drafter may append rows)
Draftable yes — draft can append rows
Drafted by no drafting provider targets this table — draft writes no row of it
Natural key gene, haplotype_a, haplotype_b, trait_efo_id, drug, clinical_context
Fact signature no
In the attestation binding yes — manifest.inputs[]

Columns

Column Type Required Values Meaning
gene str required Gene symbol, e.g. CYP2D6
haplotype_a str required First haplotype of the pair (canonicalized a <= b)
haplotype_b str required Second haplotype of the pair
trait_efo_id str | None optional EFO/MONDO/OBA/HP trait ontology id(s)
direction str | None optional one of: contested, neutral, protective, risk, unknown Effect direction
clin_sig str | None optional one of: affects, association, benign, conflicting, drug_response, likely_benign, likely_pathogenic, not_provided, other, pathogenic, protective, risk_factor, uncertain_significance Clinical significance
phenotype str | None optional Metabolizer phenotype, e.g. PM/NM
conclusion str required Human-readable interpretation for this diplotype
drug str | None optional Drug the response is about, e.g. codeine
response str | None optional Drug response / phenotype, free-form
evidence_level str | None optional one of: 1A, 1B, 2A, 2B, 3, 4 PharmGKB clinical-annotation evidence level (1A..4)
recommendation_strength str | None optional one of: moderate, no_recommendation, optional, strong How firmly the guideline recommends the action: strong|moderate|optional|no_recommendation (CPIC's classification). Empty when the source did not classify.
clinical_context str | None optional Clinical setting this row applies to — indication, age band, prior treatment, dose (e.g. 'CVI ACS PCI', 'pediatrics', 'PHT naive'). Empty means the row is unscoped. Free text: guideline bodies scope differently and a closed set would reject the next one. Part of the row key, so contexts that disagree coexist and the consumer picks.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.