diplotypes.csv¶
One row is one haplotype pair → one phenotype, and the pair is canonicalised (haplotype_a <=
haplotype_b) so a lookup is order-independent. Several rows per pair are legal — a pleiotropic
diplotype affecting several traits is several rows, not one row with a list.
There is deliberately no requires_callable column, and the absence is the decision. A diplotype
names a pair, not a locus, so the column could only mean "the variants defining these two
haplotypes were callable" — a fact about haplotypes.csv rows, restated one table over, free to drift
the moment a definition is edited. One concept, one home. extra="forbid" is what enforces it, so
adding the column by hand is a compile error rather than a silent second source of truth.
Identity¶
| Row model | DiplotypeRow (just_dna_format.pgx) |
| Becomes | diplotypes.parquet — lead parquet: diplotypes.parquet |
| Authored or derived | authored — a person writes it (a drafter may append rows) |
| Draftable | yes — draft can append rows |
| Drafted by | no drafting provider targets this table — draft writes no row of it |
| Natural key | gene, haplotype_a, haplotype_b, trait_efo_id, drug, clinical_context |
| Fact signature | no |
| In the attestation binding | yes — manifest.inputs[] |
Columns¶
| Column | Type | Required | Values | Meaning |
|---|---|---|---|---|
gene |
str |
required | Gene symbol, e.g. CYP2D6 | |
haplotype_a |
str |
required | First haplotype of the pair (canonicalized a <= b) | |
haplotype_b |
str |
required | Second haplotype of the pair | |
trait_efo_id |
str | None |
optional | EFO/MONDO/OBA/HP trait ontology id(s) | |
direction |
str | None |
optional | one of: contested, neutral, protective, risk, unknown |
Effect direction |
clin_sig |
str | None |
optional | one of: affects, association, benign, conflicting, drug_response, likely_benign, likely_pathogenic, not_provided, other, pathogenic, protective, risk_factor, uncertain_significance |
Clinical significance |
phenotype |
str | None |
optional | Metabolizer phenotype, e.g. PM/NM | |
conclusion |
str |
required | Human-readable interpretation for this diplotype | |
drug |
str | None |
optional | Drug the response is about, e.g. codeine | |
response |
str | None |
optional | Drug response / phenotype, free-form | |
evidence_level |
str | None |
optional | one of: 1A, 1B, 2A, 2B, 3, 4 |
PharmGKB clinical-annotation evidence level (1A..4) |
recommendation_strength |
str | None |
optional | one of: moderate, no_recommendation, optional, strong |
How firmly the guideline recommends the action: strong|moderate|optional|no_recommendation (CPIC's classification). Empty when the source did not classify. |
clinical_context |
str | None |
optional | Clinical setting this row applies to — indication, age band, prior treatment, dose (e.g. 'CVI ACS PCI', 'pediatrics', 'PHT naive'). Empty means the row is unscoped. Free text: guideline bodies scope differently and a closed set would reject the next one. Part of the row key, so contexts that disagree coexist and the consumer picks. |
Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.