Skip to content

haplotypes.csv

One row is one defining variant of one named haplotype — a junction table, not a description of the haplotype. Many rows per haplotype, and a variant recurs across many: CYP2D6's rs1065852 is core-defining in 22 alleles. Nobody should expect one row to describe *4.

A haplotype carries the reference allele at every site of its gene that it does not list. This is the PharmVar/CPIC reading, and the compiler's phase-ambiguity check already computes with it: an unlisted site and a row whose allele equals its own ref mean the same thing. So a sparse table is normal. A CPIC-drafted gene lists only each allele's own defining variants, and two of the four reference examples with this table do exactly that. The rule is closed-world per module: "reference" is about the sites some haplotype of this gene in this module lists, never about sites the module does not mention. There is no spelling for "unknown at this site". An author who means it has no row to write, so a haplotype whose definition is incomplete should not be defined here at all.

A star allele can be used without being defined, and that is legal. diplotypes.csv and allele_function.csv may name an allele this table never defines; the compiler warns only when haplotypes.csv is present at all, and *1 is exempt. That exemption is the literal name *1, for every gene, and it only silences that one warning. Nothing infers a definition for an undefined *1: the phase check skips any diplotype naming a haplotype this table does not define, *1 included. So "*1 is the reference at every site" is the CYP naming convention, not a rule of the format. A gene whose reference allele has another name (NAT2's is *4, and blood groups use names like wt) defines it the way hfe_compound_het defines wt, with a row per site whose allele is the ref. That makes it a haplotype the checks can see, and it clears the used-but-not-defined warning honestly.

requires_callable lives here rather than on diplotypes.csv, because a row here names a locus, so a callability claim is about a position the row actually states.

Identity

Row model HaplotypeRow (just_dna_format.pgx)
Becomes haplotypes.parquet — lead parquet: haplotypes.parquet
Authored or derived authored — a person writes it (a drafter may append rows)
Draftable yes — draft can append rows
Drafted by no drafting provider targets this table — draft writes no row of it
Natural key haplotype_name, variant_key, allele
Fact signature no
In the attestation binding yes — manifest.inputs[]
Requires at least one of rsid or chrom+start

Columns

Column Type Required Values Meaning
haplotype_name str required Named haplotype/allele, e.g. *4 or e4
rsid str | None optional dbSNP id of the defining variant
chrom str | None optional Chromosome (position-only variants)
start int | None optional 1-based position, VCF POS convention (position-only) — CPIC/PharmVar publish this convention and it is stored as-is; do not subtract one
ref str | None optional Reference allele (position-only)
allele str required The defining (variant) allele on this haplotype — bases, or a symbolic/structural allele carrying its length (e.g. for a whole-gene deletion)
gene str | None optional Gene symbol, e.g. CYP2D6
requires_callable bool | None optional True when a consumer must prove this position was callable before reading the absence of the defining allele as reference — i.e. before assigning the reference haplotype at this locus. False records the opposite claim, which is the one CPIC's star-allele system makes: an uncalled position is taken as reference. Empty says nothing either way, and is not False. Per locus rather than per module: one gene can hold a common allele defined by a single SNP beside one defined partly by a structural event, and the two do not have the same callability requirement.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.