Phasing
Phasing is the assignment of each heterozygous allele to one of the two parental chromosome copies, turning an unordered genotype list into two haplotype sequences.
Phasing is the process of deciding, for every heterozygous site in your genome, which allele sits on the chromosome you got from your mother and which sits on the one from your father. A haplotype is the resulting sequence of alleles along one physical chromosome copy. A gene is a stretch of the genome; a haplotype is one of the two versions of that stretch you carry.
A worked example. Suppose you are heterozygous at two positions in the same gene: C/T at one and G/A at the other. Unphased, that is all you know. Phased, there are two possibilities. Either C-G on one copy and T-A on the other, or C-A on one and T-G on the other. If both the T and the A are loss-of-function alleles, the first arrangement leaves you one intact copy of the gene. The second knocks out both. Same genotypes, different biology. This is why haplotypes, not single variants, are the natural unit for candidate-gene work 1.
How it works
There are three ways to get phase, and they fail differently.
Read-based phasing links alleles that appear on the same physical DNA molecule. If a single read or read pair covers two heterozygous sites, their phase is observed directly. The limit is fragment length: with 150 bp paired-end short reads, you connect sites a few hundred bases apart, and heterozygous sites in a human genome are roughly one per 1–2 kb, so most pairs never share a read. Long reads (PacBio HiFi, ONT) at 15–100 kb span many het sites at once and extend blocks by orders of magnitude. WhatsHap solves the minimum error correction problem on the read-variant matrix and is the standard tool here 2. Hi-C adds megabase-scale links and, combined with long reads, can push phase blocks to chromosome arm scale 3.
Population-based (statistical) phasing infers haplotypes from the fact that your chromosomes are mosaics of haplotypes shared with other people. Hidden Markov models over a reference panel — Beagle’s localized haplotype clustering, later Eagle2 and SHAPEIT — assign the most likely arrangement given panel haplotypes 45. Accuracy scales with panel size: reference-based phasing against the Haplotype Reference Consortium panel (~64,000 haplotypes) substantially cut switch errors relative to earlier panels 6. Modern methods now phase biobank-scale cohorts and extend the same machinery into imputation 7.
Trio phasing uses parents. If you are heterozygous and your mother is homozygous reference, your alternate allele came from your father. A mother-father-child trio resolves most het sites directly, and adding a sibling or grandparent resolves nearly all of them along with recombination breakpoints 8. A comparison across strategies on whole human genomes found read-based, statistical, and pedigree approaches each leave characteristic gaps, and that combining them beats any single one 910.
In your own data
Look at the GT field in your VCF. 0/1 is unphased. 0|1 is phased, and the order is meaningful: allele left of the bar is haplotype 1, right is haplotype 2. Phase is only valid within a block, identified by the PS FORMAT tag (an integer, usually the position of the block’s first variant). Two variants with different PS values carry no relative phase even if both show a pipe.
A practical pipeline on a HiFi BAM:
whatshap phase --reference GRCh38.fa -o phased.vcf.gz \
--ignore-read-groups sample.vcf.gz sample.bam
whatshap stats --gtf blocks.gtf phased.vcf.gz
whatshap haplotag -o tagged.bam --reference GRCh38.fa phased.vcf.gz sample.bam
whatshap stats reports block NL50 and the number of phased het variants. With 30x HiFi, expect block N50 in the hundreds of kb to megabases. With 30x Illumina alone, expect tens of kb at best. haplotag writes an HP tag per read, which lets you view haplotype-separated pileups in IGV and is the input you want for allele-specific expression or methylation.
Common mistakes: treating haplotype 1 as maternal (block orientation is arbitrary without a trio or a reference-based step); merging VCFs across tools and losing PS tags; comparing two variants in different blocks and concluding they are in trans; and using statistically phased data for a rare compound-heterozygous question, where panel information is thinnest.
Any clinical interpretation of a compound-heterozygous finding belongs with a genetics clinician, including confirmation by targeted long-range assay.
Limitations
Statistical phasing degrades exactly where you care most: rare and singleton variants have no panel support, so switch error rates rise sharply as allele frequency falls 47. Read-based phasing is bounded by fragment length and by heterozygosity, so homozygous stretches break blocks regardless of coverage. Switch errors compound: one switch mid-block silently flips everything downstream. And ancestry matters, since panel-based accuracy is lower for individuals underrepresented in the reference 6. If phase drives a decision, confirm it with a second, orthogonal method.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
In your VCF it is the difference between GT 0/1 and 0|1, plus a PS tag naming the phase block; getting it right decides whether two variants in the same gene are on one chromosome or on both.
Related Terms
References
- Andrew G. Clark. The role of haplotypes in candidate gene studies . Genetic Epidemiology, 2004. DOI
- Marcel Martin, Murray Patterson, Shilpa Garg, et al.. WhatsHap: fast and accurate read-based phasing . 2016. DOI
- Zev N. Kronenberg, Arang Rhie, Sergey Koren, et al.. Extended haplotype-phasing of long-read de novo genome assemblies using Hi-C . Nature Communications, 2021. DOI
- Sharon R. Browning, Brian L. Browning. Haplotype phasing: existing methods and new developments . Nature Reviews Genetics, 2011. DOI
- Sharon R. Browning, Brian L. Browning. Rapid and Accurate Haplotype Phasing and Missing-Data Inference for Whole-Genome Association Studies By Use of Localized Haplotype Clustering . The American Journal of Human Genetics, 2007. DOI
- Po-Ru Loh, Petr Danecek, Pier Francesco Palamara, et al.. Reference-based phasing using the Haplotype Reference Consortium panel . Nature Genetics, 2016. DOI
- Quan Sun, Yun Li. Advances in haplotype phasing and genotype imputation . Nature Reviews Genetics, 2025. DOI
- Jared C. Roach, Gustavo Glusman, Robert Hubley, et al.. Chromosomal Haplotypes by Genetic Phasing of Human Families . The American Journal of Human Genetics, 2011. DOI
- Yongwook Choi, Agnes P. Chan, Ewen Kirkness, et al.. Comparison of phasing strategies for whole human genomes . PLOS Genetics, 2018. DOI
- Matthew W. Snyder, Andrew Adey, Jacob O. Kitzman, et al.. Haplotype-resolved genome sequencing: experimental methods and applications . Nature Reviews Genetics, 2015. DOI