Imputation
Statistical inference of genotypes you did not directly measure, using haplotype patterns shared between your sample and a large reference panel.
Imputation is the statistical inference of genotypes at positions you never measured, done by matching the stretches of your chromosome you did measure against a reference panel of sequenced haplotypes and copying over what those matching haplotypes carry.
A genotype is what you have at one position on the two copies of a chromosome: the pair of alleles, written 0/0 (two reference), 0/1 (one of each), 1/1 (two alternate) 1. A genotyping array measures maybe 650,000 of those positions directly. A human genome contains tens of millions of common and rare variable positions 2. Imputation fills the gap between the two.
How it works
Chromosomes are inherited in blocks. Recombination breaks them up slowly enough that if you share a few hundred kilobases of measured markers with someone in a reference panel, you likely share the whole segment, including the variants nobody typed on your chip. The standard implementation is a hidden Markov model where the hidden state is “which reference haplotype am I copying right now,” transitions correspond to recombination, and emissions allow for mutation and genotyping error. The Li and Stephens framework underlies IMPUTE2, minimac, and Beagle alike 3.
The workflow has two stages. First phasing: resolve your unordered genotype calls into two ordered haplotypes (SHAPEIT5 or Beagle 5.4 are the current choices). Then imputation against a panel. Scaling matters here. IMPUTE2 introduced selecting a custom subset of reference haplotypes per sample, which made panels of thousands of genomes tractable 3. minimac2 got further speedups by collapsing reference haplotypes that are identical over a given region into single states, which cuts redundant HMM computation dramatically on large panels 4.
Panels grew from HapMap to 1000 Genomes to HRC to TOPMed (roughly 97,000 sequenced samples, ~300 million variants). Bigger panel, better rare-variant coverage, more ancestry representation.
In your own data
Open your imputed VCF and look at the INFO field. You will see something like:
chr9 22125504 rs1333049 C G . PASS AF=0.47;MAF=0.47;R2=0.98;IMPUTED GT:DS:GP 0|1:0.99:0.01,0.97,0.02
Three things to read:
IMPUTEDvsTYPED. TOPMed and Michigan Imputation Server tag which sites were on your input chip. Everything else was inferred.R2(orINFOfrom IMPUTE2, orDR2from Beagle). This is the estimated squared correlation between the imputed dosage and the true genotype. Standard filter is R² ≥ 0.3 for association testing, R² ≥ 0.8 if you care about a single variant. Below 0.3 the variant carries almost no information.DSandGP. DS is the alternate allele dosage, 0 to 2, a continuous quantity. GP is the posterior probability over the three genotypes. The hard callGTis just the argmax of GP and throws away the uncertainty. If you are computing a polygenic score, use DS, not GT.
Common mistakes, in the order we see them:
- Strand and reference mismatch. Array data is often on the forward strand of the probe design, not the reference. A/T and C/G sites are ambiguous and get flipped silently. Run the panel’s checking script (the HRC/1000G
HRC-1000G-check-bim.plis still the workhorse) before uploading anything. - Wrong build. Arrays ship in GRCh37 more often than you would expect. Lift over with
CrossMaporPicard LiftoverVcfand expect to lose a small percentage of sites. - Filtering on R² after computing an allele frequency across a filtered set. R² is frequency-dependent: rare variants have low R² almost by construction, so a flat threshold removes rare variants preferentially and biases anything frequency-stratified.
- Treating an imputed genotype at a clinically actionable site as a result. It is not. Confirm with targeted sequencing, and interpret with a clinician.
Limitations
Accuracy tracks how well the panel represents your ancestry. Imputation is most accurate in European-ancestry samples and degrades in African-ancestry samples, where shorter haplotype blocks and higher diversity mean each measured marker constrains less of the surrounding region 5. Admixed individuals do better with a diverse panel than with a single-population one.
Rare variants impute poorly regardless of panel. Below roughly 0.5% minor allele frequency, R² drops fast because few reference haplotypes carry the allele. Structural variants, repeat expansions, and anything in segmental duplications or the MHC are effectively unimputable without a panel built for the region.
Method choice matters less than panel choice. Comparisons of imputation algorithms find that differences between well-implemented HMM methods are small next to differences in reference panel size and match 6. Methods that skip the reference panel entirely and impute from linkage disequilibrium within the study sample exist, mainly for organisms without one 7.
The simple way out, if you can afford it: sequence. 30x whole-genome sequencing observes the variants directly, and imputation becomes a step you never run.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
If your data came from an array or low-coverage sequencing, most of the variants in your VCF were inferred, not observed. Knowing which is which, and reading the per-variant quality scores that tell you, decides whether a result is worth acting on.
Related Terms
References
- Martin Mahner, Michael Kary. What Exactly Are Genomes, Genotypes and Phenotypes? And What About Phenomes? . Journal of Theoretical Biology, 1997. DOI
- Patrick O. Brown, Leland Hartwell. Genomics and human disease—variations on variation . Nature Genetics, 1998. DOI
- Bryan Howie, Jonathan Marchini, Matthew Stephens. Genotype Imputation with Thousands of Genomes . G3 Genes|Genomes|Genetics, 2011. DOI
- Christian Fuchsberger, Gonçalo R. Abecasis, David A. Hinds. minimac2: faster genotype imputation . Bioinformatics, 2014. DOI
- Lucy Huang, Yun Li, Andrew B. Singleton, et al.. Genotype-Imputation Accuracy across Worldwide Human Populations . The American Journal of Human Genetics, 2009. DOI
- Yu-Fang Pei, Jian Li, Lei Zhang, et al.. Analyses and Comparison of Accuracy of Different Genotype Imputation Methods . PLoS ONE, 2008. DOI
- Daniel Money, Kyle Gardner, Zoë Migicovsky, et al.. LinkImpute: Fast and Accurate Genotype Imputation for Nonmodel Organisms . G3 Genes|Genomes|Genetics, 2015. DOI