Genotype
A genotype is the set of alleles an individual carries at a specific locus, written as the two (or more) sequences observed on the homologous chromosomes at that position.
A genotype is what you carry at one position in your genome: the pair of alleles, one from each parent, observed at a defined locus. The word has a broader sense too (Johannsen coined it in 1909 for the hereditary constitution of an organism, as opposed to its appearance)1, and philosophers of biology have argued about whether the term refers to the DNA itself or to a description of it2. When you are holding a VCF, the narrow sense is the one that matters.
Five concrete examples, in the notation you will encounter:
rs1801133C/T — one reference C allele and one alternate T allele at the MTHFR 677 position. Heterozygous.rs4988235T/T — homozygous for the allele associated with lactase persistence in Europeans.- ABO
O/O— two non-functional alleles at the ABO locus, which is why the common question “what is my genotype if I am O+” has a partial answer: blood group O phenotype almost always means genotype OO, and the+refers to a separate locus, RHD, where you carry at least one functional copy. - HTT
(CAG)17/(CAG)19— a short tandem repeat genotype, given as repeat counts rather than bases3. chr16:29.6–30.2Mb CN=3— a copy number genotype, three copies of a segment instead of the usual two4.
“AA genotype” is just shorthand for homozygous for whichever allele the source labeled A. It is meaningless without knowing the rsID and which strand and allele coding were used. GWAS summary tables routinely report A1/A2 in a different order from dbSNP, and some older arrays reported on the minus strand. Check before you conclude anything.
How it works
Genotyping is a measurement, and measurements have technologies with different error structures. Arrays interrogate a fixed panel of a few hundred thousand to a couple million pre-chosen sites by hybridization and produce intensity clusters that are called into AA/AB/BB5. Short-read sequencing aligns reads to a reference and calls a genotype per position from the pileup, so it sees sites nobody chose in advance. Long reads resolve repeats and structural variants that short reads collapse.
A caller like GATK HaplotypeCaller or DeepVariant emits a per-site likelihood for each of the three possible diploid states and writes the most likely one. The GT field is the answer, GQ is the confidence in phred scale, DP is depth, and AD is the per-allele read counts. All four matter. A 0/1 with AD=2,1 is noise wearing a genotype’s clothes.
In your own data
Your 30x WGS deliverable will include a BAM or CRAM and a VCF or gVCF. Start here:
bcftools view -i 'QUAL>30 && FMT/DP>10 && FMT/GQ>20' sample.vcf.gz -Oz -o filtered.vcf.gz
bcftools query -f '%CHROM\t%POS\t%ID\t%REF\t%ALT[\t%GT\t%AD\t%GQ]\n' \
-r chr1:11796321 filtered.vcf.gz
Things to check, in order:
- Reference build. A coordinate in GRCh37 is a different position in GRCh38. Read the
##referenceheader line and the contig names (chr1vs1) before you look up a single rsID. - Allele balance on heterozygous calls. For a true het at 30x, expect the alternate allele fraction near 0.5. Genome-wide, plot
AD[1]/(AD[0]+AD[1])for all0/1calls. A second mode near 0.2 usually means mapping artifacts or contamination. ./.versus0/0. A gVCF distinguishes “confidently reference” from “no data”. A plain VCF often does not. Absence of a variant line is not evidence of the reference allele unless you have the gVCF reference blocks to prove coverage.- Phase.
0|1means phased,0/1means unphased. If you want to know whether two variants in the same gene are on the same chromosome or opposite ones, you need phasing, from long reads, trio data, or statistical phasing against a reference panel that exploits haplotype block structure6.
Common mistakes: trusting an unfiltered VCF, comparing a 23andMe array call to a sequencing call without checking strand, and reading an STR locus from short reads at all.
Limitations
Genotype is not phenotype. Most common variants shift risk by a few percent, and even well-powered studies of coronary artery disease turn up loci with odds ratios in the 1.1 range7. The mapping from genotype to trait runs through networks of interacting genes, epistasis, and environment, which is why model organism work keeps finding that a single variant’s effect depends on the rest of the genome it sits in89. Penetrance varies. Genome-wide, expect roughly 4–5 million variant sites in one human genome, most of them common and benign10.
Some regions are hard: segmental duplications, the MHC, repeat expansions, and any locus where short reads multimap. If a clinically relevant genotype matters to you, it needs orthogonal confirmation in a clinical laboratory and interpretation by a clinician. A research-grade VCF line is a hypothesis about your DNA, not a result to act on.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
In your own data a genotype is a field in a VCF: GT=0/1, plus the read depth, allele balance, and quality score that tell you whether to believe it.
Related Terms
References
- W Johannsen. The genotype conception of heredity . International Journal of Epidemiology, 2014. DOI
- Martin Mahner, Michael Kary. What Exactly Are Genomes, Genotypes and Phenotypes? And What About Phenomes? . Journal of Theoretical Biology, 1997. DOI
- Hope A. Tanudisastro, Ira W. Deveson, Harriet Dashnow, et al.. Sequencing and characterizing short tandem repeats in the human genome . Nature Reviews Genetics, 2024. DOI
- Feng Zhang, Wenli Gu, Matthew E. Hurles, et al.. Copy Number Variation in Human Health, Disease, and Evolution . Annual Review of Genomics and Human Genetics, 2009. DOI
- Jiannis Ragoussis. Genotyping Technologies for Genetic Research . Annual Review of Genomics and Human Genetics, 2009. DOI
- Mark J. Daly, John D. Rioux, Stephen F. Schaffner, et al.. High-resolution haplotype structure in the human genome . Nature Genetics, 2001. DOI
- The Coronary Artery Disease (C4D) Genetics Consortium. A genome-wide association study in Europeans and South Asians identifies five new loci for coronary artery disease . Nature Genetics, 2011. DOI
- Michael Costanzo, Elena Kuzmin, Jolanda van Leeuwen, et al.. Global Genetic Networks and the Genotype-to-Phenotype Relationship . Cell, 2019. DOI
- Ben Lehner. Genotype to phenotype: lessons from model organisms for human genetics . Nature Reviews Genetics, 2013. DOI
- Patrick O. Brown, Leland Hartwell. Genomics and human disease—variations on variation . Nature Genetics, 1998. DOI