Skip to content
/genome-variants/annotation

Annotation

Annotation is the process of attaching functional, population, and clinical context to each variant call in a genome so that a list of coordinates becomes a list of interpretable statements.

Annotation is the join step of genomics: you take a variant call (chr7, 117559590, G, A) and attach everything known about that position, so the row becomes a statement about a gene, a protein consequence, a population frequency, and, sometimes, a disease. Variant calling tells you what letters you have. Annotation tells you what they might mean.

How it works

A raw VCF from a 30x whole genome has roughly 4-5 million variants relative to GRCh38, the large majority of which are common SNPs in noncoding space. Annotation attaches columns to each of those rows from several independent sources:

  • Gene models. Which transcript overlaps this position, in which exon, at which codon. Tools use RefSeq, GENCODE, or Ensembl transcript sets, and the answer changes depending on which one you pick and which transcript is designated canonical.
  • Consequence terms. Sequence Ontology labels: missense_variant, stop_gained, frameshift_variant, splice_donor_variant, synonymous_variant, intron_variant. ANNOVAR’s -operation g with refGene or ensGene produces these, as does SnpEff, which was designed around annotating and predicting effects on genes and proteins 12.
  • Population frequency. gnomAD v4 allele frequency, per-ancestry. This is the single most useful filter you have.
  • Precomputed prediction scores. CADD, REVEL, SpliceAI, AlphaMissense, packaged together in dbNSFP.
  • Clinical assertions. ClinVar star ratings and classifications, OMIM gene-disease links.
  • Regulatory context for the noncoding 98%. RegulomeDB scores variants by overlap with DNase footprints, ChIP-seq peaks, and eQTLs, which is the only tractable way to prioritize intergenic calls 3.

The standard toolchain is Ensembl VEP, ANNOVAR, or SnpEff. We use VEP as the primary annotator because of --plugin support and its handling of multiple transcripts per variant, then layer ANNOVAR’s table_annovar.pl for a wide flat TSV when we want something pandas can read directly 4. Different tools disagree on consequence calls more often than you would expect, largely because of transcript set and canonical-transcript choices 5.

In your own data

Start from the VCF your sequencing provider returns, plus the .tbi index. Check the header first: bcftools view -h your.vcf.gz | grep reference. If it says GRCh37/hg19 and your annotation databases are GRCh38, every coordinate is wrong by a few hundred to a few thousand bases and you will get plausible-looking nonsense. This is the single most common failure.

A working pipeline:

bcftools norm -m-any -f GRCh38.fa in.vcf.gz -Oz -o norm.vcf.gz
vep -i norm.vcf.gz --cache --assembly GRCh38 \
    --everything --pick_allele_gene \
    --plugin CADD,whole_genome_SNVs.tsv.gz \
    --plugin dbNSFP,dbNSFP4.5a.gz,REVEL_score,AlphaMissense_score \
    --vcf -o annotated.vcf.gz

bcftools norm -m-any is not optional. Multiallelic sites collapse several alleles into one row, and most annotators either skip them or annotate only the first ALT.

Then filter. Our first pass: keep variants with gnomAD popmax AF below 0.001, consequence in the protein-altering set, and either a ClinVar assertion or a REVEL score above ~0.7. On a personal genome that typically reduces 4.5 million rows to a few hundred. Expect 50-100 known pathogenic or likely pathogenic variants in recessive carrier genes, and roughly 1-3% chance of a reportable finding in the ACMG secondary findings gene list.

Two mistakes to avoid. First, treating a ClinVar “Pathogenic” label as final: check the review status field, because a one-star single-submitter assertion is a different object from a three-star expert panel review, and reclassifications happen. Second, reading a high CADD or REVEL score as evidence of disease. These are computational priors, not evidence about your health.

Limitations

Annotation moves the ambiguity rather than removing it. The dominant output category is the variant of uncertain significance. In a large hereditary disease testing series, VUS rates varied widely by gene and by ancestry, with underrepresented ancestry groups receiving uncertain results at higher rates because their alleles are less well sampled in reference databases 6. A VUS means the evidence is insufficient, which is the correct default position: variants should be considered uncertain until the evidence is sufficient to call them otherwise 7.

VUS rates are falling as databases grow and functional assays such as deep mutational scanning fill in the evidence gap, but they will not reach zero this decade 8. Most reclassifications move toward benign. Patients and clinicians both tend to overread these results, and the downstream anxiety is well documented 910.

Nothing in an annotated VCF is a diagnosis. If a variant in your file carries a clinical assertion in a gene with established disease association, that result needs confirmation by an accredited clinical laboratory and interpretation with a genetic counselor or clinical geneticist. Research-grade annotation is a hypothesis generator.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Computational Angle

Your VCF is 4-5 million rows of chromosome, position, ref, alt. Annotation is the join that turns those rows into gene names, consequence terms, allele frequencies, and prediction scores you can filter on.

Related Terms

References

  1. K. Wang, M. Li, H. Hakonarson. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data . Nucleic Acids Research, 2010. DOI
  2. Pablo Cingolani. Variant Annotation and Functional Prediction: SnpEff . Methods in Molecular Biology, 2012. DOI
  3. Alan P. Boyle, Eurie L. Hong, Manoj Hariharan, et al.. Annotation of functional variation in personal genomes using RegulomeDB . Genome Research, 2012. DOI
  4. Hui Yang, Kai Wang. Genomic variant annotation and prioritization with ANNOVAR and wANNOVAR . Nature Protocols, 2015. DOI
  5. Eleftherios Pilalis, Dimitrios Zisis, Christina Andrinopoulou, et al.. Genome-wide functional annotation of variants: a systematic review of state-of-the-art tools, techniques and resources . Frontiers in Pharmacology, 2025. DOI
  6. Elaine Chen, Flavia M. Facio, Kerry W. Aradhya, et al.. Rates and Classification of Variants of Uncertain Significance in Hereditary Disease Genetic Testing . JAMA Network Open, 2023. DOI
  7. Karen E. Weck. Interpretation of genomic sequencing: variants should be considered uncertain until proven guilty . Genetics in Medicine, 2018. DOI
  8. Douglas M. Fowler, Heidi L. Rehm. Will variants of uncertain significance still exist in 2030? . The American Journal of Human Genetics, 2024. DOI
  9. Lily Hoffman-Andrews. The known unknown: the challenges of genetic variants of uncertain significance in clinical practice . Journal of Law and the Biosciences, 2017. DOI
  10. Stefan Timmermans, Caroline Tietbohl, Eleni Skaperdas. Narrating uncertainty: Variants of uncertain significance (VUS) in clinical exome sequencing . BioSocieties, 2016. DOI