Skip to content
/genome-variants/indel

Indel

An insertion or deletion of one to roughly fifty bases at a single locus, represented in VCF as a REF/ALT pair of unequal length anchored on a preceding base.

An indel is an insertion or deletion of a small number of bases at one locus, conventionally capped at 50 bp, above which the event is called a structural variant. A SNP swaps one base for another and keeps the length of the sequence fixed. An indel changes the length, which is why it behaves differently in alignment, in annotation, and in coding sequence.

How it works

Most indels are not random. They concentrate in homopolymers and short tandem repeats, where polymerase slippage during replication adds or drops a repeat unit. Small insertion and deletion polymorphisms are among the most common forms of human variation after SNVs, and the majority of them fall inside repetitive sequence rather than unique sequence.1 A survey of 179 genomes defined the class as gains or losses of up to 50 nucleotides and showed that indel density, mutational origin, and functional consequence all differ by sequence context: slippage-driven events in repeats, versus non-repeat indels that look more like replication or repair errors.2 The earliest genome-wide maps established both the scale of the class and its clustering, with hundreds of thousands of distinct indel loci identified across modest sample sets.3 Larger callsets later tied specific indels to expression and protein-coding effects.4

The mutational processes behind indels are not the processes behind substitutions. Indel rates track replication timing and local sequence composition on their own terms, separately from point mutations.5 So a region that is quiet for SNVs can be noisy for indels.

Consequence depends almost entirely on where the indel lands. In an intron or intergenic region, usually nothing measurable. In a coding exon, a length change not divisible by three shifts the reading frame from that point on, and typically produces a premature stop. That is the answer to “which is worse, substitution or deletion”: a frameshift deletion in coding sequence generally destroys more protein than a single substitution, while a 3 bp in-frame deletion removes one residue and may be tolerated. Most substitutions are missense or silent. Severity is a property of the locus and the frame, not of the category.

In your own data

Indels sit in the same VCF as your SNVs. You identify them by length: REF and ALT differ in size, with one anchor base carried over from the preceding position. A 4 bp deletion at chr1:1000 looks like REF=ACGTA ALT=A at position 999.

Three things to check, in order.

Normalize before anything else. Run bcftools norm -f GRCh38.fa -m -any -c w your.vcf.gz -Oz -o norm.vcf.gz. This left-aligns indels in repeats and splits multi-allelic records. Without it, an insertion in a poly-A tract can be written at any of several positions, all correct, none matching your annotation database. Most “my variant isn’t in gnomAD” reports are unnormalized indels.

Count them. A short-read WGS callset at 30x should give you a few hundred thousand indels against low millions of SNVs. If your indel count is far outside that, the pipeline is the problem, not you.

Look at the quality stratification. Pull indels in homopolymers ≥ 8 bp and check their genotype quality. These are where short-read callers fail most often, and where Illumina and nanopore fail in different directions. If you have GIAB truth sets available for a comparable sample, hap.py or rtg vcfeval will show you indel F1 running measurably below SNV F1 on the same file, and most of the gap sits in repeats.

We would call with DeepVariant rather than GATK HaplotypeCaller for indels on short reads, because the realignment plus classifier approach handles repeat context better in practice. If you have long reads (PacBio HiFi), use those for anything in a tandem repeat and treat short-read indel calls there as provisional.

Common mistake: comparing your VCF to someone else’s with bcftools isec and getting a large private set. Almost always normalization, representation of a complex event as one indel versus two, or a liftover between GRCh37 and GRCh38 that moved coordinates. Compare with vcfeval, which does haplotype-aware matching.

Limitations

The 50 bp boundary is a convention from read length, not biology, and events near it are called inconsistently by different tools.6 A coding frameshift in your VCF is a computational prediction, not a clinical finding: confirmation requires an accredited laboratory and interpretation by a clinical geneticist. Annotation tools disagree on which transcript is canonical, so the same indel can be labeled frameshift against one transcript and intronic against another. Check the transcript ID before you believe the consequence.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Computational Angle

Indels are the variant class where your own pipeline choices show up most: caller, aligner, and normalization change which indels exist in your VCF at all, and two files describing the same event can carry different positions and different allele strings.

Related Terms

References

  1. J. M. Mullaney, R. E. Mills, W. S. Pittard, et al.. Small insertions and deletions (INDELs) in human genomes . Human Molecular Genetics, 2010. DOI
  2. Stephen B. Montgomery, David L. Goode, Erika Kvikstad, et al.. The origin, evolution, and functional impact of short insertion–deletion variants identified in 179 human genomes . Genome Research, 2013. DOI
  3. Ryan E. Mills, Christopher T. Luttig, Christine E. Larkins, et al.. An initial map of insertion and deletion (INDEL) variation in the human genome . Genome Research, 2006. DOI
  4. Ryan E. Mills, W. Stephen Pittard, Julienne M. Mullaney, et al.. Natural genetic variation caused by small insertions and deletions in the human genome . Genome Research, 2011. DOI
  5. Amnon Koren, Paz Polak, James Nemesh, et al.. Differential Relationship of DNA Replication Timing to Different Forms of Human Mutation and Variation . The American Journal of Human Genetics, 2012. DOI
  6. Benjamin D Redelings, Ian Holmes, Gerton Lunter, et al.. Insertions and Deletions: Computational Methods, Evolutionary Dynamics, and Biological Applications . Molecular Biology and Evolution, 2024. DOI