Reference Genome Build (GRCh37 vs GRCh38)
The specific assembly of the human genome your reads were aligned to, which fixes the coordinate system every variant position, annotation, and clinical report depends on.
A reference genome build is a specific, versioned assembly of the human genome that defines the coordinate system for your data: chromosome names, contig lengths, and the base at every position. GRCh37 (released 2009 by the Genome Reference Consortium, packaged by UCSC as hg19) and GRCh38 (2013, UCSC hg38) are the two builds you will encounter. “GRCh” is Genome Reference Consortium human; the number is the major build. A variant at chr1:155,205,634 means one thing in GRCh37 and a different thing in GRCh38.
How it works
The reference is a mosaic haplotype assembled from a handful of donors, not any individual’s genome. Each build release corrects assembly errors, closes gaps, and fixes the placement of sequence that was previously misassembled or missing. GRCh38’s main changes: centromeres are modeled with alpha-satellite representations instead of multi-megabase runs of N, several hundred gaps were closed, thousands of single-base and indel errors in GRCh37 were corrected, and alternate loci (ALT contigs) were added for highly polymorphic regions like MHC, plus decoy and EBV sequence in the analysis sets. Those GRCh37 base errors are the reason GRCh38 reduces a class of recurrent false-positive SNP and indel calls: reads matching the true human sequence were being called as variants against a wrong reference base.
The practical effect on alignment and calling has been measured. Comparisons of the same sequencing data processed against both builds show improved read mapping and changed variant yield on GRCh38 1, and evaluation against independent assemblies and long-read data confirmed GRCh38 resolves regions that were misrepresented in GRCh37 2. Exome reanalysis across builds found discrepant variant calls concentrated in regions where the assembly changed, including calls in clinically reported genes 3. Clinical labs that migrated report the transition as a validation exercise with real, enumerable differences rather than a version bump 4.
Naming is a separate axis from the build. GRCh37 uses 1, 2, MT; hg19 uses chr1, chr2, chrM. Worse, hg19’s chrM is the older rCRS-incompatible Yoruba mitochondrial sequence, while GRCh37’s MT is rCRS. GRCh38 and hg38 agree on chrM. Broad’s b37 is GRCh37 with Ensembl-style names and a decoy contig. These are not interchangeable even though the autosomal coordinates match.
In your own data
Check the header, not the filename. For a BAM or CRAM:
samtools view -H sample.cram | grep '^@SQ' | head -3
Look at SN: and LN:. If chr1 has LN:249250621 you are on GRCh37/hg19. LN:248956422 is GRCh38. A good header also carries AS: and M5: checksums and a UR: pointing at the exact FASTA. For a VCF, read the ##contig=<ID=chr1,length=...> lines and the ##reference= line. If ##reference is missing, treat build as unknown until you confirm it from contig lengths.
Three mistakes we see repeatedly:
- Re-annotating a GRCh37 VCF with a GRCh38 cache in VEP or SnpEff. Nothing errors. You get consequences for whatever gene happens to occupy those coordinates in the other build.
- Treating rsIDs as build-independent. They are, but the position attached to an rsID in dbSNP is build-specific, and some rsIDs have been merged or retired. If you are cross-checking a consumer array export against dbSNP, match on rsID plus alleles, not position.
- Assuming liftover is lossless.
CrossMaporpicard LiftoverVcfwith the UCSC chain file drops on the order of a small percentage of variants and can silently flip strand, changing REF/ALT. Always keep the--REJECToutput and inspect it. When you can, realign from FASTQ instead of lifting the VCF. That costs compute and gives you a genuinely GRCh38 call set.
Our recommendation: use GRCh38 for anything new, specifically the analysis set with decoys and without ALT contigs unless your caller is ALT-aware (BWA-MEM plus bwa-postalt.js, or DRAGEN). Keep GRCh37 only to reproduce an old result. If your pipeline touches immune receptor loci, note that IG and TR gene regions are poorly represented by any single linear reference and need specialized genotyping 5.
Limitations
GRCh38 is not uniformly better everywhere. Some genes were more accurately represented in GRCh37, and a few clinically relevant regions gained new problems. Inherited retinal disease analyses show that build choice plus population database version changes which variants surface for review 6. RNA-seq is not immune: reprocessing the same Alzheimer’s brain RNA-seq against an updated reference and annotation changed both which genes were called differentially expressed and, for some, the direction of the change 7. Repetitive regions like rDNA remain absent or collapsed in both builds and need custom contigs to quantify at all 8.
Also: a build is not a clinical answer. Discrepant calls across builds in a gene you care about are a reason to talk to a clinical geneticist, not to self-interpret.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Every BAM, CRAM, VCF, and expression matrix you hold is silently tied to one build; mixing builds produces variants at plausible-looking but wrong positions, and no tool will warn you.
Related Terms
References
- Yan Guo, Yulin Dai, Hui Yu, et al.. Improvements and impacts of GRCh38 human reference on high throughput sequencing data analysis . Genomics, 2017. DOI
- Valerie A. Schneider, Tina Graves-Lindsay, Kerstin Howe, et al.. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly . Genome Research, 2017. DOI
- He Li, Moez Dawood, Michael M. Khayat, et al.. Exome variant discrepancies due to reference-genome differences . The American Journal of Human Genetics, 2021. DOI
- Lisa A Lansdon, Maxime Cadieux-Dion, John C Herriges, et al.. Clinical Validation of Genome Reference Consortium Human Build 38 in a Laboratory Utilizing Next-Generation Sequencing Technologies . Clinical Chemistry, 2022. DOI
- Mao-Jan Lin, Yu-Chun Lin, Nae-Chyun Chen, et al.. Profiling genes encoding the adaptive immune receptor repertoire with gAIRR Suite . Frontiers in Immunology, 2022. DOI
- Stefan T. Stafie, Mark Lindquist, Samuel Kusher-Lenhoff, et al.. Importance of genome reference and population datasets for annotation and prioritization of disease-causing variants in inherited retinal diseases . Ophthalmic Genetics, 2025. DOI
- Anina N. Lund, Ryan C. Thompson, Ying-Chich Wang, et al.. Updates to the reference genome alter the detection and direction of genes differentially expressed in Alzheimer’s disease . Genome Biology, 2026. DOI
- Subin S. George, Maxim Pimkin, Vikram R. Paralkar. Construction and validation of customized genomes for human and mouse ribosomal DNA mapping . Journal of Biological Chemistry, 2023. DOI