Skip to content
/omics-layers/methylation-beta-value

Methylation Beta Value

The beta value is the fraction of DNA molecules methylated at a single CpG site in a sample, computed as methylated signal divided by total signal, bounded between 0 and 1.

A methylation beta value is the estimated proportion of DNA molecules in a sample that carry a methyl group at one specific CpG site: β = M / (M + U + α), where M is methylated signal, U is unmethylated signal, and α is a small offset (usually 100 on Illumina arrays) that stabilizes the ratio when both signals are near background. It runs from 0 (nothing methylated) to 1 (everything methylated).

How it works

On an array, M and U are fluorescence intensities from two probe designs. Infinium I chemistry uses two beads in the same color channel; Infinium II uses one bead with the methylated and unmethylated states read in green and red. Type II probes have a compressed dynamic range and shifted mode, which is why almost every pipeline applies some form of probe-type correction. BMIQ and SWAN are the common choices, and the data-driven comparison of 450K preprocessing options is still the best starting point for picking one 1.

On bisulfite sequencing, the beta value is simpler and more direct: methylated reads over total reads at a cytosine. No probe chemistry, no offset, but heavy dependence on coverage. At 10x, your beta is quantized to multiples of 0.1 and carries binomial noise of roughly ±0.15 at β = 0.5.

The distribution of beta across the genome is bimodal. Most CpGs sit near 0 or near 1, with promoter CpG islands mostly unmethylated and gene bodies and repeats mostly methylated. The interesting sites are the intermediate ones, and the tissue-specific variation concentrates in enhancers and shores rather than island cores 2. The EPIC array was designed around that fact, adding roughly 350,000 enhancer-region probes to the 450K content 3.

Because beta is a bounded proportion, its variance is heteroscedastic: tightest near 0 and 1, largest near 0.5. For linear models, transform to M = log2(β / (1 − β)) and report effects back on the beta scale, because a beta difference is interpretable (a 5% shift in methylated molecules) and an M-value difference is not.

In your own data

Array output arrives as IDAT files, two per sample (_Grn.idat, _Red.idat). Read them with minfi::read.metharray.exp() in R or methylprep in Python. From there:

  • Check detection p-values first. minfi::detectionP(), drop samples with mean detection p above 0.01, drop probes failing in more than 1% of samples.
  • Drop cross-reactive and SNP-overlapping probes. There is a published exclusion list for EPIC; SNPs under the probe or at the extension base produce trimodal beta distributions at roughly 0, 0.5, 1 that look like clean biology and are genotype 4.
  • Normalize: preprocessFunnorm() if you expect global differences between groups (tumor vs normal), preprocessQuantile() for homogeneous tissue.
  • Estimate cell composition. For whole blood, minfi::estimateCellCounts2() gives granulocyte, CD4T, CD8T, NK, B, monocyte fractions. Put them in your model. Almost every large blood-methylation association is partly a shift in cell proportions.

The output matrix is CpGs by samples, dimensions around 850,000 × n for EPIC v1 or 935,000 for v2. Row names are cg identifiers like cg16867657 (that one is in the ELOVL2 promoter and is one of the strongest single-CpG correlates of chronological age).

Common mistakes we see. Comparing beta values across platforms without recalibration: an EPIC beta and a WGBS beta at the same CpG can differ by 0.05 to 0.1 systematically. Treating a single CpG as a readout when the array measures about 3% of the genome’s 28 million CpGs. Underpowering: the EPIC array guidance work shows that detecting a 1% methylation difference at genome-wide significance needs hundreds of samples, and most single-person time series cannot support that kind of claim 5. Within one person sampled repeatedly, technical replicate correlation of r > 0.99 across all probes is normal, so a genome-wide r of 0.99 between two of your own timepoints means nothing about stability at the sites you care about.

Limitations

A beta value is a cell-population average. In a tumor biopsy, β = 0.4 might be 40% of cells fully methylated or every cell at partial methylation, and purity correction changes the answer substantially 6. Tools like MethylMix explicitly model this by comparing tumor beta distributions against normal tissue to find genes where methylation is driving expression 7.

Clinical uses do exist and are narrow. Episignature testing, matching a genome-wide beta pattern against reference signatures for specific Mendelian disorders, resolved a meaningful fraction of previously unsolved cases 8. Methylation markers in cervical screening have good sensitivity but the specificity and triage questions remain live 9. Both are clinician-ordered tests with validated reference panels, not something to read off your own array. Foundation models trained on millions of methylomes now impute unmeasured CpGs and predict phenotypes from beta matrices, which is promising for research use and not yet a diagnostic 10.

If a methylation result makes you think you have a specific condition, that is a conversation with a clinical geneticist, not a conclusion from your own R session.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Computational Angle

In your own files it is a single float per CpG per sample, and nearly every mistake people make with it comes from forgetting that the number is a population average across millions of cells of mixed type.

Related Terms

References

  1. Ruth Pidsley, Chloe C Y Wong, Manuela Volta, et al.. A data-driven approach to preprocessing Illumina 450K methylation array data . BMC Genomics, 2013. DOI
  2. Katherine E. Varley, Jason Gertz, Kevin M. Bowling, et al.. Dynamic DNA methylation across diverse human cell lines and tissues . Genome Research, 2013. DOI
  3. Sebastian Moran, Carles Arribas, Manel Esteller. Validation of a DNA Methylation Microarray for 850,000 CpG Sites of the Human Genome Enriched in Enhancer Sequences . Epigenomics, 2015. DOI
  4. Sergio Villicaña, Jordana T. Bell. Genetic impacts on DNA methylation: research findings and future perspectives . Genome Biology, 2021. DOI
  5. Georgina Mansell, Tyler J. Gorrie-Stone, Yanchun Bao, et al.. Guidance for DNA methylation studies: statistical insights from the Illumina EPIC array . BMC Genomics, 2019. DOI
  6. Johan Staaf, Mattias Aine. Tumor purity adjusted beta values improve biological interpretability of high-dimensional DNA methylation data . PLOS ONE, 2022. DOI
  7. Olivier Gevaert. MethylMix: an R package for identifying DNA methylation-driven genes . Bioinformatics, 2015. DOI
  8. Erfan Aref-Eshghi, Eric G. Bend, Samantha Colaiacovo, et al.. Diagnostic Utility of Genome-wide DNA Methylation Testing in Genetically Unsolved Individuals with Suspected Hereditary Conditions . The American Journal of Human Genetics, 2019. DOI
  9. Attila T. Lorincz. Virtues and Weaknesses of DNA Methylation as a Test for Cervical Cancer Prevention . Acta Cytologica, 2016. DOI
  10. Lucas Paulo de Lima Camillo, Raghav Sehgal, Jenel Armstrong, et al.. CpGPT: a Foundation Model for DNA Methylation . 2024. DOI