Skip to content
/statistical-modeling/multiple-testing-correction

Multiple Testing Correction

A family of procedures that adjust significance thresholds or p-values so that testing thousands or millions of hypotheses at once does not produce a flood of false positives.

Multiple testing correction is the adjustment you apply when a single analysis produces many p-values, so that the error rate you control is a property of the whole set of tests rather than of one test in isolation. Test 20,000 genes at α = 0.05 with nothing true, and you expect about 1,000 “significant” results.

How it works

Two error rates matter, and they answer different questions.

The family-wise error rate (FWER) is the probability of making even one false positive across the entire set. Bonferroni controls it by testing each hypothesis at α/m for m tests. It is valid under any dependence structure, which is why it survives: no assumptions, one line of arithmetic. Šidák (1 − (1 − α)^(1/m)) is marginally less conservative and assumes independence. Holm’s step-down procedure is uniformly more powerful than Bonferroni and still controls FWER under arbitrary dependence, so there is no good reason to use plain Bonferroni over Holm when you have the p-value vector in hand.

The false discovery rate (FDR) is the expected proportion of your declared discoveries that are false. Benjamini-Hochberg sorts p-values ascending, finds the largest k with p(k) ≤ (k/m)·α, and rejects the first k. If you call 100 genes at FDR 0.05, you accept that roughly 5 are wrong. That is a reasonable bargain for a screen whose output is a list of candidates for follow-up, and a bad bargain when a single false claim is expensive 1.

The choice follows the downstream use. One confirmatory test, or a result you will act on directly, argues for FWER. A genome-wide scan feeding a ranked shortlist argues for FDR. Goeman and Solari give a clear treatment of why these are different questions and where the intermediate options (closed testing, selective inference on a chosen subset) sit 2.

The complication in molecular data is that tests are not independent. Variants in linkage disequilibrium and co-expressed genes carry correlated signal, so m overstates the number of independent tests and Bonferroni becomes conservative. The standard fixes estimate an effective number of tests, M_eff, from the eigenvalues of the correlation matrix and apply α/M_eff 3. Refinements to that eigenvalue estimator matter: the original spectral decomposition approach can misbehave, and corrected versions give effective numbers that track permutation results more closely 45. For millions of correlated markers, permutation is the reference standard but is too slow, so multivariate-normal approximations such as SLIDE were developed to reproduce permutation thresholds at a fraction of the cost 6. In eQTL studies, where every gene has its own set of correlated cis-variants, an eigenvalue-based per-gene adjustment (eigenMT) approximates gene-level permutation p-values in orders of magnitude less compute 7.

In your own data

Where it shows up concretely:

  • Differential expression from your RNA-seq. DESeq2 writes padj (BH by default) alongside pvalue. Filter on padj, not pvalue. DESeq2 also sets padj to NA for genes filtered by independent filtering or flagged as count outliers, and those genes are excluded from m. Check sum(is.na(res$padj)) before you conclude a gene “wasn’t significant”.
  • Variant association. PLINK’s --adjust emits BONF, HOLM, FDR_BH, and FDR_BY columns plus a genomic inflation λ. For a single individual you are usually not running association tests, but you will read GWAS summary statistics, and 5 × 10⁻⁸ is the FWER threshold derived from roughly one million independent common-variant tests in Europeans. It is not a universal constant. A denser panel or a different ancestry changes M_eff.
  • Proteomics panels. An Olink or SomaScan run gives you ~3,000-7,000 analytes. Correlations across analytes are strong, so BH on the raw p-value vector is defensible and Bonferroni is heavily conservative.

Common mistakes we see. Correcting after you have already picked the interesting genes by eye, which is the selection you were supposed to correct for. Applying BH separately to each of several subgroups and then pooling the “significant” lists. Forgetting that repeated sampling over time multiplies m: 40 biomarkers measured at 12 visits is 480 tests if you scan all of them. And reporting q-values as if they were per-hypothesis probabilities. A q of 0.05 describes the list, not the gene.

Limitations

Correction controls false positives and costs you power, and with n = 1 the power is often low to begin with. Nothing here converts a self-run screen into a clinical finding. A corrected p-value says the pattern is unlikely under the null, not that the effect is large, causal, or actionable. Region-based tests that aggregate variants across a gene shift the question from “which variant” to “does this gene carry signal”, changing both m and interpretation 8. And correlation-aware thresholds depend on the LD structure of the reference panel you used, which may not match your ancestry. Anything you would act on medically goes to a clinician and, for germline variants, to an accredited confirmatory assay.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Computational Angle

Every scan you run over your own genome, transcriptome, or proteome is a multiple testing problem: 20,000 genes, ~7,000 plasma proteins, or millions of variants, each with its own p-value. Choosing the correction determines how many of your top hits are real.

Related Terms

References

  1. William S Noble. How does multiple testing correction work? . Nature Biotechnology, 2009. DOI
  2. Jelle J. Goeman, Aldo Solari. Multiple hypothesis testing in genomics . Statistics in Medicine, 2014. DOI
  3. Dale R. Nyholt. A Simple Correction for Multiple Testing for Single-Nucleotide Polymorphisms in Linkage Disequilibrium with Each Other . The American Journal of Human Genetics, 2004. DOI
  4. Valentina Moskvina, Karl Michael Schmidt. On multiple‐testing correction in genome‐wide association studies . Genetic Epidemiology, 2008. DOI
  5. Xiaoyi Gao, Joshua Starmer, Eden R. Martin. A multiple testing correction method for genetic association studies using correlated single nucleotide polymorphisms . Genetic Epidemiology, 2008. DOI
  6. Buhm Han, Hyun Min Kang, Eleazar Eskin. Rapid and Accurate Multiple Testing Correction and Power Estimation for Millions of Correlated Markers . PLoS Genetics, 2009. DOI
  7. Joe R. Davis, Laure Fresard, David A. Knowles, et al.. An Efficient Multiple-Testing Adjustment for eQTL Studies that Accounts for Linkage Disequilibrium between Variants . The American Journal of Human Genetics, 2016. DOI
  8. Audrey E Hendricks, Josée Dupuis, Mark W Logue, et al.. Correction for multiple testing in a gene region . European Journal of Human Genetics, 2013. DOI