Confounding
A third variable that causes both the thing you measured and the outcome you care about, making a non-causal association look real.
Confounding is when a third variable causes both your predictor and your outcome, so the association between them is real in the data and wrong as a causal claim.
Bias and confounding get conflated, and they are different failures. Bias comes from how you collected or measured: a pulse oximeter that reads high on darker skin, a self-report questionnaire people answer aspirationally, a cohort where sicker patients drop out. You cannot fix it after the fact with a covariate. Confounding comes from the causal structure of the world: the variable exists, it affects both sides, and if you measured it you can adjust for it. Bias is a defect in the measurement pipeline. Confounding is a defect in the inference.
How it works
The textbook case: coffee drinkers have more lung cancer. Coffee drinking associates with smoking, smoking causes lung cancer, so coffee inherits a false effect. A confounder has to satisfy three conditions. It is associated with the exposure, it independently affects the outcome, and it is not on the causal path between them. That last condition matters. If you adjust for something downstream of your exposure, you are not removing confounding, you are deleting the effect you are trying to see.
Molecular data has its own recurring confounders. Population stratification is the canonical one: allele frequencies differ by ancestry, disease prevalence differs by ancestry, and a case-control study that does not match on ancestry will produce associations at loci with nothing to do with the disease. The CYP3A4-V variant and prostate cancer in African Americans is a clean worked example of exactly this ambiguity between a causal and a stratification-driven association.1 Self-identified race is a coarse proxy for genetic structure and does not control it reliably.2 The same problem recurs in studies of metabolic gene polymorphisms, where ethnicity, smoking, and diet track together and with the genotype of interest.3
For expression, proteomics, and any high-dimensional assay, the dominant confounders are technical and demographic. Age and sex alone shift blood microRNA profiles enough to generate false disease signatures if left unmodeled.4 Inflammatory biomarkers like CRP and IL-6 move with BMI, smoking, infection, circadian phase, and time since last meal, which is why isolated readings support very little.5 And the general shape of the problem in omics, where confounders are often unrecorded and correlated with the group labels by accident of study design, is well characterized.6
In your own data
Start by building a sample metadata table before you touch the matrices: collection date, draw time, fasting hours, sequencing run, flow cell lane, library prep batch, plate and well for proteomics, RIN, and any illness or hard workout in the prior 72 hours. In a longitudinal profile this is your entire defense.
Concrete checks:
- Run PCA on the log-transformed expression matrix (
prcomp(t(log2(cpm+1)))) and regress each of PC1–PC10 against every metadata column. If PC1 correlates with batch at r > 0.5, batch is your leading signal, not biology. - Model it rather than subtract it. In DESeq2 use
design = ~ batch + sex + condition, not a corrected matrix fed back in as if it were raw counts.limma::removeBatchEffect()is for visualization only; using its output as input to a test understates variance and inflates significance. - For unrecorded confounders,
sva::sva()or PEER factors estimate latent components. Include them as covariates. Check first that the surrogate variables are not collinear with your variable of interest, or you will regress away the signal. - For genotype work, compute ancestry PCs with PLINK on LD-pruned common variants (
--indep-pairwise 50 5 0.2, then--pca 10) and carry the top PCs in every association model. - For CGM, an apparent food effect is usually confounded by time of day, prior-meal carryover, and activity. Compare the same food at the same hour on different days before concluding anything.
The standard test for whether a variable is confounding your estimate is the change-in-estimate rule: fit with and without the covariate and see if the effect size moves more than about 10%. Covariate adjustment changes which genes come out as associated, and it changes downstream meta-analysis results too.7 If you are predicting one molecular readout from another, check whether the apparent skill comes from a shared upstream driver rather than the biology you assume. Image-based biomarker predictors show this plainly: performance that looks specific turns out to ride on interdependencies between biomarkers and site-level artifacts.8
Limitations
Adjustment only works for confounders you measured, measured well, and modeled in the right functional form. Residual confounding survives all three. Observational epidemiology has a long record of confidently adjusted associations that failed in trials.9 Genetic instruments are one partial escape, since alleles are assigned at conception and are largely independent of the social and behavioral clustering that drives conventional confounding.10 They come with their own assumptions, including pleiotropy and stratification, and they need large samples, so they are a tool for reading the literature more than for analyzing n=1 data.
For any result that would change a medical decision, bring it to a clinician who can see the rest of your chart.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
In personal molecular data, the confounders are usually mundane and recorded somewhere in your own metadata: draw time, fasting state, sequencing batch, plate position, sleep, recent exercise, age, sex, and genetic ancestry. Most of the work is joining that metadata to your matrices and putting it in the model.
Related Terms
References
- Rick A. Kittles, Weidong Chen, Ramesh K. Panguluri, et al.. CYP3A4-V and prostate cancer in African Americans: causal or confounding association because of population stratification? . Human Genetics, 2002. DOI
- Hua Tang, Tom Quertermous, Beatriz Rodriguez, et al.. Genetic Structure, Self-Identified Race/Ethnicity, and Confounding in Case-Control Association Studies . The American Journal of Human Genetics, 2005. DOI
- Emanuela Taioli, Seymour Garte. Covariates and confounding in epidemiologic studies using metabolic gene polymorphisms . International Journal of Cancer, 2002. DOI
- Benjamin Meder, Christina Backes, Jan Haas, et al.. Influence of the Confounding Factors Age and Sex on MicroRNA Profiles from Peripheral Blood . Clinical Chemistry, 2014. DOI
- Qurrat Ul Ain, Mehak Sarfraz, Gayuk Kalih Prasesti, et al.. Confounders in Identification and Analysis of Inflammatory Biomarkers in Cardiovascular Diseases . Biomolecules, 2021. DOI
- Wilson Wen Bin Goh, Limsoon Wong. Dealing with Confounders in Omics Analysis . Trends in Biotechnology, 2018. DOI
- Xingbin Wang, Yan Lin, Chi Song, et al.. Detecting disease-associated genes with confounding variable adjustment and the impact on genomic meta-analysis: With application to major depressive disorder . BMC Bioinformatics, 2012. DOI
- Muhammad Dawood, Kim Branson, Sabine Tejpar, et al.. Confounding factors and biases abound when predicting molecular biomarkers from histological images . Nature Biomedical Engineering, 2026. DOI
- Shah Ebrahim, George Davey Smith. Mendelian randomization: can genetic epidemiology help redress the failures of observational epidemiology? . Human Genetics, 2007. DOI
- George Davey Smith, Debbie A Lawlor, Roger Harbord, et al.. Clustered Environments and Randomized Genes: A Fundamental Distinction between Conventional and Genetic Epidemiology . PLoS Medicine, 2007. DOI