Skip to content
/genetic-scoring/gwas

GWAS (Genome-Wide Association Study)

A study design that tests millions of common genetic variants across a cohort for statistical association with a trait or disease, producing per-variant effect sizes and p-values rather than a causal mechanism.

A genome-wide association study is a regression run millions of times: for each common variant in the genome, test whether carrying an extra copy of one allele shifts a trait value or disease odds across a cohort, then correct hard for having asked the question a million times. The output is a table, one row per variant, with an effect size, a standard error, and a p-value. That table, called summary statistics, is the actual product.

How it works

The unit of analysis is a single nucleotide polymorphism (SNP) with a minor allele frequency high enough to be measurable, typically above 1%. Each person’s genotype at that SNP is coded 0, 1, or 2 copies of the effect allele. For a continuous trait like LDL-C or height you fit a linear model; for a case/control disease you fit logistic regression, in both cases with covariates for age, sex, genotyping batch, and the top 5 to 20 principal components of the genotype matrix to absorb population structure 1. Modern studies use linear mixed models (BOLT-LMM, REGENIE, SAIGE) instead, because they handle relatedness and unbalanced case/control ratios without inflating the test statistic.

The significance threshold is 5×10⁻⁸, a Bonferroni correction for roughly one million independent common-variant tests in European-ancestry samples 2. That number is the reason GWAS needs enormous cohorts. Effect sizes for common variants on complex traits are small, odds ratios frequently in the 1.05 to 1.20 range, so power comes from n, not from signal strength 34.

Almost no hit is the causal variant. Genotyping arrays measure 500k to 2M markers and the rest is imputed against a reference panel (1000 Genomes, HRC, TOPMed), so a significant SNP is usually a tag for a block of correlated variants in linkage disequilibrium. Fine-mapping and functional work come after 3.

Canonical examples: complement factor H and age-related macular degeneration, an early hit with an unusually large effect; the 2007 Wellcome Trust Case Control Consortium study of 7 diseases in 14,000 cases; FTO and body mass index; hundreds of loci for type 2 diabetes, schizophrenia, Crohn’s disease, coronary artery disease, lipid levels, and height 54. The NHGRI-EBI GWAS Catalog is the canonical index of published associations and increasingly hosts the full summary statistics files 6. GWAS Central serves a comparable role for comparing studies and querying results across them 7.

In your own data

You do not run a GWAS on yourself. n=1 has no variance to regress on. What you do is consume other people’s summary statistics against your own genotypes.

Concretely: from whole-genome sequencing you have a VCF (or the gVCF it came from). From the GWAS Catalog you download a summary statistics file, typically tab-separated with columns for chromosome, base_pair_location, effect_allele, other_allele, beta or odds_ratio, standard_error, p_value, and effect_allele_frequency. The join is on position and alleles.

Things that break, in the order they bite:

  • Genome build. Catalog files are a mix of GRCh37 and GRCh38. Your VCF is probably GRCh38. Lift the summary stats with CrossMap.py vcf or liftOver, never assume. A silent build mismatch produces a score that is pure noise and looks entirely plausible.
  • Strand and allele coding. A/T and C/G SNPs are ambiguous on strand. The standard fix is to drop palindromic SNPs with allele frequency near 0.5, which plink2 --score will not do for you.
  • Effect allele versus reference allele. These are not the same field and swapping them flips the sign of every contribution.
  • Direction of beta versus log(OR). Check the file header, not the paper abstract.

For scoring, plink2 --score file.txt 1 2 3 header cols=+scoresums with an explicit allele column is the baseline. For anything serious, use PRS-CS or LDpred2, which reweight effects using an LD reference panel instead of naive clumping.

The ancestry check matters most. Roughly 80% of GWAS participants have been of European ancestry, and effect estimates transfer poorly across ancestries because LD structure and allele frequencies differ 38. If your genome is not European, a European-derived score is systematically miscalibrated, usually attenuated, and there is no correction factor you can apply at home.

Limitations

GWAS identifies loci, not genes and not mechanisms. Most hits sit in non-coding regions and the nearest gene is often the wrong gene 3. The effects are population averages of common variants, so rare high-penetrance variants are invisible to the design, and heritability estimated from GWAS hits usually falls well short of family-study heritability 48. Multiple-testing correction, imputation quality, and the sheer scale of the data are ongoing sources of error 910.

A GWAS-derived score is a population-level statistic applied to you. It is not a diagnosis and it does not tell you what to do. Anything you would act on clinically belongs with a physician or a genetic counselor, using a validated clinical test rather than a research-grade pipeline you assembled yourself.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Computational Angle

GWAS summary statistics are the input files behind almost every polygenic score you will ever compute on your own genome; knowing how they were produced tells you which ones are usable with your VCF and which will silently mislead you.

Related Terms

References

  1. Ben Hayes. Overview of Statistical Methods for Genome-Wide Association Studies (GWAS) . Methods in Molecular Biology, 2013. DOI
  2. Thomas A. Pearson. How to Interpret a Genome-wide Association Study . JAMA, 2008. DOI
  3. Vivian Tam, Nikunj Patel, Michelle Turcotte, et al.. Benefits and limitations of genome-wide association studies . Nature Reviews Genetics, 2019. DOI
  4. Barbara E Stranger, Eli A Stahl, Towfique Raj. Progress and Promise of Genome-Wide Association Studies for Human Complex Trait Genetics . Genetics, 2011. DOI
  5. A. Corvin, N. Craddock, P. F. Sullivan. Genome-wide association studies: a primer . Psychological Medicine, 2009. DOI
  6. Annalisa Buniello, Jacqueline A L MacArthur, Maria Cerezo, et al.. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019 . Nucleic Acids Research, 2018. DOI
  7. Tim Beck, Robert K Hastings, Sirisha Gollapudi, et al.. GWAS Central: a comprehensive resource for the comparison and interrogation of genome-wide association studies . European Journal of Human Genetics, 2013. DOI
  8. Yuval B. Simons, Kevin Bullaughey, Richard R. Hudson, et al.. A population genetic interpretation of GWAS findings for human quantitative traits . PLOS Biology, 2018. DOI
  9. Jason H. Moore, Folkert W. Asselbergs, Scott M. Williams. Bioinformatics challenges for genome-wide association studies . Bioinformatics, 2010. DOI
  10. A. Uitterlinden. An Introduction to Genome-Wide Association Studies: GWAS for Dummies . Seminars in Reproductive Medicine, 2016. DOI