Skip to content

Whole Genome Testing: What You Get, What It Costs, and What It Can Answer

Woolf Software
A complete iridescent feathered reptile on a specimen stand, lit against darkness, beside a second mount holding only scattered pinned feathers.

Whole genome sequencing reads essentially all of your roughly 3.1 billion base pairs, usually at 30x average depth, meaning each position is covered by about thirty independent sequencing reads on average. What comes back is a set of files rather than a conclusion: FASTQ or CRAM alignments holding the reads themselves, a VCF (variant call format file) listing small differences from the reference genome, and, if the lab is doing its job, separate calls for structural variants, copy number changes, and repeat expansions. The test is worth buying if you want the data and intend to work with it, or to hand it to something that can. It is a poor purchase if what you want is a one-page verdict on your health, because the hard and expensive part is the interpretation rather than the sequencing 1. For comparison, consumer array tests genotype roughly 600,000 to 1 million preselected sites. WGS calls on the order of 4 to 5 million variants per person, most of them noncoding and most of them uninterpretable today.

What a whole genome test produces

Before evaluating any vendor, it helps to know which files retain their value over time, since some of them can be regenerated later and some cannot. The artifacts you should insist on, ordered by durability, are these:

  • FASTQ (gzipped), paired-end 2x150. Roughly 40-60 GB for a 30x human genome. This is the rawest thing you will get and the only input you can fully re-analyze with a future pipeline.
  • CRAM aligned to GRCh38 (or T2T-CHM13 v2.0), with the reference FASTA and its .fai/.dict. CRAM is reference-compressed: about 15-25 GB at 30x versus ~90-110 GB for the equivalent BAM. Store the exact reference; a CRAM without its reference is a brick.
  • gVCF, not just VCF. A plain VCF tells you where you differ from the reference. A gVCF also encodes reference-confidence blocks, so you can tell “homozygous reference, 34 reads” apart from “no coverage, no call.” That distinction matters every time you check whether a specific position was interrogated at all.
  • Separate SV/CNV VCFs and a repeat-expansion catalog output.

The practical test of a vendor is what they will hand over. If all you can get is a PDF and a filtered “clinically relevant variants” VCF, you are buying a report rather than a genome.

Coverage and chemistry: 30x PCR-free is the floor

Average depth is the number that appears in marketing copy, and it hides the distribution underneath. What you care about is the fraction of the genome at ≥10x and ≥20x, and whether the exons of genes you might one day query are covered at all. Ask for mosdepth or Picard CollectWgsMetrics output alongside the CRAM so you can check. At 30x mean on a PCR-free library, expect more than 95% of the genome at ≥20x. PCR-amplified libraries lose GC-extreme regions and inflate duplicate rates, with 15-20% duplicates being common under PCR versus 1-3% for PCR-free preparation, so your effective depth ends up lower than advertised.

Uniform coverage is precisely where WGS outperforms targeted panels. In a hypertrophic cardiomyopathy cohort, whole genome sequencing achieved more even coverage of cardiomyopathy genes than panel capture, whose exons sometimes dropped below callable depth 2. The same technology has clear blind spots. Short reads cannot resolve segmental duplications, where near-identical sequence appears in more than one place in the genome. The standard examples are PMS2 exons 11-15 against PMS2CL, SMN1/SMN2, and the HBA locus. CGG-rich repeats such as FMR1 remain unreliable at any depth with 150 bp reads. If one of those loci is the actual question, a targeted or long-read assay is the right instrument.

A pipeline we would run

Given FASTQs, the following stack is what we would use, with the reasoning for each choice.

  1. QC: fastp or FastQC, then alignment with bwa-mem2 mem -K 100000000 -Y -R '@RG\tID:...\tSM:...\tPL:ILLUMINA'. The -K flag fixes the batch size so output is deterministic across thread counts, which you want if you ever need to reproduce a run.
  2. Duplicate marking with samtools markdup or Picard, then CRAM with samtools view -T GRCh38.fa -C -o sample.cram.
  3. Small variants with DeepVariant (--model_type=WGS) rather than GATK HaplotypeCaller. Benchmark samples come from Genome in a Bottle (GIAB), the reference materials used to measure caller accuracy. On those samples, DeepVariant at 30x typically posts SNV F1 above 99.5% and indel F1 in the high 99s inside high-confidence regions, and it needs no base quality score recalibration step. If you want joint calling across family members, GLnexus merges DeepVariant gVCFs.
  4. Structural variants: Manta plus a second caller (Delly or GRIDSS), taking the intersection seriously and the union skeptically. Short-read SV calling has genuinely poor precision for insertions and for anything inside repeats.
  5. Copy number: cn.MOPS or GATK gCNV if you have a panel of normals; a single-sample WGS CNV call with no control cohort is noisy.
  6. Repeat expansions: ExpansionHunter with the standard variant catalog (HTT, ATXN1-3, C9orf72, DMPK and friends). Treat FMR1 output as a screening signal only.
  7. Annotation: VEP with --everything --assembly GRCh38, plus ClinVar, gnomAD v4 (730,947 exomes and 76,215 genomes), SpliceAI, and REVEL. Then filter.
  8. Pharmacogenomics: PharmCAT on the VCF, with Aldy or Stargazer for CYP2D6 star alleles. CYP2D6 has structural variation such as deletions, duplications, and CYP2D6/CYP2D7 hybrids that a SNV-only caller will get wrong. PharmCAT output is a set of predicted metabolizer phenotypes. Take them to a clinician or pharmacist; they are not dosing instructions and we will not give any.
  9. Polygenic scores: plink2 --score against PGS Catalog weight files. With WGS you skip imputation entirely, which removes one large source of error, though ancestry-transfer bias in the weights remains and dominates the uncertainty.

The resource requirements are modest by modern standards. On a 32-core machine, alignment through variant calling runs 8-20 hours, and on a cloud instance it amounts to a few dollars of compute per genome.

What the genome answers well, and what it does not

Setting expectations correctly is most of the value of this section, because the same file can be highly informative for one question and nearly silent on another. The genome answers well on carrier status for recessive conditions with well-characterized variants, confirmed pathogenic variants in the ACMG secondary-findings list (v3.2, 81 genes), pharmacogenomic star alleles, ancestry and relatedness, and the provision of a personal reference sequence you never have to buy again.

It answers badly on most common-disease risk, where effect sizes are small and polygenic scores are calibrated to cohorts that may not match your ancestry. It also answers badly on any variant classified as uncertain significance, which describes most of them. Interpretation is the rate limiter here, rather than data generation. A systematic review of WGS in healthy adults found that the yield of clearly actionable findings is low and that the burden of uncertain results is substantial 3. Early clinical WGS work quantified this directly: manual curation of variants in a modest set of disease-associated genes took hours of specialist time per person, and reviewers frequently disagreed about pathogenicity 1. The gap between having sequence and having meaning has been the recurring theme since WGS entered clinical use 4, and it persists in current implementation reviews 5.

Quantitative traits deserve a note of their own. Genome-wide rare-variant analysis needs aggregation across many people to have statistical power, so inference from a single genome about a lipid level or a blood pressure reading is weak 6.

Cost and insurance

Prices in this field are widely misquoted, so it is worth separating the reagent cost from what a patient pays. The $1,000 genome was a 2005 target and a reagent-cost figure rather than a clinical price 7. Today the sequencing itself costs a few hundred dollars at scale, clinical WGS with interpretation and a signed report typically runs $2,000-$5,000 in the US, and research-grade 30x WGS with raw data delivery runs several hundred dollars to about $1,500.

Insurance denials follow a predictable set of reasons. There may be no documented clinical indication, or a phenotype that a cheaper targeted panel would address. Others are ordering by a non-genetics provider or a payer policy that classifies WGS in asymptomatic adults as screening. Cost-effectiveness evidence for WGS is heterogeneous and depends heavily on the clinical setting and the comparator, which is part of why payers hedge 8.

Downsides worth naming

Any honest account of whole genome testing has to include the costs that are not financial. Variants of uncertain significance are the main one: you will receive a long list, and the follow-up testing they prompt costs money and creates anxiety without changing anything. Secondary findings arrive whether or not you asked for them. Family implications are real, since your genome is partly your siblings’ and your children’s.

Privacy and insurance discrimination protections vary by jurisdiction. In the US the Genetic Information Nondiscrimination Act (GINA) does not cover life, disability, or long-term care insurance. These tradeoffs are argued most sharply in the debate over sequencing newborns, where the same data raises questions about consent and the right to decide later 9, and they underlie professional guidance that sequencing in minors be handled conservatively 10.

Questions people also ask

Is whole genome testing worth it? If you want the raw data and will analyze it, yes; the marginal cost over an exome or panel is small and the files remain reusable for decades. If what you want is a health verdict, the yield of clearly actionable findings in a healthy adult is low 3.

How much does a whole genome sequencing test cost? Roughly $400-$1,500 for research-grade 30x WGS with raw data, and $2,000-$5,000 for a clinical-grade test with an interpreted report.

Does insurance cover whole genome testing? Sometimes, for diagnostic use in a patient with an unexplained phenotype, particularly in pediatrics and critical care. Predictive testing in an asymptomatic adult is usually self-pay.

What is the downside of genetic testing? Uncertain results you cannot act on, incidental findings you did not seek, and downstream testing costs. There are also implications for relatives who did not consent.

Is WGS better than a panel? For coverage uniformity and future reanalysis, yes 2. For loci inside segmental duplications or long repeats, short-read WGS performs worse than a purpose-built assay.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Frederick E. Dewey, Megan E. Grove, Cuiping Pan, et al. Clinical Interpretation and Implications of Whole-Genome Sequencing. JAMA, 2014. https://doi.org/10.1001/jama.2014.1717 2

  2. Allison L. Cirino, Neal K. Lakdawala, Barbara McDonough, et al. A Comparison of Whole Genome Sequencing to Multigene Panel Testing in Hypertrophic Cardiomyopathy Patients. Circulation Cardiovascular Genetics, 2017. https://doi.org/10.1161/circgenetics.117.001768 2

  3. Noralane M. Lindor, Stephen N. Thibodeau, Wylie Burke. Whole-Genome Sequencing in Healthy People. Mayo Clinic Proceedings, 2017. https://doi.org/10.1016/j.mayocp.2016.10.019 2

  4. Kelly E. Ormond, Matthew T. Wheeler, Louanne Hudgins, et al. Challenges in the clinical application of whole-genome sequencing. The Lancet, 2010. https://doi.org/10.1016/s0140-6736(10)60599-5

  5. Petar Brlek, Luka Bulić, Matea Bračić, et al. Implementing Whole Genome Sequencing (WGS) in Clinical Practice: Advantages, Challenges, and Future Perspectives. Cells, 2024. https://doi.org/10.3390/cells13060504

  6. Alanna C. Morrison, Zhuoyi Huang, Bing Yu, et al. Practical Approaches for Whole-Genome Sequence Analysis of Heart- and Blood-Related Traits. The American Journal of Human Genetics, 2017. https://doi.org/10.1016/j.ajhg.2016.12.009

  7. Simon T. Bennett, C.L. Barnes, Anthony J. Cox, et al. Toward the $1000 Human Genome. Pharmacogenomics, 2005. https://doi.org/10.1517/14622416.6.4.373

  8. Katharina Schwarze, James Buchanan, Jenny C. Taylor, et al. Are whole-exome and whole-genome sequencing approaches cost-effective? A systematic review of the literature. Genetics in Medicine, 2018. https://doi.org/10.1038/gim.2017.247

  9. Csaba Szalai. Arguments for and against the whole-genome sequencing of newborns. PubMed, 2023. https://pubmed.ncbi.nlm.nih.gov/37969196

  10. Wayne W. Grody, Barry H. Thompson, Louanne Hudgins. Whole-Exome/Genome Sequencing and Genomics. PEDIATRICS, 2013. https://doi.org/10.1542/peds.2013-1032e