Microbiome Relative Abundance
The fraction of sequencing reads in a sample assigned to a given taxon, which measures each microbe's share of the community rather than its absolute concentration.
Relative abundance is the proportion of a sample’s sequencing reads assigned to a taxon, after normalizing by the total assigned reads. It tells you a microbe’s share of the community. It tells you nothing, on its own, about how many cells of that microbe were in the stool.
How it works
The raw output of a taxonomic classifier is a count table: rows are taxa, columns are samples, cells are read counts. Relative abundance is that table divided by its column sums. With Kraken2 plus Bracken you get this directly:
kraken2 --db k2_pluspf --paired --confidence 0.1 \
--report sample.kreport R1.fq.gz R2.fq.gz > sample.kraken
bracken -d k2_pluspf -i sample.kreport -o sample.bracken -r 150 -l S -t 10
The fraction_total_reads column in sample.bracken is the relative abundance at species level. MetaPhlAn4 skips the count step entirely and emits percentages from marker-gene coverage in a *_profiled.tsv file.
Two corrections matter and are usually skipped. First, read counts scale with genome length: a 6 Mb Bacteroides genome yields roughly twice the reads of a 3 Mb genome at equal cell counts. Bracken’s Bayesian reassignment handles classification ambiguity but not genome-length normalization, so convert to cell fractions by dividing each taxon’s read fraction by its genome size and renormalizing. Methods built for this estimate genome relative abundance directly from shotgun reads with explicit length and coverage modeling 1, and extensions handle closely related species where reads map ambiguously to several references 2. Second, average genome size differs between communities, which shifts functional-gene-per-genome estimates if you ignore it 3.
The sum-to-one constraint is the structural fact. If Bifidobacterium doubles in absolute terms and nothing else changes, every other taxon’s relative abundance falls. You cannot tell that case apart from a uniform decline in everything else. Compositional methods address this by working in log-ratios: center log-ratio (CLR) values, or additive log-ratios against a chosen reference taxon. LinDA fits linear models on CLR-transformed data and corrects the resulting bias, which keeps the false discovery rate controlled while staying fast on large tables 4. Causal mediation methods for microbiome data build on the same log-ratio geometry 5.
In your own data
Start from the Bracken or MetaPhlAn table and check four things before you plot anything.
Sequencing depth per sample. Below roughly 5 million classified reads, rare taxa become noise. Record depth as a covariate and inspect whether the taxa you care about correlate with it.
Unclassified fraction. Kraken2 with a standard database often leaves 20-50% of stool reads unassigned. Relative abundances are fractions of the classified portion, so a shift in the unclassified share silently rescales everything else. Report it in every figure.
Zeros. A taxon at 0 is usually below detection, not absent. Before a log-ratio transform you need a replacement. We use a multiplicative simple replacement at 65% of the per-sample detection limit rather than adding a flat pseudocount of 1, which distorts low-depth samples the most.
Contamination and transit artifacts. Oral taxa such as Streptococcus and Veillonella appearing at high relative abundance in stool can reflect depletion of the resident gut community rather than true oral expansion, and that pattern has been associated with clinical outcomes in hospitalized patients 6. Interpretation of any such finding belongs with a clinician.
For longitudinal comparison of your own samples, the transform we would use is CLR per sample, then a paired test or mixed model across timepoints. Plot CLR trajectories, not stacked bar charts. Stacked bars encode the constraint visually and make every taxon look like it moved.
Limitations
Relative abundance cannot answer “did my total bacterial load change.” Only an external measure can: qPCR of 16S copies, flow cytometry cell counts, or a spiked-in synthetic cell standard added before extraction. These quantitative microbiome profiling approaches give different answers from each other and from relative profiling, and the choice of method changes which taxa appear to differ between groups 7. Absolute quantification is appropriate for load questions and unnecessary for questions about community structure 8.
Reference database coverage bounds everything. Species present in your gut but absent from the database are either unclassified or misassigned to a relative, and prevalence across individuals varies enormously at the strain level 9. Stool sampling also captures the lumen, not the mucosa, and a single stool sample represents one day of a variable system 10. Take at least three samples over two weeks before treating any number as a baseline.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Your taxonomic profile is a vector of proportions summing to 1, so any change you see in one taxon is entangled with every other taxon; comparing samples across time requires log-ratio transforms or an external quantification step, not raw percentages.
References
- Li C. Xia, Jacob A. Cram, Ting Chen, et al.. Accurate Genome Relative Abundance Estimation Based on Shotgun Metagenomic Reads . PLoS ONE, 2011. DOI
- Michael B Sohn, Lingling An, Naruekamol Pookhao, et al.. Accurate genome relative abundance estimation for closely related species in a metagenomic sample . BMC Bioinformatics, 2014. DOI
- Stephen Nayfach, Katherine S Pollard. Average genome size estimation improves comparative metagenomics and sheds light on the functional ecology of the human microbiome . Genome Biology, 2015. DOI
- Huijuan Zhou, Kejun He, Jun Chen, et al.. LinDA: linear models for differential abundance analysis of microbiome compositional data . Genome Biology, 2022. DOI
- Chan Wang, Jiyuan Hu, Martin J Blaser, et al.. Estimating and testing the microbial causal mediation effect with high-dimensional and compositional microbiome data . Bioinformatics, 2019. DOI
- Chen Liao, Thierry Rolling, Ana Djukovic, et al.. Oral bacteria relative abundance in faeces increases due to gut microbiota depletion and is linked with patient outcomes . Nature Microbiology, 2024. DOI
- Gianluca Galazzo, Niels van Best, Birke J. Benedikter, et al.. How to Count Our Microbes? The Effect of Different Quantitative Microbiome Profiling Approaches . Frontiers in Cellular and Infection Microbiology, 2020. DOI
- Xiaofan Wang, Samantha Howe, Feilong Deng, et al.. Current Applications of Absolute Bacterial Quantification in Microbiome Studies and Decision-Making Regarding Different Biological Questions . Microorganisms, 2021. DOI
- Laurens Kraal, Sahar Abubucker, Karthik Kota, et al.. The Prevalence of Species and Strains in the Human Microbiome: A Resource for Experimental Efforts . PLoS ONE, 2014. DOI
- Andrea D Tyler, Michelle I Smith, Mark S Silverberg. Analyzing the Human Microbiome: A “How To” guide for Physicians . American Journal of Gastroenterology, 2014. DOI