Reference Incompleteness in Human Gut Metagenomics Is a Geographically Structured Measurement Bias
Gut microbiome research turns sequencing reads into taxa and functions by matching them against a reference database of isolate and metagenome-assembled genomes. That database is the measuring instrument, yet its sensitivity is never reported. Reference incompleteness is therefore not a computational nuisance but a measurement bias; because the gaps are geographically structured, it behaves as differential misclassification rather than random error. The magnitudes are substantial. On identical data, the share of reads receiving a taxonomic assignment runs from 44% to 94.5% depending only on whether a standard Kraken2 database, UHGG, or a gut-specific catalog is used; missing genomes are enriched in non-Westernized populations. Formally, expected observed richness is true richness multiplied by population-specific coverage, so comparing two populations estimates the true ratio times the ratio of their coverages, a second term no study reports. In a simulation, a 5-percentage-point coverage gap raises the false-positive rate for a between-population richness comparison to 55%, a proof-of-concept rather than a real-world estimate. Reference bias distorts between-population comparisons in a determinate direction rather than simply blurring them and does so invisibly: batch correction leaves it untouched and quality control never sees it. We distinguish it from sampling bias, batch effects, and non-differential misclassification, and propose Minimum Information for Reference Reporting (MIRR).
Authors
- Soumok Sadhu (ORCID: https://orcid.org/0000-0003-4186-254X)
- Ram Hari Dahal (ORCID: https://orcid.org/0000-0001-5887-6898)
Institutions
- University of Minnesota (US)
Publication Details
- Journal
- Journal of Genome Biotechnology and Genetics
- Published
- 2026-10-04
- DOI
- https://doi.org/10.3390/jgbg1030018
- Primary Topic
- Gut microbiota and health
- Type
- article
- Field-Weighted Citation Impact
- 0.00