Removing sample-to-sample cross-contamination in high-throughput sequencing data
While high-throughput sequencing (HTS) has enabled the rapid and inexpensive acquisition of large quantities of genetic sequencing data, HTS data may not fully reflect characteristics of the environment being sequenced. A common step in generating HTS data is multiplexing, which reduces the cost of studying many samples by enabling sample processing in a single sequencing batch. Unfortunately multiplexing can incur sample-to-sample misclassification of observations, with misclassification rates varying by sequencing chemistry and batch. In this paper, we propose a statistical model to link the source sample to observed count data. Because many latent misclassification matrices could connect the source and observed data, directly computing likelihoods is infeasible, and so we propose an importance sampling method to estimate the likelihood of a candidate source matrices. We demonstrate the performance of our proposed method using a simulation study and an application to a vaginal microbiome dataset.
Authors
- Amy D. Willis (ORCID: https://orcid.org/0000-0002-2802-4317)
- Xiaochuan Cecilia Shi
Institutions
- University of Toronto (CA)
- University of Washington (US)
Publication Details
- Journal
- Journal of Applied Statistics
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1080/02664763.2026.2708838
- Primary Topic
- Genomics and Phylogenetic Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00