Exploring transcriptomic and genomic latent variable correction approaches in differential expression analysis

Abstract Differential expression analysis is a central tool for studying the biological processes altered in human diseases via transcriptomic signatures. However, transcriptomic datasets are systematically confounded by latent variables from two distinct sources: unmeasured technical and biological heterogeneity within the expression data, and expression differences driven by population stratification. Correction using expression-based surrogate variables (SVs) and genotype-based principal components (PCs) addresses these sources independently, yet no study has directly evaluated their combined use against either method alone within a differential expression framework. In this study we hypothesised that simultaneously including both correction layers would produce more biologically valid and reproducible results than either approach alone, and tested this in two independent post-mortem RNA-seq datasets of amyotrophic lateral sclerosis (ALS) cases and controls with paired genotype data. Four nested differential expression models (corrected for PC-only, SV-only, both SV and PC, and neither PCs nor SVs) were evaluated across the KCLBB (96 cases and 52 controls) and ALS Consortium (272 cases and 35 controls) datasets. Models were evaluated on cross-dataset effect size concordance, cross-dataset replicability quantified by the Jaccard Similarity Index, and biological recall against a curated reference set of 66 known ALS genes. The combined SV + PC framework outperformed simpler models across most metrics. Replicability improved nearly ten-fold compared to the non-corrected model, (Jaccard index: 2.28% to 19.5%), and the combined framework exhibited a statistically significant 2.2% gain over the SV-only model. Biological recall doubled compared to SV correction alone. Effect size magnitude consistency was preserved, though a reduction in Spearman’s rank correlation indicates PC correction refines magnitude rather than gene rankings. These findings remained generally robust to genotype PC number and sequencing platform adjustment, with more variable performance observed under ancestry restriction and upon extension to an independent tissue. This study found that SVs and genotype PCs address non-redundant sources of confounding, and we recommend their combined use as standard practice in differential expression analysis where paired genotype data are available, with the greatest benefit expected in ancestrally diverse cohorts. Notably PCs capturing population structure can also be derived directly from RNA-seq data, potentially extending this framework’s applicability to studies lacking paired genotype data. Although this analysis was restricted to ALS datasets, we expect these findings to generalise to other traits.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-11
DOI
https://doi.org/10.1038/s41598-026-69619-8
Primary Topic
Gene expression and cancer classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Exploring transcriptomic and genomic latent variable correction approaches in differential expression analysis

Kerstin Lindblad‐Toh, William J. Salerno, Oleg Butovsky, Vilas Menon et al.
Scientific Reports
Gene expression and cancer classification
article

Exploring transcriptomic and genomic latent variable correction approaches in differential expression analysis

Kerstin Lindblad‐Toh, William J. Salerno, Oleg Butovsky, Vilas Menon, Stephen Muljo, Michael Cantor, Matt Brauer, Sabrina Paganoni, Lyle W. Ostrow, Efthimios Dardiotis, Timothy M. Miller, Suma Babu, Dhruv Sareen, Kendall R. Van Keuren‐Jensen, Bryan J. Traynor, James D. Berry, Sulev Kõks, Marc Gotkine, Mickey Atwal, Eran Hornstein, Vivianna M. Van Deerlin, Leslie Michels Thompson, Ximena Arcila-Londono, Ernest Fraenkel, David A. Frendewey, Darius J. Adams, Towfique Raj, Daphne Koller, Dale J. Lange, Pietro Fratta, Noah Zaitlen, Jeongho Lee, Edward B. Lee, Giovanni Coppola, T. Nickerson, Gregory A. Cox, Nikolaos A. Patsopoulos, Jennifer E. Phillips‐Cremins, Orit Rozenblatt–Rosen, Eli Stahl, Ophir Shalem, Neil A. Shneider, Robert Moccia, Robert H. Baloh, Daniel J. L. MacGowan, Oliver Pain, Leonidas Stefanis, Suvankar Pal, Matt Anderson, Andrew Deubler, Vivian E. Drory, Yossef Lerner, Mary Rozenman, John Jammal, Steven Altschuler, Colin Smith, Corey McMillan, John Ravits, John Crary, Yadusayan Appulingam, Simon Topp, Hemali P. Phatnani, Shameek Biswas, Alfredo Iacoangeli, Molly Hammell, James R. Broach, Iris Broce, Rita Sattler, Frank Baas, Matthew Harms, Mary Poss, Nazem Atassi, Zachary Simmons, Bin Zhang, Kimberly A. Wilson, Peter Gregersen, Andrea Malaspina, Robert Bowser, Eleonora Aronica, Lani Wu, Seng Cheng, Aminah Ali, Justin Kwan, Terry Heiman-Patterson, Steve Finkbeiner, Thomas Blanchard, Katharine Nicholson, Brent Harris, Joshua Dubnau, Avindra Nath, Siddharthan Chandran
article en

Abstract

Abstract Differential expression analysis is a central tool for studying the biological processes altered in human diseases via transcriptomic signatures. However, transcriptomic datasets are systematically confounded by latent variables from two distinct sources: unmeasured technical and biological heterogeneity within the expression data, and expression differences driven by population stratification. Correction using expression-based surrogate variables (SVs) and genotype-based principal components (PCs) addresses these sources independently, yet no study has directly evaluated their combined use against either method alone within a differential expression framework. In this study we hypothesised that simultaneously including both correction layers would produce more biologically valid and reproducible results than either approach alone, and tested this in two independent post-mortem RNA-seq datasets of amyotrophic lateral sclerosis (ALS) cases and controls with paired genotype data. Four nested differential expression models (corrected for PC-only, SV-only, both SV and PC, and neither PCs nor SVs) were evaluated across the KCLBB (96 cases and 52 controls) and ALS Consortium (272 cases and 35 controls) datasets. Models were evaluated on cross-dataset effect size concordance, cross-dataset replicability quantified by the Jaccard Similarity Index, and biological recall against a curated reference set of 66 known ALS genes. The combined SV + PC framework outperformed simpler models across most metrics. Replicability improved nearly ten-fold compared to the non-corrected model, (Jaccard index: 2.28% to 19.5%), and the combined framework exhibited a statistically significant 2.2% gain over the SV-only model. Biological recall doubled compared to SV correction alone. Effect size magnitude consistency was preserved, though a reduction in Spearman’s rank correlation indicates PC correction refines magnitude rather than gene rankings. These findings remained generally robust to genotype PC number and sequencing platform adjustment, with more variable performance observed under ancestry restriction and upon extension to an independent tissue. This study found that SVs and genotype PCs address non-redundant sources of confounding, and we recommend their combined use as standard practice in differential expression analysis where paired genotype data are available, with the greatest benefit expected in ancestrally diverse cohorts. Notably PCs capturing population structure can also be derived directly from RNA-seq data, potentially extending this framework’s applicability to studies lacking paired genotype data. Although this analysis was restricted to ALS datasets, we expect these findings to generalise to other traits.

Scientific Reports
Openalex Percentile: Top 23%
Gene expression and cancer classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.