ToxCompl Completion of the DrugMatrix Toxicogenomics Database: An Integrated Resource for Toxicological Hypothesis Generation

The DrugMatrix database contains systematically generated toxicogenomics data from short-term in vivo studies for over 600 chemicals. However, most potential endpoints are missing due to a lack of experimental measurements. Therefore, we leveraged matrix factorization and machine learning methods to predict the missing values, which includes gene expression across eight tissues on two expression platforms along with paired clinical chemistry, hematology, and histopathology. We propose a method, ToxCompl, that applies systematic hybrid sampling guided by Bayesian optimization in conjunction with low-rank matrix factorization to predict the missing values. In-depth validation of the ToxCompl predicted data from machine learning, biological, and toxicological perspectives shows that the predicted differential gene expression aligns well with what would be anticipated. This includes examining the connectivity pattern of predicted gene expression responses, characterizing molecular pathway-level responses from sets of differentially expressed genes, evaluating known transcriptional biomarkers of tissue toxicity, and characterizing predicted apical endpoints. For example, we identified kidney toxicants using the transcriptional biomarker Havcr1. All measured and predicted DrugMatrix data (i.e., gene expression, clinical chemistry, hematology, and histopathology) are available to the public (https://rstudio.niehs.nih.gov/toxcompl/). Notably, predicted clinical chemistry of subtle effects and histopathological prediction are two areas we will continue to improve. The main advantage of the ToxCompl approach is that it drastically extends the toxicogenomic landscape into many data-poor tissues in the absence of acquiring additional experimental data, thereby allowing researchers to formulate mechanistic hypotheses about effects in tissues that have been underrepresented in the literature.

Authors

Institutions

Publication Details

Journal
Toxicological Sciences
Published
2026-09-04
DOI
https://doi.org/10.1093/toxsci/kfag113
Primary Topic
Gene expression and cancer classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ToxCompl Completion of the DrugMatrix Toxicogenomics Database: An Integrated Resource for Toxicological Hypothesis Generation

Daniel Svoboda, Robert M. Patton, Jeremy N. Erickson, Laura J Word et al.
Toxicological Sciences
Gene expression and cancer classification
article

ToxCompl Completion of the DrugMatrix Toxicogenomics Database: An Integrated Resource for Toxicological Hypothesis Generation

Daniel Svoboda, Robert M. Patton, Jeremy N. Erickson, Laura J Word, Warren Casey, Charles Schmitt, Guojing Cong, Frank Chao, Parker Combs, Scott S Auerbach
article en

Abstract

The DrugMatrix database contains systematically generated toxicogenomics data from short-term in vivo studies for over 600 chemicals. However, most potential endpoints are missing due to a lack of experimental measurements. Therefore, we leveraged matrix factorization and machine learning methods to predict the missing values, which includes gene expression across eight tissues on two expression platforms along with paired clinical chemistry, hematology, and histopathology. We propose a method, ToxCompl, that applies systematic hybrid sampling guided by Bayesian optimization in conjunction with low-rank matrix factorization to predict the missing values. In-depth validation of the ToxCompl predicted data from machine learning, biological, and toxicological perspectives shows that the predicted differential gene expression aligns well with what would be anticipated. This includes examining the connectivity pattern of predicted gene expression responses, characterizing molecular pathway-level responses from sets of differentially expressed genes, evaluating known transcriptional biomarkers of tissue toxicity, and characterizing predicted apical endpoints. For example, we identified kidney toxicants using the transcriptional biomarker Havcr1. All measured and predicted DrugMatrix data (i.e., gene expression, clinical chemistry, hematology, and histopathology) are available to the public (https://rstudio.niehs.nih.gov/toxcompl/). Notably, predicted clinical chemistry of subtle effects and histopathological prediction are two areas we will continue to improve. The main advantage of the ToxCompl approach is that it drastically extends the toxicogenomic landscape into many data-poor tissues in the absence of acquiring additional experimental data, thereby allowing researchers to formulate mechanistic hypotheses about effects in tissues that have been underrepresented in the literature.

Toxicological Sciences
American Medical Informatics Association (US), Oak Ridge National Laboratory (US), National Institute of Environmental Health Sciences (US), Institute of Contemporary History (SI), Depomed (United States) (US)
Openalex Percentile: Top 17%
Gene expression and cancer classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.