Biological interaction fingerprints improve prediction of drug-induced steatosis

Abstract Drug-induced liver injury (DILI) is a major safety concern in drug development, with steatosis constituting a relevant toxicological phenotype within DILI. Current in vitro assays typically rely on lipid accumulation measurements in hepatocyte models, capturing only selected aspects of the underlying biology. Computational approaches complement these assays by leveraging machine learning to detect multivariate and nonlinear patterns relevant to steatosis. In this study, we developed binary classification models using steatosis-positive drugs curated from the Food and Drug Administration Adverse Event Reporting System (FAERS) and steatosis-negative compounds from the DILIst dataset. Separate models were trained on chemical (PubChem fingerprint) and biological features (target and pathway fingerprints). Biological fingerprints significantly outperformed the chemical fingerprint, highlighting the added predictive value of target- and pathway-level information for steatosis prediction. The outputs of all three models were combined into a consensus classifier using majority voting, which achieved a balanced accuracy of 0.72 and a sensitivity of 0.69 in 5×2 repeated cross-validation, comparable to the individual biological fingerprint models. To challenge model generalizability, the 10% most dissimilar compounds based on maximum Tanimoto similarity were withheld as an external test set. Feature importance analysis highlighted transporter proteins and cytochrome P450 enzymes, with enriched pathways primarily linked to absorption, distribution, metabolism, and excretion (ADME) processes. Moreover, an interaction count for liver-expressed targets differentiates steatosis-positive from -negative compounds, suggesting a simple metric that could be worth exploring for other complex toxicity endpoints. To illustrate practical utility, the framework was applied to compounds with sparse FAERS evidence for steatosis. Overall, integrating chemical and biological features within a consensus machine-learning framework enables a robust steatosis prediction while offering interpretable patterns that help contextualize biological processes associated with steatosis.

Authors

Institutions

Publication Details

Journal
Discover Pharmaceutical Sciences
Published
2026-10-05
DOI
https://doi.org/10.1007/s44395-026-00057-1
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Biological interaction fingerprints improve prediction of drug-induced steatosis

Gerhard Franz Ecker, Florian Oehler, Palle Steen Helmke
Discover Pharmaceutical Sciences
Computational Drug Discovery Methods
article

Biological interaction fingerprints improve prediction of drug-induced steatosis

Gerhard Franz Ecker, Florian Oehler, Palle Steen Helmke
article en

Abstract

Abstract Drug-induced liver injury (DILI) is a major safety concern in drug development, with steatosis constituting a relevant toxicological phenotype within DILI. Current in vitro assays typically rely on lipid accumulation measurements in hepatocyte models, capturing only selected aspects of the underlying biology. Computational approaches complement these assays by leveraging machine learning to detect multivariate and nonlinear patterns relevant to steatosis. In this study, we developed binary classification models using steatosis-positive drugs curated from the Food and Drug Administration Adverse Event Reporting System (FAERS) and steatosis-negative compounds from the DILIst dataset. Separate models were trained on chemical (PubChem fingerprint) and biological features (target and pathway fingerprints). Biological fingerprints significantly outperformed the chemical fingerprint, highlighting the added predictive value of target- and pathway-level information for steatosis prediction. The outputs of all three models were combined into a consensus classifier using majority voting, which achieved a balanced accuracy of 0.72 and a sensitivity of 0.69 in 5×2 repeated cross-validation, comparable to the individual biological fingerprint models. To challenge model generalizability, the 10% most dissimilar compounds based on maximum Tanimoto similarity were withheld as an external test set. Feature importance analysis highlighted transporter proteins and cytochrome P450 enzymes, with enriched pathways primarily linked to absorption, distribution, metabolism, and excretion (ADME) processes. Moreover, an interaction count for liver-expressed targets differentiates steatosis-positive from -negative compounds, suggesting a simple metric that could be worth exploring for other complex toxicity endpoints. To illustrate practical utility, the framework was applied to compounds with sparse FAERS evidence for steatosis. Overall, integrating chemical and biological features within a consensus machine-learning framework enables a robust steatosis prediction while offering interpretable patterns that help contextualize biological processes associated with steatosis.

Discover Pharmaceutical SciencesVol. 2(1)
University of Vienna (AT)
European Commission, National Centre for the Replacement Refinement and Reduction of Animals in Research
Good health and well-being
Openalex Percentile: Top 12%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.