Biological interaction fingerprints improve prediction of drug-induced steatosis
Abstract Drug-induced liver injury (DILI) is a major safety concern in drug development, with steatosis constituting a relevant toxicological phenotype within DILI. Current in vitro assays typically rely on lipid accumulation measurements in hepatocyte models, capturing only selected aspects of the underlying biology. Computational approaches complement these assays by leveraging machine learning to detect multivariate and nonlinear patterns relevant to steatosis. In this study, we developed binary classification models using steatosis-positive drugs curated from the Food and Drug Administration Adverse Event Reporting System (FAERS) and steatosis-negative compounds from the DILIst dataset. Separate models were trained on chemical (PubChem fingerprint) and biological features (target and pathway fingerprints). Biological fingerprints significantly outperformed the chemical fingerprint, highlighting the added predictive value of target- and pathway-level information for steatosis prediction. The outputs of all three models were combined into a consensus classifier using majority voting, which achieved a balanced accuracy of 0.72 and a sensitivity of 0.69 in 5×2 repeated cross-validation, comparable to the individual biological fingerprint models. To challenge model generalizability, the 10% most dissimilar compounds based on maximum Tanimoto similarity were withheld as an external test set. Feature importance analysis highlighted transporter proteins and cytochrome P450 enzymes, with enriched pathways primarily linked to absorption, distribution, metabolism, and excretion (ADME) processes. Moreover, an interaction count for liver-expressed targets differentiates steatosis-positive from -negative compounds, suggesting a simple metric that could be worth exploring for other complex toxicity endpoints. To illustrate practical utility, the framework was applied to compounds with sparse FAERS evidence for steatosis. Overall, integrating chemical and biological features within a consensus machine-learning framework enables a robust steatosis prediction while offering interpretable patterns that help contextualize biological processes associated with steatosis.
Authors
- Gerhard Franz Ecker (ORCID: https://orcid.org/0000-0003-4209-6883)
- Florian Oehler
- Palle Steen Helmke (ORCID: https://orcid.org/0009-0003-8393-9047)
Institutions
- University of Vienna (AT)
Publication Details
- Journal
- Discover Pharmaceutical Sciences
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1007/s44395-026-00057-1
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- European Commission
- National Centre for the Replacement Refinement and Reduction of Animals in Research