Validating and Correcting Graph Neural Network Attention for Drug-Induced Liver Injury Prediction
Abstract Attention-based graph neural networks (GNNs) are increasingly used for drug-induced liver injury (DILI) and other toxicity prediction tasks on the claimed strength of built-in interpretability, but this claim is almost always supported by a handful of hand-selected examples rather than being tested systematically. Using a curated, scaffold-split gold-standard DILI data set (1,079 compounds), we show that a graph attention network’s attention weights do not, in general, concentrate on seven literature-derived hepatotoxicity structural alerts across the full data set (pooled enrichment ratio of 0.35) and that the network itself does not outperform simple descriptor-based baselines under rigorous repeated-split evaluation (mean AUROC 0.625-0.649 vs. 0.714 for random forest, p = 0.011). We then show both gaps can be addressed: a chemistry-informed multitask extension, trained with an auxiliary structural-alert-recognition loss, improves attention alignment for two alerts by more than 5-fold (both Bonferroni/FDR-corrected p < 10–14), confirmed across five independent scaffold splits, without a robust cost to classification performance. All curated data, trained models, and analysis code are released to support reproducibility and further benchmarking of both predictive and interpretability claims for molecular GNNs.
Authors
- Youssef M. Hassan (ORCID: https://orcid.org/0009-0005-3615-4137)
- Ibrahim Hassan Ali (ORCID: https://orcid.org/0000-0002-8378-6103)
- Hala El-Tantawi
- Mohamed S. Attia
Institutions
- Ain Shams University (EG)
- Imam Mohammad ibn Saud Islamic University (SA)
- Theodor Bilharz Research Institute (EG)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1021/acs.jcim.6c02332
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00