Language of Toxicity: An eXplainable Artificial Intelligence Approach
Abstract Toxicity prediction in small molecules represents a fundamental challenge in drug development and chemical safety assessment. Traditional approaches heavily rely on predefined molecular descriptors or fingerprints, potentially limiting the ability to capture complex and nonlinear structure–activity relationships. Here, we present a descriptor-free, language-inspired framework that can be applied to different toxicity prediction tasks within a unified architecture. The model proposed combines a multiscale Convolutional Neural Network (CNN) layer to capture chemical patterns at different scales and a Gated Recurrent Unit (GRU) layer to capture the sequential nature of these patterns. This architecture also exploits an attention mechanism that computes attention weights across the sequence, enabling the model to focus on the most relevant molecular substructures for toxicity prediction. Toxic and nontoxic chemicals, represented by canonical SMILES, are investigated as the words of two languages which have to be discriminated; using eight different end points, the model provided an accurate description of toxicity patterns, with an average Area under the ROC curve (AUC) of 0.83 (min: 0.70, max: 0.94) under repeated cross-validation. The models were trained on relatively small data sets (∼1000 samples) and often strongly imbalanced, two important challenges that highlight the opportunities for future improvement; moreover, the proposed attention-based framework offers a representation of the molecular regions influencing model predictions, providing a basis for future investigations into toxicity-related structural patterns and potentially supporting hypothesis generation in drug design or drug repurposing applications.
Authors
- Fabrizio Mastrolorito (ORCID: https://orcid.org/0000-0001-6753-8997)
- Nicola Amoroso (ORCID: https://orcid.org/0000-0003-0211-0783)
- Alfonso Monaco (ORCID: https://orcid.org/0000-0002-5968-8642)
- R. Bellotti (ORCID: https://orcid.org/0000-0003-3198-2708)
- Orazio Nicolotti (ORCID: https://orcid.org/0000-0001-6533-5539)
- Nicola Gambacorta (ORCID: https://orcid.org/0000-0003-1965-1519)
- Ester Pantaleo (ORCID: https://orcid.org/0000-0001-8407-9032)
- Fulvio Ciriaco (ORCID: https://orcid.org/0000-0002-0695-6607)
- Francesca Cutropia (ORCID: https://orcid.org/0009-0008-9454-4282)
- Angelica Orfino
Institutions
- Istituto Nazionale di Fisica Nucleare, Sezione di Bari (IT)
- Merck Serono S.A. (Switzerland) (CH)
- University of Bari Aldo Moro (IT)
- Merck Serono S.p.A. (Italy) (IT)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1021/acs.jcim.5c03214
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00