Research EnzyPred: a general model for predicting enzyme–substrate interactions using deep learning, parallel LoRA, and, natural language processing-inspired approaches
Enzyme–substrate interaction (ESI) prediction is a critical challenge in biocatalysis, with applications in metabolic engineering and drug design. This prediction is valuable because it allows avoiding experimental characterization, a process that can be complex and expensive. In this work, we present EnzyPred, an ESI prediction tool based on Natural Language Processing (NLP), Deep Learning (DL), fine-tuning, Low Range Parallel Adapter (LoRA-P), and Inference-time Prioritized Fine-tuning (PFT) techniques, using representations generated through Evolutionary Scale Modeling (ESM) for enzymes and Extended Connectivity Fingerprinting (ECFP) for substrates. This approach allows training the model with evolutionary, structural, and chemical features, improving performance in identifying interactions. The model was trained on a dataset of 32,833 positive enzyme–substrate pairs, complemented by 32,833 negative pairs generated using soft clustering, molecular descriptors, and Tanimoto similarity. EnzyPred achieved an average accuracy of 86.97%, proving to be competitive with previously presented tools, standing out for its generalization capacity, computational efficiency, and adaptability. Two case studies were used to validate the tool experimentally. In the first, the enzyme that catalyzes the conversion of 4'-O-methylnorbelladine to N-desmethylnarwedine in Amaryllidaceae was sought. For this, the substrate was fixed and tested with predicted enzymes from different transcriptomes reported in the literature for this family. The predicted ones were tested in the laboratory, identifying the enzyme, which contributed to the knowledge of the synthesis pathway of galantamine, a drug used in the treatment of Alzheimer's. In the second case, we searched for compounds that may interact with the enzyme acetylcholinesterase (AChE), for which the enzyme was fixed and we used the ZINC22 database as a search space. 747 compounds with activity were predicted, of which nine were prioritized through clustering for experimental testing, and all presented activity. These results confirm the relevance of this approach to predict molecular interactions, opening new possibilities for the design of enzymatic and synthetic pathways. Scientific contribution This study introduces EnzyPred, an adaptive deep learning framework for enzyme–substrate interaction prediction that jointly integrates protein language model, molecular fingerprints, and a novel activity-free negative sampling strategy to address data imbalance and false negatives. Unlike previous ESI models, EnzyPred incorporates low-rank adaptation and Inference-time Prioritized Fine-tuning, enabling rapid and targeted model specialization for specific enzyme–substrate pairs without retraining from scratch. The framework is validated across enzyme identification and virtual screening tasks, with experimental confirmation, demonstrating its potential to accelerate biocatalysis discovery and pathway elucidation beyond existing methods.
Authors
- Óscar Alberto Álvarez Solano
- Edison Osorio (ORCID: https://orcid.org/0000-0001-7636-8168)
- Luis Fernando Salas Nuñez (ORCID: https://orcid.org/0000-0002-7335-3523)
- Álvaro Barrera-Ocampo
- María F. Villegas-Torres (ORCID: https://orcid.org/0000-0001-7067-4113)
- Natalie Cortés
- Paola A. Caicedo (ORCID: https://orcid.org/0000-0002-4517-427X)
- Andrés Fernando González Barrios González Barrios
Institutions
- Universidad de Los Andes (CO)
- Universidad Católica Luis Amigó (CO)
- Antioquia Institute of Technology (CO)
- Universidad de Ibagué (CO)
- Icesi University (CO)
Publication Details
- Journal
- Journal of Cheminformatics
- Published
- 2026-08-25
- DOI
- https://doi.org/10.1186/s13321-026-01282-7
- Primary Topic
- Chemical synthesis and alkaloids
- Type
- article
- Field-Weighted Citation Impact
- 0.00