Research EnzyPred: a general model for predicting enzyme–substrate interactions using deep learning, parallel LoRA, and, natural language processing-inspired approaches

Enzyme–substrate interaction (ESI) prediction is a critical challenge in biocatalysis, with applications in metabolic engineering and drug design. This prediction is valuable because it allows avoiding experimental characterization, a process that can be complex and expensive. In this work, we present EnzyPred, an ESI prediction tool based on Natural Language Processing (NLP), Deep Learning (DL), fine-tuning, Low Range Parallel Adapter (LoRA-P), and Inference-time Prioritized Fine-tuning (PFT) techniques, using representations generated through Evolutionary Scale Modeling (ESM) for enzymes and Extended Connectivity Fingerprinting (ECFP) for substrates. This approach allows training the model with evolutionary, structural, and chemical features, improving performance in identifying interactions. The model was trained on a dataset of 32,833 positive enzyme–substrate pairs, complemented by 32,833 negative pairs generated using soft clustering, molecular descriptors, and Tanimoto similarity. EnzyPred achieved an average accuracy of 86.97%, proving to be competitive with previously presented tools, standing out for its generalization capacity, computational efficiency, and adaptability. Two case studies were used to validate the tool experimentally. In the first, the enzyme that catalyzes the conversion of 4'-O-methylnorbelladine to N-desmethylnarwedine in Amaryllidaceae was sought. For this, the substrate was fixed and tested with predicted enzymes from different transcriptomes reported in the literature for this family. The predicted ones were tested in the laboratory, identifying the enzyme, which contributed to the knowledge of the synthesis pathway of galantamine, a drug used in the treatment of Alzheimer's. In the second case, we searched for compounds that may interact with the enzyme acetylcholinesterase (AChE), for which the enzyme was fixed and we used the ZINC22 database as a search space. 747 compounds with activity were predicted, of which nine were prioritized through clustering for experimental testing, and all presented activity. These results confirm the relevance of this approach to predict molecular interactions, opening new possibilities for the design of enzymatic and synthetic pathways. Scientific contribution This study introduces EnzyPred, an adaptive deep learning framework for enzyme–substrate interaction prediction that jointly integrates protein language model, molecular fingerprints, and a novel activity-free negative sampling strategy to address data imbalance and false negatives. Unlike previous ESI models, EnzyPred incorporates low-rank adaptation and Inference-time Prioritized Fine-tuning, enabling rapid and targeted model specialization for specific enzyme–substrate pairs without retraining from scratch. The framework is validated across enzyme identification and virtual screening tasks, with experimental confirmation, demonstrating its potential to accelerate biocatalysis discovery and pathway elucidation beyond existing methods.

Authors

Institutions

Publication Details

Journal
Journal of Cheminformatics
Published
2026-08-25
DOI
https://doi.org/10.1186/s13321-026-01282-7
Primary Topic
Chemical synthesis and alkaloids
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Research EnzyPred: a general model for predicting enzyme–substrate interactions using deep learning, parallel LoRA, and, natural language processing-inspired approaches

Óscar Alberto Álvarez Solano, Edison Osorio, Luis Fernando Salas Nuñez, Álvaro Barrera-Ocampo et al.
Journal of Cheminformatics
Chemical synthesis and alkaloids
article

Research EnzyPred: a general model for predicting enzyme–substrate interactions using deep learning, parallel LoRA, and, natural language processing-inspired approaches

Óscar Alberto Álvarez Solano, Edison Osorio, Luis Fernando Salas Nuñez, Álvaro Barrera-Ocampo, María F. Villegas-Torres, Natalie Cortés, Paola A. Caicedo, Andrés Fernando González Barrios González Barrios
article en

Abstract

Enzyme–substrate interaction (ESI) prediction is a critical challenge in biocatalysis, with applications in metabolic engineering and drug design. This prediction is valuable because it allows avoiding experimental characterization, a process that can be complex and expensive. In this work, we present EnzyPred, an ESI prediction tool based on Natural Language Processing (NLP), Deep Learning (DL), fine-tuning, Low Range Parallel Adapter (LoRA-P), and Inference-time Prioritized Fine-tuning (PFT) techniques, using representations generated through Evolutionary Scale Modeling (ESM) for enzymes and Extended Connectivity Fingerprinting (ECFP) for substrates. This approach allows training the model with evolutionary, structural, and chemical features, improving performance in identifying interactions. The model was trained on a dataset of 32,833 positive enzyme–substrate pairs, complemented by 32,833 negative pairs generated using soft clustering, molecular descriptors, and Tanimoto similarity. EnzyPred achieved an average accuracy of 86.97%, proving to be competitive with previously presented tools, standing out for its generalization capacity, computational efficiency, and adaptability. Two case studies were used to validate the tool experimentally. In the first, the enzyme that catalyzes the conversion of 4'-O-methylnorbelladine to N-desmethylnarwedine in Amaryllidaceae was sought. For this, the substrate was fixed and tested with predicted enzymes from different transcriptomes reported in the literature for this family. The predicted ones were tested in the laboratory, identifying the enzyme, which contributed to the knowledge of the synthesis pathway of galantamine, a drug used in the treatment of Alzheimer's. In the second case, we searched for compounds that may interact with the enzyme acetylcholinesterase (AChE), for which the enzyme was fixed and we used the ZINC22 database as a search space. 747 compounds with activity were predicted, of which nine were prioritized through clustering for experimental testing, and all presented activity. These results confirm the relevance of this approach to predict molecular interactions, opening new possibilities for the design of enzymatic and synthetic pathways. Scientific contribution This study introduces EnzyPred, an adaptive deep learning framework for enzyme–substrate interaction prediction that jointly integrates protein language model, molecular fingerprints, and a novel activity-free negative sampling strategy to address data imbalance and false negatives. Unlike previous ESI models, EnzyPred incorporates low-rank adaptation and Inference-time Prioritized Fine-tuning, enabling rapid and targeted model specialization for specific enzyme–substrate pairs without retraining from scratch. The framework is validated across enzyme identification and virtual screening tasks, with experimental confirmation, demonstrating its potential to accelerate biocatalysis discovery and pathway elucidation beyond existing methods.

Journal of Cheminformatics
Universidad de Los Andes (CO), Universidad Católica Luis Amigó (CO), Antioquia Institute of Technology (CO), Universidad de Ibagué (CO), Icesi University (CO)
Quality Education
Openalex Percentile: Top 18%
Chemical synthesis and alkaloids
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.