Enzymatic Reaction Feasibility Classification Using Machine Learning Methods

Abstract With the advancement of computer-aided retrobiosynthesis, numerous biosynthetic pathways have been predicted, exceeding the capacity of experimental validation. Effective classifiers are needed to identify feasible reactions. In this study, we collected 75,864 feasible reactions from biocatalysis databases and generated an equal number of infeasible reactions based on reaction rules. Following atom mapping focused on reaction centers and systematic reaction preprocessing, three datasets for training were constructed: the “stereo” dataset, which retained reaction stereochemical information; the “non-stereo” dataset, which was a stereochemistry-agnostic version of the “stereo” dataset; and the “mixed” dataset, which comprised both. We established a total of 22 individual enzymatic reaction feasibility classification models, which include: eXtreme Gradient Boosting (XGBoost) and Deep Neural Network (DNN) models utilizing Reaction Fingerprints (RXNFP), Differential Reaction Fingerprint (DRFP), and our constructed Combined ECFP4 Reaction Fingerprints (c_ECFP4) for reaction representation, and Transformer models and fine-tuned ChemBERTa-77M-MLM (ChemMLM) models using reaction SMILES strings as the direct input. The results indicate that models utilizing the c_ECFP4 representation achieved the highest predictive performance, which effectively captured underlying enzymatic reaction mechanisms. Among them, Model 1A-M (based on XGBoost and “mixed” dataset) was identified as the optimal individual model, achieving Matthews Correlation Coefficient (MCC) values of 0.865 and 0.853 and Area Under Curve (AUC) values of 0.981 and 0.980 on the “stereo” and “non-stereo” test sets, respectively. Furthermore, a consensus model enzymatic reaction feasibility classification (ERFC) integrating four reaction representations further improved predictive performance, achieving MCC values of 0.894 and 0.886 and an AUC of 0.986 on both test sets. Moreover, both models (Model 1A-M and ERFC) successfully validated a five-step biosynthetic pathway, demonstrating higher prediction accuracy than the previously reported DeepRFC and DORA-XGB models in identifying feasible reactions. All data, the individual Model 1A-M, and the consensus model ERFC are openly available, offering a reliable and flexible framework for predicting enzymatic reaction feasibility in the presence or absence of stereochemical information.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-09-29
DOI
https://doi.org/10.1021/acs.jcim.6c02294
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Enzymatic Reaction Feasibility Classification Using Machine Learning Methods

Yushan Zhu, Aixia Yan, Igor V. Tetko, Yekai Shen et al.
Journal of Chemical Information and Modeling
Machine Learning in Materials Science
article

Enzymatic Reaction Feasibility Classification Using Machine Learning Methods

Yushan Zhu, Aixia Yan, Igor V. Tetko, Yekai Shen, Hongyan Yin, Xin Wang
article en

Abstract

Abstract With the advancement of computer-aided retrobiosynthesis, numerous biosynthetic pathways have been predicted, exceeding the capacity of experimental validation. Effective classifiers are needed to identify feasible reactions. In this study, we collected 75,864 feasible reactions from biocatalysis databases and generated an equal number of infeasible reactions based on reaction rules. Following atom mapping focused on reaction centers and systematic reaction preprocessing, three datasets for training were constructed: the “stereo” dataset, which retained reaction stereochemical information; the “non-stereo” dataset, which was a stereochemistry-agnostic version of the “stereo” dataset; and the “mixed” dataset, which comprised both. We established a total of 22 individual enzymatic reaction feasibility classification models, which include: eXtreme Gradient Boosting (XGBoost) and Deep Neural Network (DNN) models utilizing Reaction Fingerprints (RXNFP), Differential Reaction Fingerprint (DRFP), and our constructed Combined ECFP4 Reaction Fingerprints (c_ECFP4) for reaction representation, and Transformer models and fine-tuned ChemBERTa-77M-MLM (ChemMLM) models using reaction SMILES strings as the direct input. The results indicate that models utilizing the c_ECFP4 representation achieved the highest predictive performance, which effectively captured underlying enzymatic reaction mechanisms. Among them, Model 1A-M (based on XGBoost and “mixed” dataset) was identified as the optimal individual model, achieving Matthews Correlation Coefficient (MCC) values of 0.865 and 0.853 and Area Under Curve (AUC) values of 0.981 and 0.980 on the “stereo” and “non-stereo” test sets, respectively. Furthermore, a consensus model enzymatic reaction feasibility classification (ERFC) integrating four reaction representations further improved predictive performance, achieving MCC values of 0.894 and 0.886 and an AUC of 0.986 on both test sets. Moreover, both models (Model 1A-M and ERFC) successfully validated a five-step biosynthetic pathway, demonstrating higher prediction accuracy than the previously reported DeepRFC and DORA-XGB models in identifying feasible reactions. All data, the individual Model 1A-M, and the consensus model ERFC are openly available, offering a reliable and flexible framework for predicting enzymatic reaction feasibility in the presence or absence of stereochemical information.

Journal of Chemical Information and Modeling
Center for Environmental Health (US), Helmholtz Zentrum München (DE), Beijing University of Chemical Technology (CN)
Openalex Percentile: Top 26%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.