Deep Learning in Enzyme Function Prediction and Novel Enzyme Discovery

Abstract The rapid growth of genomic and metagenomic data has created a large gap between protein sequence discovery and enzyme functional characterization. Deep learning has become an important tool for enzyme-related prediction tasks, including EC and GO annotation, substrate specificity, catalytic-residue identification, kinetic-parameter estimation, thermostability, pH optima, and candidate discovery. This review summarizes recent progress with an emphasis on benchmark design rather than reported scores alone. We compare representative methods by input modality, training data, split strategy, leakage control, evaluation metric, and validation evidence. Across tasks, apparent performance gains can arise from dataset scale, annotation quality, homolog overlap, negative sampling, or endpoint definition rather than architecture alone. Current models are valuable for annotation and candidate prioritization, but robust discovery of remote homologs, rare functions, or new catalytic activities still requires low-identity or out-of-distribution evaluation and experimental validation.

Authors

Institutions

Publication Details

Journal
ACS Synthetic Biology
Published
2026-09-04
DOI
https://doi.org/10.1021/acssynbio.6c00633
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Deep Learning in Enzyme Function Prediction and Novel Enzyme Discovery

Tatsuhisa Tsuboi, Rongsheng Gao, Youmeng Liu, Zhen Chen et al.
ACS Synthetic Biology
Machine Learning in Bioinformatics
article

Deep Learning in Enzyme Function Prediction and Novel Enzyme Discovery

Tatsuhisa Tsuboi, Rongsheng Gao, Youmeng Liu, Zhen Chen, Kunyu Liu, Nan Qin, Chengye Duan
article en

Abstract

Abstract The rapid growth of genomic and metagenomic data has created a large gap between protein sequence discovery and enzyme functional characterization. Deep learning has become an important tool for enzyme-related prediction tasks, including EC and GO annotation, substrate specificity, catalytic-residue identification, kinetic-parameter estimation, thermostability, pH optima, and candidate discovery. This review summarizes recent progress with an emphasis on benchmark design rather than reported scores alone. We compare representative methods by input modality, training data, split strategy, leakage control, evaluation metric, and validation evidence. Across tasks, apparent performance gains can arise from dataset scale, annotation quality, homolog overlap, negative sampling, or endpoint definition rather than architecture alone. Current models are valuable for annotation and candidate prioritization, but robust discovery of remote homologs, rare functions, or new catalytic activities still requires low-identity or out-of-distribution evaluation and experimental validation.

ACS Synthetic Biology
University Town of Shenzhen (CN), Tsinghua–Berkeley Shenzhen Institute (CN), Tsinghua University (CN)
National Natural Science Foundation of China, Tsinghua University, Beijing Municipal Science and Technology Commission, National Key Research and Development Program of China
Openalex Percentile: Top 17%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Deep Learning in Enzyme Function Prediction and Novel Enzyme Discovery — Tatsuhisa Tsuboi, Rongsheng Gao, et al. · ACS Synthetic Biology (2026) | TGRS Research Map | TGRS