CatRange enables robust prediction of enzyme variant kinetic regimes

Abstract Predicting enzyme kinetics directly from sequence remains a central challenge in computational biology, particularly in resolving the effects of mutations at catalytically essential residues. Existing models frequently overlook the functional consequences of such perturbations, defaulting to wild-type predictions even in cases of substantial activity loss, thereby limiting their reliability for enzyme design and mechanistic inference. Here, we introduce CatRange, a machine learning framework trained on CatLog-27k, a human-in-the-loop, AI-agent trustworthy dataset of 27,176 in vitro enzyme–substrate kinetic records created by a systematic audit and correction of BRENDA and SABIO-RK. All mutant entries are manually reconciled against 2,158 source articles. CatRange reframes kinetic prediction from exact numerical regression into classification over log10-spaced bins for catalytic turnover (kcat) and substrate affinity (KM), matching the order-of-magnitude scale at which experimental enzyme kinetic measurements are commonly interpreted. This biologically grounded formulation mitigates assay-level variability while preserving distinctions among functional catalytic and binding states. Using joint enzyme–substrate representations and gradient-boosted classifiers, CatRange predicts kinetic ranges for wild-type and mutant enzymes across standard held-out, out-of-distribution, and few-shot mutation settings. The model shows robust order-of-magnitude recovery with class-balanced discrimination and captures mutation-induced movement across kinetic regimes, including losses associated with perturbation of annotated catalytic residues. CatRange detects non-enzyme sequence inputs and emphasizes rigorous data curation, transparent training data dissemination (CatLog), biochemically informed task formulation, and balanced evaluation metrics. These position CatRange as an interpretable, mutation-sensitive framework with utility in enzyme engineering and kinetic metabolic modeling.

Authors

Institutions

Publication Details

Journal
PNAS Nexus
Published
2026-09-16
DOI
https://doi.org/10.1093/pnasnexus/pgag309
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CatRange enables robust prediction of enzyme variant kinetic regimes

Karuna Anna Sajeevan, Rahil Salehi, Sakib Ferdous, Supantha Dey et al.
PNAS Nexus
Machine Learning in Materials Science
article

CatRange enables robust prediction of enzyme variant kinetic regimes

Karuna Anna Sajeevan, Rahil Salehi, Sakib Ferdous, Supantha Dey, Abraham Osinuga, Ankur Mali, Mohammed Sakib Noor, Rajib Saha, Nabia Shahreen, Ratul Chowdhury, Randy Aryee, Brisa Calderon-Lopez, Shashank Koneru, Niaz B. Chowdhury, Laura Mariana Santos-Correa, B Arunraj
article en

Abstract

Abstract Predicting enzyme kinetics directly from sequence remains a central challenge in computational biology, particularly in resolving the effects of mutations at catalytically essential residues. Existing models frequently overlook the functional consequences of such perturbations, defaulting to wild-type predictions even in cases of substantial activity loss, thereby limiting their reliability for enzyme design and mechanistic inference. Here, we introduce CatRange, a machine learning framework trained on CatLog-27k, a human-in-the-loop, AI-agent trustworthy dataset of 27,176 in vitro enzyme–substrate kinetic records created by a systematic audit and correction of BRENDA and SABIO-RK. All mutant entries are manually reconciled against 2,158 source articles. CatRange reframes kinetic prediction from exact numerical regression into classification over log10-spaced bins for catalytic turnover (kcat) and substrate affinity (KM), matching the order-of-magnitude scale at which experimental enzyme kinetic measurements are commonly interpreted. This biologically grounded formulation mitigates assay-level variability while preserving distinctions among functional catalytic and binding states. Using joint enzyme–substrate representations and gradient-boosted classifiers, CatRange predicts kinetic ranges for wild-type and mutant enzymes across standard held-out, out-of-distribution, and few-shot mutation settings. The model shows robust order-of-magnitude recovery with class-balanced discrimination and captures mutation-induced movement across kinetic regimes, including losses associated with perturbation of annotated catalytic residues. CatRange detects non-enzyme sequence inputs and emphasizes rigorous data curation, transparent training data dissemination (CatLog), biochemically informed task formulation, and balanced evaluation metrics. These position CatRange as an interpretable, mutation-sensitive framework with utility in enzyme engineering and kinetic metabolic modeling.

PNAS Nexus
University of Nebraska–Lincoln (US), Iowa State University (US), University of South Florida (US), Ramakrishna Mission Vidyamandira (IN), Ames National Laboratory (US)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 24%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.