Classification of pre-metabolic syndrome and metabolic syndrome using machine learning and routine clinical variables in Thai adults

Background Metabolic syndrome (MetS) and pre-metabolic syndrome (pre-MetS) are increasingly recognized as important stages of metabolic dysfunction in Asian populations. Conventional diagnostic approaches rely on fixed threshold criteria, whereas metabolic dysfunction occurs along a continuous spectrum. Machine learning (ML) can integrate multidimensional clinical data and model nonlinear relationships among clinical variables, potentially providing complementary information for distinguishing metabolic states. This study evaluated ten ML algorithms for distinguishing pre-MetS from MetS in Thai adults and identified features contributing most strongly to model-based classification. Methods In this single-center retrospective cross-sectional study, 657 Thai adults were classified as pre-MetS ( n = 359) or MetS ( n = 298). Pre-MetS was operationally defined as the presence of at least one IDF-defined metabolic abnormality without fulfilling the complete International Diabetes Federation (IDF) criteria for MetS. Twenty-four clinical and biochemical features selected using the Boruta algorithm were evaluated across ten supervised ML algorithms. Models were trained using an 80% training dataset with repeated 10-fold cross-validation (10 repetitions) for hyperparameter tuning and subsequently evaluated using an independent 20% hold-out test dataset ( n = 132). All preprocessing steps, including z-score normalization and feature selection, were performed exclusively within the training dataset to prevent data leakage. Results Among the ten evaluated algorithms, ensemble methods achieved the highest classification performance. XGBoost demonstrated the highest discrimination (AUC = 0.986, 95% CI: 0.970–0.997), highest accuracy (0.932, 95% CI: 0.885–0.969), and lowest Brier score (0.050, 95% CI: 0.026–0.080), followed by the neural network (AUC = 0.948, 95% CI: 0.906–0.983) and random forest (AUC = 0.942, 95% CI: 0.905–0.974). SHAP analysis identified systolic blood pressure, waist circumference, fasting plasma glucose, sex, and TyG-WHtR as the features contributing most strongly to XGBoost classification. Sensitivity analyses demonstrated progressively reduced discrimination after excluding predictors that overlapped with MetS diagnostic components and related derived indices, although substantial discrimination was retained after removal of these overlapping features. Conclusions Among the ten evaluated algorithms, classification performance varied substantially across models, with XGBoost achieving the highest performance and support vector machine models showing the lowest discrimination. These findings indicate that algorithm choice may influence classification performance in metabolic classification tasks. However, the observed performance should be interpreted cautiously because several influential predictors overlap with, or are derived from, the MetS diagnostic criteria, and the study lacked external validation. The findings are therefore exploratory and require validation in multicenter cohorts using predictor sets that minimize overlap with the diagnostic definition before clinical implementation can be considered.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-28
DOI
https://doi.org/10.1371/journal.pone.0359283
Primary Topic
Artificial Intelligence in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Classification of pre-metabolic syndrome and metabolic syndrome using machine learning and routine clinical variables in Thai adults

Somlak Chuengsamarn, Metha Yaikwawong, Khanittha Kamdee
PLoS ONE
Artificial Intelligence in Healthcare
article

Classification of pre-metabolic syndrome and metabolic syndrome using machine learning and routine clinical variables in Thai adults

Somlak Chuengsamarn, Metha Yaikwawong, Khanittha Kamdee
article en

Abstract

Background Metabolic syndrome (MetS) and pre-metabolic syndrome (pre-MetS) are increasingly recognized as important stages of metabolic dysfunction in Asian populations. Conventional diagnostic approaches rely on fixed threshold criteria, whereas metabolic dysfunction occurs along a continuous spectrum. Machine learning (ML) can integrate multidimensional clinical data and model nonlinear relationships among clinical variables, potentially providing complementary information for distinguishing metabolic states. This study evaluated ten ML algorithms for distinguishing pre-MetS from MetS in Thai adults and identified features contributing most strongly to model-based classification. Methods In this single-center retrospective cross-sectional study, 657 Thai adults were classified as pre-MetS ( n = 359) or MetS ( n = 298). Pre-MetS was operationally defined as the presence of at least one IDF-defined metabolic abnormality without fulfilling the complete International Diabetes Federation (IDF) criteria for MetS. Twenty-four clinical and biochemical features selected using the Boruta algorithm were evaluated across ten supervised ML algorithms. Models were trained using an 80% training dataset with repeated 10-fold cross-validation (10 repetitions) for hyperparameter tuning and subsequently evaluated using an independent 20% hold-out test dataset ( n = 132). All preprocessing steps, including z-score normalization and feature selection, were performed exclusively within the training dataset to prevent data leakage. Results Among the ten evaluated algorithms, ensemble methods achieved the highest classification performance. XGBoost demonstrated the highest discrimination (AUC = 0.986, 95% CI: 0.970–0.997), highest accuracy (0.932, 95% CI: 0.885–0.969), and lowest Brier score (0.050, 95% CI: 0.026–0.080), followed by the neural network (AUC = 0.948, 95% CI: 0.906–0.983) and random forest (AUC = 0.942, 95% CI: 0.905–0.974). SHAP analysis identified systolic blood pressure, waist circumference, fasting plasma glucose, sex, and TyG-WHtR as the features contributing most strongly to XGBoost classification. Sensitivity analyses demonstrated progressively reduced discrimination after excluding predictors that overlapped with MetS diagnostic components and related derived indices, although substantial discrimination was retained after removal of these overlapping features. Conclusions Among the ten evaluated algorithms, classification performance varied substantially across models, with XGBoost achieving the highest performance and support vector machine models showing the lowest discrimination. These findings indicate that algorithm choice may influence classification performance in metabolic classification tasks. However, the observed performance should be interpreted cautiously because several influential predictors overlap with, or are derived from, the MetS diagnostic criteria, and the study lacked external validation. The findings are therefore exploratory and require validation in multicenter cohorts using predictor sets that minimize overlap with the diagnostic definition before clinical implementation can be considered.

PLoS ONEVol. 21(9)
Mahidol University (TH), Srinakharinwirot University (TH)
Peace, Justice and strong institutions
Openalex Percentile: Top 3%
Artificial Intelligence in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.