Machine learning-based multiclass classification of cognitive decline using ACE-III assessment data: a comparative study with SHAP explainability

Abstract Dementia and its earlier stages mild cognitive impairment (MCI) and subjective cognitive decline (SCD) are increasingly common, yet automated tools that distinguish all four cognitive states simultaneously remain rare. Most published models address a simpler binary problem; the four-class case, which is what clinicians actually face, has attracted far less attention. We developed and benchmarked twelve machine learning classifiers spanning linear, kernel-based, tree-based, boosting, probabilistic, distance-based, and ensemble meta-learning paradigms for four-class cognitive stratification (Cognitively Normal, SCD, MCI, Dementia) using ACE-III assessment-derived clinical and demographic data ( n = 140, 28 features). Class imbalance was mitigated through pipeline-integrated Synthetic Minority Over-sampling Technique (SMOTE), with model evaluation conducted via stratified five-fold cross-validation. All performance metrics (accuracy, macro F1-score, ROC-AUC, balanced accuracy, Cohen’s Kappa, Matthews Correlation Coefficient) were computed across folds and reported as mean ± standard deviation. SHAP (SHapley Additive exPlanations) analysis was employed to quantify global and local feature contributions across all model types. LightGBM achieved the highest macro F1-score (0.723 ± 0.114), ROC-AUC (0.944 ± 0.030), and Cohen’s Kappa (0.729 ± 0.138) among the twelve classifiers, with a balanced accuracy of 0.716 ± 0.121. Voting Classifier attained a near-identical balanced accuracy (0.716 ± 0.068) and the highest single-model MCC (0.734 ± 0.091), closely matching LightGBM’s own MCC (0.733 ± 0.139), and ranked second by macro F1-score (0.714 ± 0.058). k-NN and SVM ranked third and fourth by macro F1-score (0.640 ± 0.136 and 0.631 ± 0.095, respectively), while Gaussian Naïve Bayes performed least favourably (F1: 0.480 ± 0.143). A paired t-test found LightGBM’s macro F1-score advantage statistically significant against one of the eleven remaining classifiers, Logistic Regression (t = 6.783, p = 0.0025, Holm-adjusted p = 0.0271), while differences from the other ten classifiers — including its closest competitors, Voting Classifier, k-NN, and Extra Trees — did not reach significance at this sample size. SHAP analysis identified age, health-condition status, educational attainment, depression, blood pressure, and family history of MCI/Alzheimer’s disease as the most consistently influential predictors across model types. LightGBM classifier trained on Egyptian Arabic ACE-III data classifies patients into four cognitive categories with accuracy that is promising for pre-screening or triage role, though external validation on independent cohorts is needed before any claim of clinical utility can be made, and SHAP makes its reasoning transparent. The framework requires no specialist equipment and runs on data already collected in routine Egyptian encounters, making primary care deployment in Egypt and the MENA region practical pending that validation.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-71441-1
Primary Topic
Dementia and Cognitive Impairment Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Machine learning-based multiclass classification of cognitive decline using ACE-III assessment data: a comparative study with SHAP explainability

Gamal Attiya, Salah Eldin S. E. Abdulrahman, Heba M. Tawfik, Nabil A. Ismail et al.
Scientific Reports
Dementia and Cognitive Impairment Research
article

Machine learning-based multiclass classification of cognitive decline using ACE-III assessment data: a comparative study with SHAP explainability

Gamal Attiya, Salah Eldin S. E. Abdulrahman, Heba M. Tawfik, Nabil A. Ismail, Gamal A. El-Sheikh, Mohamed Sherif N. Abd Rabou
article en

Abstract

Abstract Dementia and its earlier stages mild cognitive impairment (MCI) and subjective cognitive decline (SCD) are increasingly common, yet automated tools that distinguish all four cognitive states simultaneously remain rare. Most published models address a simpler binary problem; the four-class case, which is what clinicians actually face, has attracted far less attention. We developed and benchmarked twelve machine learning classifiers spanning linear, kernel-based, tree-based, boosting, probabilistic, distance-based, and ensemble meta-learning paradigms for four-class cognitive stratification (Cognitively Normal, SCD, MCI, Dementia) using ACE-III assessment-derived clinical and demographic data ( n = 140, 28 features). Class imbalance was mitigated through pipeline-integrated Synthetic Minority Over-sampling Technique (SMOTE), with model evaluation conducted via stratified five-fold cross-validation. All performance metrics (accuracy, macro F1-score, ROC-AUC, balanced accuracy, Cohen’s Kappa, Matthews Correlation Coefficient) were computed across folds and reported as mean ± standard deviation. SHAP (SHapley Additive exPlanations) analysis was employed to quantify global and local feature contributions across all model types. LightGBM achieved the highest macro F1-score (0.723 ± 0.114), ROC-AUC (0.944 ± 0.030), and Cohen’s Kappa (0.729 ± 0.138) among the twelve classifiers, with a balanced accuracy of 0.716 ± 0.121. Voting Classifier attained a near-identical balanced accuracy (0.716 ± 0.068) and the highest single-model MCC (0.734 ± 0.091), closely matching LightGBM’s own MCC (0.733 ± 0.139), and ranked second by macro F1-score (0.714 ± 0.058). k-NN and SVM ranked third and fourth by macro F1-score (0.640 ± 0.136 and 0.631 ± 0.095, respectively), while Gaussian Naïve Bayes performed least favourably (F1: 0.480 ± 0.143). A paired t-test found LightGBM’s macro F1-score advantage statistically significant against one of the eleven remaining classifiers, Logistic Regression (t = 6.783, p = 0.0025, Holm-adjusted p = 0.0271), while differences from the other ten classifiers — including its closest competitors, Voting Classifier, k-NN, and Extra Trees — did not reach significance at this sample size. SHAP analysis identified age, health-condition status, educational attainment, depression, blood pressure, and family history of MCI/Alzheimer’s disease as the most consistently influential predictors across model types. LightGBM classifier trained on Egyptian Arabic ACE-III data classifies patients into four cognitive categories with accuracy that is promising for pre-screening or triage role, though external validation on independent cohorts is needed before any claim of clinical utility can be made, and SHAP makes its reasoning transparent. The framework requires no specialist equipment and runs on data already collected in routine Egyptian encounters, making primary care deployment in Egypt and the MENA region practical pending that validation.

Scientific ReportsVol. 16(1)
Openalex Percentile: Top 12%
Dementia and Cognitive Impairment Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.