Predicting time-to-progression in Alzheimer’s disease using artificial intelligence for survival analysis models: a systematic review and meta-analysis
Machine learning (ML) and deep learning (DL) for survival analysis models are increasingly used to predict Alzheimer’s disease (AD) progression, yet their comparative effectiveness remains unclear. This review evaluates the performance of ML and DL for survival analysis models across AD stages and identifies methodological gaps for clinical translation. We conducted a systematic review and meta-analysis following PRISMA guidelines. We searched PubMed, IEEE Xplore, Cochrane Library, EMBASE, Google Scholar, and Scopus for studies published between January 1, 2015 and May 12, 2026. Eligible studies employed ML/DL survival analysis for AD progression stages. Two independent reviewers assessed study inclusion using eligibility and quality criteria. Pooled concordance indices (C-index) with 95% confidence intervals were estimated using random-effects meta-analysis, with heterogeneity assessed using Cochran’s \\(Q\\) and \\(I^2\\) . Primary analysis pooled C-index by progression stage and contrasted ML vs DL. Secondary analysis of horizon-specific pooling (short term \\( < \\) 1 year; medium term 1–3 years; long term 3–5 years) was conducted within the MCI to AD transition. We assessed the reporting quality and risk of bias using TRIPOD-AI and PROBAST, respectively. Twenty-three studies were included in this systematic review. Pooled C-index estimates were: CN to MCI (0.83; 95% CI: 0.72–0.94; \\(\\text{I}^2\\) = 92.4%), MCI to AD (0.82; 95% CI: 0.77–0.87; \\(\\text{I}^2\\) = 89.1%), CN to AD (0.85; 95% CI: 0.77–0.94; \\(\\text{I}^2\\) = 94.7%), and CN/MCI to AD (0.85; 95% CI 0.83–0.86: \\(\\text{I}^2\\) = 0.0%). Subgroup analyses did not identify statistically significant differences between ML and DL models for MCI to AD ( \\(p\\) = 0.715), CN to MCI ( \\(p\\) = 0.201), or CN/MCI to AD ( \\(p\\) = 0.340). Horizon-specific analysis of MCI to AD was limited to long-term predictions for both ML (0.83, 95% CI 0.74–0.93) and DL (0.81, 95% CI 0.75–0.87), so prediction horizons could not be compared. Meta-regression analyses did not identify significant effects of model type or sample size on predictive performance. Funnel plot asymmetry tests and trim-and-fill analyses did not provide clear evidence of publication bias. Reporting completeness ranged from 64 to 82%, and 5/23 studies were judged at high risk of bias. Both ML and DL survival models demonstrated moderate to good discrimination for predicting AD progression, but no consistent superiority of one model class over the other was observed. Reported performance appeared to be influenced more by disease stage, cohort characteristics, predictor modalities, and validation practices than by algorithm choice alone. Methodological heterogeneity, limited external validation, and inconsistent reporting remain important barriers to clinical translation. Future studies should emphasize standardized reporting, transparent model development, and rigorous external validation across independent cohorts. CRD42024629585
Authors
- Ferial Abuhantash
- Leontios Hadjileontiadis
- Natnael Tumzghi
- Mohamed Seghier
- Aamna AlShehhi
Institutions
- Abu Dhabi University (AE)
- Khalifa University of Science and Technology (AE)
- Aristotle University of Thessaloniki (GR)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1186/s12911-026-03822-5
- Primary Topic
- Machine Learning in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00