Multi-modal machine and deep learning framework for integrated diagnosis and prognosis of diabetic retinopathy
Abstract Diabetic retinopathy (DR) remains a major cause of preventable visual impairment, yet existing automated systems predominantly focus on diagnosis and provide limited support for disease progression assessment. We present a unified multimodal machine and deep learning framework for joint DR diagnosis and prognostic modeling by integrating retinal image representations, clinically interpretable handcrafted vascular and lesion descriptors, and systemic clinical variables. A Vision Transformer (ViT) branch captures high-level retinal representations, while a complementary machine-learning branch incorporates vascular, lesion, texture, and clinical features, which are subsequently integrated through late feature fusion for DR grading and pseudo-temporal progression modeling using an LSTM-based module. The framework further incorporates Grad-CAM and SHAP to provide complementary visual and feature-level interpretability. Experiments across four publicly available retinal fundus datasets APTOS, MESSIDOR, IDRiD, and EyePACS demonstrate competitive diagnostic performance, achieving a mean AUC of 0.976, while the prognostic component achieved a concordance index (C-index) of 0.872 and a Brier score of 0.104. SHAP analysis identified HbA1c, vascular tortuosity, and microaneurysm area ratio as prominent prognostic contributors, whereas haemorrhage area ratio and exudate coverage were among the most influential diagnostic features. Grad-CAM analysis further demonstrated lesion-focused attention, with an attention-localization IoU of 0.86 against ophthalmologist-annotated lesion masks. Importantly, the clinical variables used in the multimodal analysis were synthetically generated from published epidemiological distributions rather than obtained from patient-linked clinical records; therefore, the prognostic findings should be interpreted as a methodological proof of concept rather than clinical evidence. Overall, the proposed framework demonstrates the feasibility of combining deep retinal representations with interpretable image-derived and synthetic clinical features within a unified diagnostic–prognostic architecture, while prospective validation using real-world longitudinal, EHR-linked datasets remains essential before clinical translation.
Authors
- Jaikumar M. Patil (ORCID: https://orcid.org/0000-0002-9466-5462)
- Vikram Ratan
- Priyanka V. Deshmukh
Institutions
- Symbiosis International University (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1038/s41598-026-71929-w
- Primary Topic
- Retinal Imaging and Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00