Hierarchical deep learning with stability-driven normalization and regularization for accurate multi-class diabetic retinopathy diagnosis

Diabetic retinopathy (DR) is one of the major causes of vision loss worldwide; early detection can help prevent permanent vision loss. Several deep learning approaches have been developed for retinal image analysis, yet they often face practical challenges in accurately diagnosing diabetic retinopathy. Class imbalance is a common issue in DR, where disease stages share subtle visual traits, and models tend to overfit to the training data. To mitigate these issues in existing deep learning approaches, we developed a novel compact hierarchical vision encoder that combines convolutional features with tailored normalization and regularization modules. A normalization technique is applied to stabilize model training, followed by structural risk–based regularization to enhance generalization. A pooling module is embedded in the model to learn subtle retinal details for distinguishing disease severity. The model is trained and evaluated on the APTOS dataset for five-class DR classes. Model performance was evaluated using accuracy, ROC analysis, confusion matrices, and an ablation study. The proposed model achieved an accuracy of 98.91%, with micro-averaged precision, recall, and F1 of 98.91%, whereas macro-averaged measures are above 97.8%. The model achieved class wise precision and recall above 93% for positive cases and classified the no DR cases without error. The ROC analysis demonstrates perfect separability with AUC value between 0.99 and 1.00. The proposed framework demonstrated an approximately 15% improvement in accuracy compared with pretrained CNN and transformer models. The model consistently recognized the AD classes with higher recognition accuracy, with slight variations. Although the proposed model is evaluated only on the APTOS dataset, the results suggest that the framework has potential for broader applications in DR screening.

Authors

Institutions

Publication Details

Journal
Ain Shams Engineering Journal
Published
2026-09-21
DOI
https://doi.org/10.1016/j.asej.2026.104453
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hierarchical deep learning with stability-driven normalization and regularization for accurate multi-class diabetic retinopathy diagnosis

Mohd Khalid Awang, Tariq M. Ali, Mohammad Hijji, Ghulam Ali et al.
Ain Shams Engineering Journal
Retinal Imaging and Analysis
article

Hierarchical deep learning with stability-driven normalization and regularization for accurate multi-class diabetic retinopathy diagnosis

Mohd Khalid Awang, Tariq M. Ali, Mohammad Hijji, Ghulam Ali, Muhammad Hamza, Muhammad Ayaz, Azqa Jafar
article en

Abstract

Diabetic retinopathy (DR) is one of the major causes of vision loss worldwide; early detection can help prevent permanent vision loss. Several deep learning approaches have been developed for retinal image analysis, yet they often face practical challenges in accurately diagnosing diabetic retinopathy. Class imbalance is a common issue in DR, where disease stages share subtle visual traits, and models tend to overfit to the training data. To mitigate these issues in existing deep learning approaches, we developed a novel compact hierarchical vision encoder that combines convolutional features with tailored normalization and regularization modules. A normalization technique is applied to stabilize model training, followed by structural risk–based regularization to enhance generalization. A pooling module is embedded in the model to learn subtle retinal details for distinguishing disease severity. The model is trained and evaluated on the APTOS dataset for five-class DR classes. Model performance was evaluated using accuracy, ROC analysis, confusion matrices, and an ablation study. The proposed model achieved an accuracy of 98.91%, with micro-averaged precision, recall, and F1 of 98.91%, whereas macro-averaged measures are above 97.8%. The model achieved class wise precision and recall above 93% for positive cases and classified the no DR cases without error. The ROC analysis demonstrates perfect separability with AUC value between 0.99 and 1.00. The proposed framework demonstrated an approximately 15% improvement in accuracy compared with pretrained CNN and transformer models. The model consistently recognized the AD classes with higher recognition accuracy, with slight variations. Although the proposed model is evaluated only on the APTOS dataset, the results suggest that the framework has potential for broader applications in DR screening.

Ain Shams Engineering JournalVol. 17(12)
University of Okara (PK), Sultan Zainal Abidin University (MY), University of Tabuk (SA)
Openalex Percentile: Top 11%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.