A Fair Evaluation of Lightweight CNNs and Ordinal Regression for Diabetic Retinopathy Grading: The Double-Rebalancing Pitfall

Diabetic retinopathy (DR) is a leading cause of preventable blindness worldwide, affecting approximately one-third of diabetic patients. Early detection through automated retinal screening systems can significantly reduce the burden of vision loss. In this study, we present a comparative analysis of three lightweight convolutional neural networks—EfficientNet-B0, MobileNetV3-Small, and DenseNet-121—for five-class DR severity grading using the APTOS 2019 dataset. We employ a two-phase transfer learning strategy with ImageNet-pretrained backbones, enhanced by CLAHE and Ben Graham preprocessing. To examine how the class-balancing configuration and the loss function affect ordinal DR grading, particularly for Moderate DR detection, we conduct ablation studies comparing CrossEntropy, Focal Loss, and CORN (Conditional Ordinal Regression) loss functions. Our central finding is methodological: combining a class-balanced sampler with an inverse-frequency weighted loss (“double rebalancing”) silently collapses the second-largest class (Moderate DR) to an F1-score of 0.00, whereas removing either component restores Moderate-DR F1 to 0.63–0.75. Under a fair single-rebalancing configuration, well-configured lightweight CNNs are strong and generalizable, with the best (CrossEntropy) configuration reaching APTOS QWK ≈ 0.88 and, zero-shot on IDRiD, QWK ≈ 0.72 with referable-DR AUC ≈ 0.95 (CORN slightly lower). Ordinal (CORN) and standard (CE) losses perform comparably and are statistically indistinguishable on APTOS, although on IDRiD CrossEntropy is significantly better on mean absolute error; CORN’s practical advantage is robustness, reaching comparable performance without class-weight tuning and never collapsing a class. Grad-CAM is used to qualitatively illustrate model attention. These results indicate that correct class-balancing configuration matters more than the choice of loss function for ordinal DR grading.

Authors

Publication Details

Journal
Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri Dergisi
Published
2026-09-17
DOI
https://doi.org/10.65520/erciyesfen.1970251
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Fair Evaluation of Lightweight CNNs and Ordinal Regression for Diabetic Retinopathy Grading: The Double-Rebalancing Pitfall

Fatih ŞAHİN
Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri Dergisi
Retinal Imaging and Analysis
article

A Fair Evaluation of Lightweight CNNs and Ordinal Regression for Diabetic Retinopathy Grading: The Double-Rebalancing Pitfall

Fatih ŞAHİN
article en

Abstract

Diabetic retinopathy (DR) is a leading cause of preventable blindness worldwide, affecting approximately one-third of diabetic patients. Early detection through automated retinal screening systems can significantly reduce the burden of vision loss. In this study, we present a comparative analysis of three lightweight convolutional neural networks—EfficientNet-B0, MobileNetV3-Small, and DenseNet-121—for five-class DR severity grading using the APTOS 2019 dataset. We employ a two-phase transfer learning strategy with ImageNet-pretrained backbones, enhanced by CLAHE and Ben Graham preprocessing. To examine how the class-balancing configuration and the loss function affect ordinal DR grading, particularly for Moderate DR detection, we conduct ablation studies comparing CrossEntropy, Focal Loss, and CORN (Conditional Ordinal Regression) loss functions. Our central finding is methodological: combining a class-balanced sampler with an inverse-frequency weighted loss (“double rebalancing”) silently collapses the second-largest class (Moderate DR) to an F1-score of 0.00, whereas removing either component restores Moderate-DR F1 to 0.63–0.75. Under a fair single-rebalancing configuration, well-configured lightweight CNNs are strong and generalizable, with the best (CrossEntropy) configuration reaching APTOS QWK ≈ 0.88 and, zero-shot on IDRiD, QWK ≈ 0.72 with referable-DR AUC ≈ 0.95 (CORN slightly lower). Ordinal (CORN) and standard (CE) losses perform comparably and are statistically indistinguishable on APTOS, although on IDRiD CrossEntropy is significantly better on mean absolute error; CORN’s practical advantage is robustness, reaching comparable performance without class-weight tuning and never collapsing a class. Grad-CAM is used to qualitatively illustrate model attention. These results indicate that correct class-balancing configuration matters more than the choice of loss function for ordinal DR grading.

Erciyes Üniversitesi Fen Bilimleri Enstitüsü Fen Bilimleri DergisiVol. 42(3)
Openalex Percentile: Top 12%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.