Knowledge Distillation and Explainability Analysis for Lightweight Retinal Disease Classification Using MultiEYE Fundus Images
Background/Objectives: Class imbalance and probabilistic prediction quality are important considerations in lightweight retinal classification. We evaluated whether knowledge distillation (KD) from a ConvNeXtV2-Base teacher improved an EfficientNet-B0 student in nine-class MultiEYE fundus classification. Methods: Exact-content screening quarantined 131 of 58,036 records because of exact-content duplication or label conflicts. Six prespecified seeds were evaluated in no-KD, label-smoothing, and KD arms (18 runs) under matched training conditions. Within the controlled experiment, TEST data were not used for training, hyperparameter tuning, checkpoint or model selection, or protocol modification. The primary endpoint was the KD−no-KD macro-F1 difference on decontaminated TEST (n = 11,573), assessed by paired image-level bootstrap (10,000 replicates). Grad-CAM and Grad-CAM++ were examined qualitatively using one deterministically selected case per class and six-seed consensus maps. Results: Mean macro-F1 was 0.593, 0.597, and 0.633 for no-KD, label smoothing, and KD. KD exceeded no-KD by 0.039 (95% CI, 0.025–0.053; p < 0.001), with positive paired differences in all six seeds, and label smoothing by 0.036 (Holm-adjusted p < 0.001); label smoothing did not differ from no-KD (Holm-adjusted p = 0.554). Class-wise F1 differences were non-negative across all nine classes, with six remaining significant after correction for multiple comparisons. The Brier score difference was −0.077 (95% CI, −0.081 to −0.072; Holm-adjusted p < 0.001), while ECE decreased descriptively from 0.171 to 0.084. Nine-class Grad-CAM and direct Grad-CAM++ consensus maps enabled qualitative no-KD versus KD comparison without establishing lesion-localization superiority. Conclusions: KD improved macro-F1 and lowered the Brier score while retaining the same EfficientNet-B0 student at inference. The evaluated label-smoothing configuration did not reproduce the macro-F1 gain; clinical superiority and external generalizability remain unestablished.
Authors
- Arif Koyun (ORCID: https://orcid.org/0000-0001-6701-363X)
- Gülşen Türker (ORCID: https://orcid.org/0000-0003-3210-5061)
Institutions
- Süleyman Demirel Üniversitesi (TR)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-11
- DOI
- https://doi.org/10.3390/diagnostics16182945
- Primary Topic
- Retinal Imaging and Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00