Knowledge Distillation and Explainability Analysis for Lightweight Retinal Disease Classification Using MultiEYE Fundus Images

Background/Objectives: Class imbalance and probabilistic prediction quality are important considerations in lightweight retinal classification. We evaluated whether knowledge distillation (KD) from a ConvNeXtV2-Base teacher improved an EfficientNet-B0 student in nine-class MultiEYE fundus classification. Methods: Exact-content screening quarantined 131 of 58,036 records because of exact-content duplication or label conflicts. Six prespecified seeds were evaluated in no-KD, label-smoothing, and KD arms (18 runs) under matched training conditions. Within the controlled experiment, TEST data were not used for training, hyperparameter tuning, checkpoint or model selection, or protocol modification. The primary endpoint was the KD−no-KD macro-F1 difference on decontaminated TEST (n = 11,573), assessed by paired image-level bootstrap (10,000 replicates). Grad-CAM and Grad-CAM++ were examined qualitatively using one deterministically selected case per class and six-seed consensus maps. Results: Mean macro-F1 was 0.593, 0.597, and 0.633 for no-KD, label smoothing, and KD. KD exceeded no-KD by 0.039 (95% CI, 0.025–0.053; p < 0.001), with positive paired differences in all six seeds, and label smoothing by 0.036 (Holm-adjusted p < 0.001); label smoothing did not differ from no-KD (Holm-adjusted p = 0.554). Class-wise F1 differences were non-negative across all nine classes, with six remaining significant after correction for multiple comparisons. The Brier score difference was −0.077 (95% CI, −0.081 to −0.072; Holm-adjusted p < 0.001), while ECE decreased descriptively from 0.171 to 0.084. Nine-class Grad-CAM and direct Grad-CAM++ consensus maps enabled qualitative no-KD versus KD comparison without establishing lesion-localization superiority. Conclusions: KD improved macro-F1 and lowered the Brier score while retaining the same EfficientNet-B0 student at inference. The evaluated label-smoothing configuration did not reproduce the macro-F1 gain; clinical superiority and external generalizability remain unestablished.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-11
DOI
https://doi.org/10.3390/diagnostics16182945
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowledge Distillation and Explainability Analysis for Lightweight Retinal Disease Classification Using MultiEYE Fundus Images

Arif Koyun, Gülşen Türker
Diagnostics
Retinal Imaging and Analysis
article

Knowledge Distillation and Explainability Analysis for Lightweight Retinal Disease Classification Using MultiEYE Fundus Images

Arif Koyun, Gülşen Türker
article en

Abstract

Background/Objectives: Class imbalance and probabilistic prediction quality are important considerations in lightweight retinal classification. We evaluated whether knowledge distillation (KD) from a ConvNeXtV2-Base teacher improved an EfficientNet-B0 student in nine-class MultiEYE fundus classification. Methods: Exact-content screening quarantined 131 of 58,036 records because of exact-content duplication or label conflicts. Six prespecified seeds were evaluated in no-KD, label-smoothing, and KD arms (18 runs) under matched training conditions. Within the controlled experiment, TEST data were not used for training, hyperparameter tuning, checkpoint or model selection, or protocol modification. The primary endpoint was the KD−no-KD macro-F1 difference on decontaminated TEST (n = 11,573), assessed by paired image-level bootstrap (10,000 replicates). Grad-CAM and Grad-CAM++ were examined qualitatively using one deterministically selected case per class and six-seed consensus maps. Results: Mean macro-F1 was 0.593, 0.597, and 0.633 for no-KD, label smoothing, and KD. KD exceeded no-KD by 0.039 (95% CI, 0.025–0.053; p < 0.001), with positive paired differences in all six seeds, and label smoothing by 0.036 (Holm-adjusted p < 0.001); label smoothing did not differ from no-KD (Holm-adjusted p = 0.554). Class-wise F1 differences were non-negative across all nine classes, with six remaining significant after correction for multiple comparisons. The Brier score difference was −0.077 (95% CI, −0.081 to −0.072; Holm-adjusted p < 0.001), while ECE decreased descriptively from 0.171 to 0.084. Nine-class Grad-CAM and direct Grad-CAM++ consensus maps enabled qualitative no-KD versus KD comparison without establishing lesion-localization superiority. Conclusions: KD improved macro-F1 and lowered the Brier score while retaining the same EfficientNet-B0 student at inference. The evaluated label-smoothing configuration did not reproduce the macro-F1 gain; clinical superiority and external generalizability remain unestablished.

DiagnosticsVol. 16(18)
Süleyman Demirel Üniversitesi (TR)
Quality Education
Openalex Percentile: Top 11%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Knowledge Distillation and Explainability Analysis for Lightweight Retinal Disease Classification Using MultiEYE Fundus Images — Arif Koyun, Gülşen Türker · Diagnostics (2026) | TGRS Research Map | TGRS