Robust and interpretable retinal disease classification: a cross-source validation study
Abstract Retinal diseases such as hypertensive retinopathy (HR), myopia, and retinitis pigmentosa (RP) may remain undiagnosed until irreversible visual impairment occurs. Retinal multi disease classification models are often developed by integrating multiple public fundus datasets and evaluated using internal validation. We explored whether such evaluation amply reflects disease-related performance or may obscure acquisition-source effects. A total of 5,373 fundus images representing HR, myopia, normal retina, and RP were analyzed following content-level deduplication. ImageNet-pretrained ResNet-18 embeddings were combined with 26 scale-invariant handcrafted features and assessed using group-aware cross-validation and a sealed held-out test set. The image-only model achieved 0.872 test accuracy (95% CI: 0.850–0.890), while the handcrafted features alone achieved 0.810 accuracy. Further, the source-aware analysis discovered substantial acquisition-related confounding. All RP images originated from a single acquisition source, and image resolution alone separated RP from the other classes with an F1 score of 1.000. The source identity was also predictable (97.4%) from the retained features. External evaluation was conducted using two independent datasets that resulted in noticeably lower accuracies of 0.580 and 0.601. Particularly, the model assigned RP to 37.5% of images in a dataset containing no RP cases. Thus, conventional internal validation did not reveal vulnerabilities that became evident under cross-source and external evaluation. In contrast, myopia showed more consistent cross-source recognition, with leave-one-source-out sensitivity ranging from 0.80 to 0.97 across four independent sources. Amongst the handcrafted features, tessellation index and optic disc area ratio showed consistent associations with myopia across permutation importance, gradient attribution, and FDR-corrected statistical testing, while neither feature predicted acquisition source more accurately than disease class. These findings demonstrate that high internal performance in multi-source retinal classification can coexist with substantial vulnerability to acquisition-source differences. Source-aware partitioning and independent external evaluation should therefore be incorporated when assessing the robustness and generalizability of models developed from heterogeneous retinal datasets.
Authors
- Zoya Khalid (ORCID: https://orcid.org/0000-0003-2075-462X)
- Osman Uğur Sezerman (ORCID: https://orcid.org/0000-0003-0905-6783)
- Fathima Hafsa
- Azkiya Manal
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1038/s41598-026-73785-0
- Primary Topic
- Retinal Imaging and Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00