Diabetic retinopathy screening using artificial intelligence: a comprehensive performance comparison on datasets from diverse populations and imaging modalities

Objective Diabetic retinopathy (DR) is the leading cause of preventable blindness in adults and poses significant challenges in low- and middle-income regions due to limited access to skilled clinicians and diagnostic facilities. Automated screening solutions using artificial intelligence (AI) have emerged as an efficient alternative, achieving high diagnostic accuracy. However, these solutions are often developed using data from specific populations obtained using relatively expensive high-end devices. This study addresses the potential scope limitations by evaluating the AI-based screening tool retina.help across eight datasets representing diverse populations, imaging modalities, and geographic regions. Approach The datasets include both public and private sources, with images captured using tabletop and handheld fundus cameras. Key performance metrics for detecting binary referable DR - sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC) - were calculated on an image-by-image basis. Main results Retina.help demonstrated high accuracy on tabletop images, achieving AUROC values of 0.97 on the BRSET and DeepDRiD datasets. Handheld device performance was more variable, with AUROC ranging from 0.88 (Filipino dataset) to 0.99 (Finnish dataset). Sensitivity declined with increased retinal pigmentation, as evidenced by lower values for datasets from Tanzania (62.2%) and Brazil (76.7%) compared to Finland (89.9%). Images from handheld devices often yielded lower sensitivity due to challenges related to low-contrast images. Nonetheless, retina.help generalized well across diverse datasets, showcasing its robustness. Significance The study highlights the impact of imaging equipment, demographics, and image quality on diagnostic performance. These findings underscore the need for benchmarking AI-based DR screening tools using standardized datasets that encompass diverse populations and imaging conditions. Such evaluations can guide the development of equitable, reliable and robust screening solutions. .

Authors

Institutions

Publication Details

Journal
Physiological Measurement
Published
2026-09-04
DOI
https://doi.org/10.1088/1361-6579/aea2ee
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diabetic retinopathy screening using artificial intelligence: a comprehensive performance comparison on datasets from diverse populations and imaging modalities

Paolo S. Silva, Leo Anthony Celi, Luis Filipe Nakayama, Philipp Schapotschnikow et al.
Physiological Measurement
Retinal Imaging and Analysis
article

Diabetic retinopathy screening using artificial intelligence: a comprehensive performance comparison on datasets from diverse populations and imaging modalities

Paolo S. Silva, Leo Anthony Celi, Luis Filipe Nakayama, Philipp Schapotschnikow, Murat Osswald, Vladyslav Dusiak
article en

Abstract

Objective Diabetic retinopathy (DR) is the leading cause of preventable blindness in adults and poses significant challenges in low- and middle-income regions due to limited access to skilled clinicians and diagnostic facilities. Automated screening solutions using artificial intelligence (AI) have emerged as an efficient alternative, achieving high diagnostic accuracy. However, these solutions are often developed using data from specific populations obtained using relatively expensive high-end devices. This study addresses the potential scope limitations by evaluating the AI-based screening tool retina.help across eight datasets representing diverse populations, imaging modalities, and geographic regions. Approach The datasets include both public and private sources, with images captured using tabletop and handheld fundus cameras. Key performance metrics for detecting binary referable DR - sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC) - were calculated on an image-by-image basis. Main results Retina.help demonstrated high accuracy on tabletop images, achieving AUROC values of 0.97 on the BRSET and DeepDRiD datasets. Handheld device performance was more variable, with AUROC ranging from 0.88 (Filipino dataset) to 0.99 (Finnish dataset). Sensitivity declined with increased retinal pigmentation, as evidenced by lower values for datasets from Tanzania (62.2%) and Brazil (76.7%) compared to Finland (89.9%). Images from handheld devices often yielded lower sensitivity due to challenges related to low-contrast images. Nonetheless, retina.help generalized well across diverse datasets, showcasing its robustness. Significance The study highlights the impact of imaging equipment, demographics, and image quality on diagnostic performance. These findings underscore the need for benchmarking AI-based DR screening tools using standardized datasets that encompass diverse populations and imaging conditions. Such evaluations can guide the development of equitable, reliable and robust screening solutions. .

Physiological Measurement
Joslin Diabetes Center (US), RIKEN Center for Brain Science (JP), Moscow Institute of Thermal Technology (RU), Universidade Federal de São Paulo (BR)
National Science Foundation
No poverty
Openalex Percentile: Top 11%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.