Performance of AI-based diabetic retinopathy screening is highly dependent on evaluation setting: a five-year, multi-framework study

Abstract Despite near-perfect performance reported on benchmark datasets, the real-world behavior of artificial intelligence (AI) systems for diabetic retinopathy (DR) screening remains insufficiently characterized. Here, we report a multi-framework evaluation of OphtAI, a CE-marked AI system for automated detection of referable DR and diabetic macular edema from color fundus photographs. The system was initially validated on the Messidor-2 benchmark dataset and subsequently assessed across three independent evaluation settings: a large-scale masked comparative study (US Veterans Affairs), an external validation using handheld fundus photography in Finland, and an open, large-scale comparative evaluation within the UK National Health Service. Across these evaluation frameworks, substantial variability in observed sensitivity and specificity was found. These variations occurred across settings differing in imaging devices, population characteristics, referral definitions, and handling of ungradable images. In masked settings, interpretation of comparative performance was limited by the lack of system-level attribution, while open evaluations remained sensitive to protocol design. Taken together, these results show that performance in AI-based DR screening is not a fixed property of a system, but an emergent property of the evaluation framework. This work calls for context-aware validation strategies and more standardized evaluation protocols for clinical AI systems.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-09
DOI
https://doi.org/10.1038/s41598-026-70676-2
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance of AI-based diabetic retinopathy screening is highly dependent on evaluation setting: a five-year, multi-framework study

Petri Huhtinen, Alexandre Le Guilcher, Gwenolé Quellec, Pascale Massin et al.
Scientific Reports
Retinal Imaging and Analysis
article

Performance of AI-based diabetic retinopathy screening is highly dependent on evaluation setting: a five-year, multi-framework study

Petri Huhtinen, Alexandre Le Guilcher, Gwenolé Quellec, Pascale Massin, Ali Erginay, Pierre Deman, Laurent Borderie, Béatrice Cochener, Philippe Zhang, Nina Hautala, Sarah Matta, Paola Ortega Saborio, Anna-Maria Kubin, Mathieu Lamard
article en

Abstract

Abstract Despite near-perfect performance reported on benchmark datasets, the real-world behavior of artificial intelligence (AI) systems for diabetic retinopathy (DR) screening remains insufficiently characterized. Here, we report a multi-framework evaluation of OphtAI, a CE-marked AI system for automated detection of referable DR and diabetic macular edema from color fundus photographs. The system was initially validated on the Messidor-2 benchmark dataset and subsequently assessed across three independent evaluation settings: a large-scale masked comparative study (US Veterans Affairs), an external validation using handheld fundus photography in Finland, and an open, large-scale comparative evaluation within the UK National Health Service. Across these evaluation frameworks, substantial variability in observed sensitivity and specificity was found. These variations occurred across settings differing in imaging devices, population characteristics, referral definitions, and handling of ungradable images. In masked settings, interpretation of comparative performance was limited by the lack of system-level attribution, while open evaluations remained sensitive to protocol design. Taken together, these results show that performance in AI-based DR screening is not a fixed property of a system, but an emergent property of the evaluation framework. This work calls for context-aware validation strategies and more standardized evaluation protocols for clinical AI systems.

Scientific Reports
Good health and well-being
Openalex Percentile: Top 11%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.