Decomposing Cross-Modality Failure in Diabetic Retinopathy Grading

Deep learning detects referable diabetic retinopathy (DR) from colour fundus photography at a level comparable with specialists, but such models are rarely tested across a change of imaging modality. We evaluate a colour-trained DR grader zero-shot on near-infrared reflectance images acquired by confocal scanning laser ophthalmoscopy, and report two findings. First, the transfer fails without any detectable signal in the model output: 96.1% of images collapse to the no-DR class while mean confidence rises from 0.619 to 0.675 (Cohen’s d = 0.41), so an abstention policy calibrated in domain would pass the collapsed predictions through. The apparent severe-grade tail correlates with image brightness and is unrelated to pathology. Second, and centrally, the failure is not monolithic. A deterministic green-channel intensity transform, which by construction cannot synthesise structure, recovers substantial transferable signal: collapse falls to 53.2%, grade-to-thickness concordance flips from -0.492 to +0.332, and sensitivity for a macular fluid composite rises from 5.2% to 65.2%, at a cost of 0.020 in in-domain validation quadratic weighted kappa. Because the transform cannot fabricate structure, the recovered signal is attributable to distributional alignment alone, which places a measurable lower bound on the preprocessing-addressable share of the gap and bounds the remainder. The dominant barrier is representational: single-wavelength reflectance replicated across three channels and normalised with colour-image statistics occupies a region of input space the encoder never encountered in training, and the chromatic contribution is comparatively small. We validate indirectly against clinical proxies, since no public near-infrared dataset carries ground-truth DR grades, and we are explicit that this constitutes convergent evidence and does not amount to proof.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22725651
Primary Topic
Retinal Diseases and Treatments
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Decomposing Cross-Modality Failure in Diabetic Retinopathy Grading

Savita Nair
Zenodo (CERN European Organization for Nuclear Research)
Retinal Diseases and Treatments
preprint

Decomposing Cross-Modality Failure in Diabetic Retinopathy Grading

Savita Nair
preprint en

Abstract

Deep learning detects referable diabetic retinopathy (DR) from colour fundus photography at a level comparable with specialists, but such models are rarely tested across a change of imaging modality. We evaluate a colour-trained DR grader zero-shot on near-infrared reflectance images acquired by confocal scanning laser ophthalmoscopy, and report two findings. First, the transfer fails without any detectable signal in the model output: 96.1% of images collapse to the no-DR class while mean confidence rises from 0.619 to 0.675 (Cohen’s d = 0.41), so an abstention policy calibrated in domain would pass the collapsed predictions through. The apparent severe-grade tail correlates with image brightness and is unrelated to pathology. Second, and centrally, the failure is not monolithic. A deterministic green-channel intensity transform, which by construction cannot synthesise structure, recovers substantial transferable signal: collapse falls to 53.2%, grade-to-thickness concordance flips from -0.492 to +0.332, and sensitivity for a macular fluid composite rises from 5.2% to 65.2%, at a cost of 0.020 in in-domain validation quadratic weighted kappa. Because the transform cannot fabricate structure, the recovered signal is attributable to distributional alignment alone, which places a measurable lower bound on the preprocessing-addressable share of the gap and bounds the remainder. The dominant barrier is representational: single-wavelength reflectance replicated across three channels and normalised with colour-image statistics occupies a region of input space the encoder never encountered in training, and the chromatic contribution is comparatively small. We validate indirectly against clinical proxies, since no public near-infrared dataset carries ground-truth DR grades, and we are explicit that this constitutes convergent evidence and does not amount to proof.

Zenodo (CERN European Organization for Nuclear Research)
University of Bath (GB)
Retinal Diseases and Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.