Comparing Clinical Metadata Fusion Strategies for Diabetic and Hypertensive Retinopathy Classification

Structured clinical variables may complement fundus photographs, but the most appropriate point at which to integrate them into a retinal foundation model remains uncertain. This study compares an image-only baseline with three strategies for adding age, sex, diabetes status and hypertension status to a DINOv3-based classifier, namely a manually encoded visual metadata patch, learned metadata tokens and late fusion by feature concatenation. DINOv3 ViT-B/16 and ViT-L/16 backbones were adapted with Low-Rank Adaptation under a matched training protocol and evaluated on three patient-wise folds of the Brazilian Multilabel Ophthalmological Dataset. The diabetic retinopathy task comprised 9531 images with 1070 positive cases and the hypertensive retinopathy task comprised 8745 images with 284 positive cases. For diabetic retinopathy with ViT-L/16, all three strategies improved the F1-score over the image-only model. Late fusion reached an ROC-AUC of 0.9943, an AUC-PR of 0.9700 and an F1-score of 0.9326; the visual patch had a similar AUC-PR and learned tokens had the highest sensitivity. For hypertensive retinopathy, late fusion had the highest mean ROC-AUC of 0.8970, AUC-PR of 0.4441 and F1-score of 0.4338, but its paired advantage over the image-only model was limited to threshold-dependent metrics. Removing the target-related systemic variable reduced the late-fusion F1-score from 0.9326 to 0.9228 for diabetic retinopathy and from 0.4338 to 0.3776 for hypertensive retinopathy, showing that part of the multimodal gain came from disease-related clinical context. Models using only clinical metadata performed substantially worse than the image-based models. Metadata integration was beneficial in several matched comparisons, but early integration was not uniformly superior and late fusion yielded the most consistently favorable mean results.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-09-28
DOI
https://doi.org/10.3390/make8100299
Primary Topic
Retinal Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparing Clinical Metadata Fusion Strategies for Diabetic and Hypertensive Retinopathy Classification

Rodolfo Deusdará, Thiago Alves Martins, Vinícius de Carvalho Rispoli, Weverson Garcia Medeiros et al.
Machine Learning and Knowledge Extraction
Retinal Imaging and Analysis
article

Comparing Clinical Metadata Fusion Strategies for Diabetic and Hypertensive Retinopathy Classification

Rodolfo Deusdará, Thiago Alves Martins, Vinícius de Carvalho Rispoli, Weverson Garcia Medeiros, Alan Müller, Marília Miranda Forte Gomes
article en

Abstract

Structured clinical variables may complement fundus photographs, but the most appropriate point at which to integrate them into a retinal foundation model remains uncertain. This study compares an image-only baseline with three strategies for adding age, sex, diabetes status and hypertension status to a DINOv3-based classifier, namely a manually encoded visual metadata patch, learned metadata tokens and late fusion by feature concatenation. DINOv3 ViT-B/16 and ViT-L/16 backbones were adapted with Low-Rank Adaptation under a matched training protocol and evaluated on three patient-wise folds of the Brazilian Multilabel Ophthalmological Dataset. The diabetic retinopathy task comprised 9531 images with 1070 positive cases and the hypertensive retinopathy task comprised 8745 images with 284 positive cases. For diabetic retinopathy with ViT-L/16, all three strategies improved the F1-score over the image-only model. Late fusion reached an ROC-AUC of 0.9943, an AUC-PR of 0.9700 and an F1-score of 0.9326; the visual patch had a similar AUC-PR and learned tokens had the highest sensitivity. For hypertensive retinopathy, late fusion had the highest mean ROC-AUC of 0.8970, AUC-PR of 0.4441 and F1-score of 0.4338, but its paired advantage over the image-only model was limited to threshold-dependent metrics. Removing the target-related systemic variable reduced the late-fusion F1-score from 0.9326 to 0.9228 for diabetic retinopathy and from 0.4338 to 0.3776 for hypertensive retinopathy, showing that part of the multimodal gain came from disease-related clinical context. Models using only clinical metadata performed substantially worse than the image-based models. Metadata integration was beneficial in several matched comparisons, but early integration was not uniformly superior and late fusion yielded the most consistently favorable mean results.

Machine Learning and Knowledge ExtractionVol. 8(10)
Universidade de Brasília (BR)
Good health and well-being
Openalex Percentile: Top 12%
Retinal Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.