Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations

Abstract Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores > 95 vs. 84.9; p < 0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and ChatGPT-5 achieved higher scores than students, whereas students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. Medical students demonstrated the smallest performance gap between text-only and image-based components, indicating greater cross-modal consistency. Sub-domain analyses revealed that some models achieved accurate visual description and diagnosis but showed reduced performance in etiological and treatment reasoning. Gemini 3 Pro exhibited the highest overall accuracy and the lowest output variability across repeated evaluations. These findings indicate that while VLMs excel in text-based dermatologic assessment, multimodal competence remains uneven and model-dependent, supporting their use as complementary rather than standalone tools in dermatology education.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-05
DOI
https://doi.org/10.1038/s41598-026-68943-3
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations

Ozan Erdem, Ece Gökyayla, Buğra Burç Dağtaş, Melek Aslan Kayıran et al.
Scientific Reports
Artificial Intelligence in Healthcare and Education
article

Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations

Ozan Erdem, Ece Gökyayla, Buğra Burç Dağtaş, Melek Aslan Kayıran, Abdurrahim Yilmaz, Ahmet Sait Şahin, Vefa Aslı Erdemir, Mehmet Salih Gurel
article en

Abstract

Abstract Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores > 95 vs. 84.9; p < 0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and ChatGPT-5 achieved higher scores than students, whereas students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. Medical students demonstrated the smallest performance gap between text-only and image-based components, indicating greater cross-modal consistency. Sub-domain analyses revealed that some models achieved accurate visual description and diagnosis but showed reduced performance in etiological and treatment reasoning. Gemini 3 Pro exhibited the highest overall accuracy and the lowest output variability across repeated evaluations. These findings indicate that while VLMs excel in text-based dermatologic assessment, multimodal competence remains uneven and model-dependent, supporting their use as complementary rather than standalone tools in dermatology education.

Scientific Reports
Ege University (TR), İstanbul Eğitim ve Araştırma Hastanesi (TR), Sağlık Bilimleri Üniversitesi (TR), Imperial College London (GB), Istanbul Medeniyet University (TR), University of Health Sciences Antigua (AG)
Imperial College London
Quality Education
Openalex Percentile: Top 75%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.