Image‐Based Diagnosis of Oral Lesions: Performance of a Vision‐Language Model versus Human Clinicians

OBJECTIVE: To evaluate the real-world diagnostic performance of a multimodal large language model (LLM) for image-based assessment of oral mucosal lesions compared with clinicians of varying expertise. STUDY DESIGN: Prospective international multicenter diagnostic accuracy study. SETTING: Twenty university and tertiary head and neck centers in Italy, Belgium, France, Spain, and Israel. METHODS: We enrolled 350 consecutive patients (320 with oral lesions, 30 with normal mucosa). Clinical photographs and basic epidemiologic data were analyzed using Gemini 2.5 Advanced with a standardized prompt. Model outputs for lesion detection, malignancy versus benign versus normal, precise histologic diagnosis, and urgency class were compared with histopathology and with 4 clinicians. Sensitivity, specificity, accuracy, and agreement were calculated. RESULTS: AI-Gemini achieved 97.1% accuracy for lesion detection and malignancy classification, with sensitivity 98.5% and specificity 96.2% for malignancy, and 88.0% accuracy for precise histologic diagnosis. The head and neck surgeon achieved the highest accuracy for precise diagnosis (97.7%). Three-class diagnostic accuracy was 94.2% for AI-Gemini and 67.0% to 86.5% for nonsurgeon clinicians. Urgency assignment was correct in 70% of cases (κ = 0.716), with a conservative tendency to overestimate risk. Agreement with histologic diagnosis was almost perfect (κ = 0.929). CONCLUSION: In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. However, given the exploratory single-reader design, these findings should not be interpreted as evidence of equivalence or superiority relative to clinicians and require further prospective external validation before any potential clinical deployment in telemedicine or primary care settings.

Authors

Institutions

Publication Details

Journal
Otolaryngology
Published
2026-09-21
DOI
https://doi.org/10.1002/ohn.70437
Primary Topic
Head and Neck Cancer Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Image‐Based Diagnosis of Oral Lesions: Performance of a Vision‐Language Model versus Human Clinicians

Guido Gabriele, Fabiana Allevi, Alberto Maria Saibene, Andrea Frosolini et al.
Otolaryngology
Head and Neck Cancer Studies
article

Image‐Based Diagnosis of Oral Lesions: Performance of a Vision‐Language Model versus Human Clinicians

Guido Gabriele, Fabiana Allevi, Alberto Maria Saibene, Andrea Frosolini, Fabiola Giudici, Thomas Radulesco, Miguel Mayo-Yáñez, Resi Pucci, Fabio Maglitto, Marzia Petrocelli, Luigi Angelo Vaira, Salvatore Crimi, Giacomo De Riu, Valentino Vellone, Umberto Committeri, Francesco Laganà, Giulio Cirignaco, Fabrizio Spallaccia, Sholem Hack, Paolo Boscolo-Rizzo, Eleonora M. C. Trecca, Jerome R. Lechien, Giovanni M. Soro, Andrea Gasparato, Carlos M. Chiesa‐Estomba, Bernardo Bianchi, Giovanni Salzano, Andrea De Vito, Giuseppe Consorti, Giannicola Iannella, Antonino Maniaci, Stefania Troise, Jerry Polesel
article en

Abstract

OBJECTIVE: To evaluate the real-world diagnostic performance of a multimodal large language model (LLM) for image-based assessment of oral mucosal lesions compared with clinicians of varying expertise. STUDY DESIGN: Prospective international multicenter diagnostic accuracy study. SETTING: Twenty university and tertiary head and neck centers in Italy, Belgium, France, Spain, and Israel. METHODS: We enrolled 350 consecutive patients (320 with oral lesions, 30 with normal mucosa). Clinical photographs and basic epidemiologic data were analyzed using Gemini 2.5 Advanced with a standardized prompt. Model outputs for lesion detection, malignancy versus benign versus normal, precise histologic diagnosis, and urgency class were compared with histopathology and with 4 clinicians. Sensitivity, specificity, accuracy, and agreement were calculated. RESULTS: AI-Gemini achieved 97.1% accuracy for lesion detection and malignancy classification, with sensitivity 98.5% and specificity 96.2% for malignancy, and 88.0% accuracy for precise histologic diagnosis. The head and neck surgeon achieved the highest accuracy for precise diagnosis (97.7%). Three-class diagnostic accuracy was 94.2% for AI-Gemini and 67.0% to 86.5% for nonsurgeon clinicians. Urgency assignment was correct in 70% of cases (κ = 0.716), with a conservative tendency to overestimate risk. Agreement with histologic diagnosis was almost perfect (κ = 0.929). CONCLUSION: In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. However, given the exploratory single-reader design, these findings should not be interpreted as evidence of equivalence or superiority relative to clinicians and require further prospective external validation before any potential clinical deployment in telemedicine or primary care settings.

Otolaryngology
University of Siena (IT), University of Foggia (IT), Marche Polytechnic University (IT), Centre National de la Recherche Scientifique (FR), AOL (United States) (US), University of Mons (BE), University of Trieste (IT), University of Sassari (IT), Aix-Marseille Université (FR), Università degli Studi di Enna Kore (IT), Santa Maria Nuova Hospital (IT), Sheba Medical Center (IL), Casa Sollievo della Sofferenza (IT), University of Catania (IT), Hôpital de la Conception (FR), Biogipuzkoa Health Research Institute (ES), Ospedale San Paolo (IT), Azienda USL di Bologna (IT), Centro di Riferimento Oncologico (IT), IRCCS San Camillo Hospital (IT), Complexo Hospitalario Universitario A Coruña (ES), Ospedale Policlinico San Martino (IT), Ospedale Bellaria (IT), Istituti di Ricovero e Cura a Carattere Scientifico (IT), Ospedali Riuniti Umberto I (IT), Link Campus University (IT), University of Naples Federico II (IT), Sapienza University of Rome (IT)
Openalex Percentile: Top 8%
Head and Neck Cancer Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.