Diagnostic performance of multimodal large language models in detecting pulp stones on panoramic radiographs

Abstract Background Pulp stones are calcified structures within the dental pulp that may increase the complexity of endodontic treatment and are commonly identified through radiographic examinations. Despite recent advances in artificial intelligence applied to dental imaging, the diagnostic capability of publicly available multimodal large language models (LLMs) for detecting pulp stones under zero-shot conditions remains unclear. Therefore, this study evaluated the diagnostic performance and response consistency of multimodal LLMs in detecting pulp stones on panoramic radiographs of posterior teeth. Methods This comparative and exploratory study evaluated the GPT-5 and Gemini 2.5 Flash multimodal LLMs using panoramic radiograph image crops containing posterior teeth. A total of 64 image crops, comprising 192 posterior teeth, were analyzed. The models were tested under standardized zero-shot conditions using independent sessions and a standardized English-language prompt. Each image was submitted six times to both models to assess response consistency. Diagnostic performance was evaluated using accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1-score with 95% confidence intervals. Intra-model consistency was assessed using majority agreement, complete agreement across repetitions, and Fleiss’ Kappa coefficient. Results The models demonstrated similar overall accuracy for binary detection of pulp stones, with 56.5% for ChatGPT and 57.0% for Gemini. ChatGPT showed high specificity (96.2%) and greater intra-model stability, whereas Gemini demonstrated higher sensitivity (81.8%) and superior performance in identifying the affected tooth and dental group. However, Gemini also presented lower specificity and greater variability across repeated evaluations. Fleiss’ Kappa coefficients indicated low intra-model agreement beyond chance for both models, particularly for Gemini. Conclusion Multimodal LLMs demonstrated limited capability in detecting pulp stones on panoramic radiographs, with distinct diagnostic profiles between the evaluated models. Although the findings suggest potential future applications of LLMs as auxiliary tools for dental image interpretation, their current performance does not yet support consistent clinical applicability for pulp stone detection.

Authors

Publication Details

Journal
BMC Oral Health
Published
2026-10-08
DOI
https://doi.org/10.1186/s12903-026-10119-6
Primary Topic
Dental Radiography and Imaging
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Diagnostic performance of multimodal large language models in detecting pulp stones on panoramic radiographs

Flares Baratto‐Filho, Michelle Nascimento Meger, Luíz Fernando Fariniuk, Érika Calvano Küchler et al.
BMC Oral Health
Dental Radiography and Imaging
article

Diagnostic performance of multimodal large language models in detecting pulp stones on panoramic radiographs

Flares Baratto‐Filho, Michelle Nascimento Meger, Luíz Fernando Fariniuk, Érika Calvano Küchler, Rafaela Scariot, Christian Kirschneck, Bianca Marques de Mattos de Araujo, Allan Abuabara, Maria Eduarda Cavassin, Maria Luisa Giacon
article en

Abstract

Abstract Background Pulp stones are calcified structures within the dental pulp that may increase the complexity of endodontic treatment and are commonly identified through radiographic examinations. Despite recent advances in artificial intelligence applied to dental imaging, the diagnostic capability of publicly available multimodal large language models (LLMs) for detecting pulp stones under zero-shot conditions remains unclear. Therefore, this study evaluated the diagnostic performance and response consistency of multimodal LLMs in detecting pulp stones on panoramic radiographs of posterior teeth. Methods This comparative and exploratory study evaluated the GPT-5 and Gemini 2.5 Flash multimodal LLMs using panoramic radiograph image crops containing posterior teeth. A total of 64 image crops, comprising 192 posterior teeth, were analyzed. The models were tested under standardized zero-shot conditions using independent sessions and a standardized English-language prompt. Each image was submitted six times to both models to assess response consistency. Diagnostic performance was evaluated using accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1-score with 95% confidence intervals. Intra-model consistency was assessed using majority agreement, complete agreement across repetitions, and Fleiss’ Kappa coefficient. Results The models demonstrated similar overall accuracy for binary detection of pulp stones, with 56.5% for ChatGPT and 57.0% for Gemini. ChatGPT showed high specificity (96.2%) and greater intra-model stability, whereas Gemini demonstrated higher sensitivity (81.8%) and superior performance in identifying the affected tooth and dental group. However, Gemini also presented lower specificity and greater variability across repeated evaluations. Fleiss’ Kappa coefficients indicated low intra-model agreement beyond chance for both models, particularly for Gemini. Conclusion Multimodal LLMs demonstrated limited capability in detecting pulp stones on panoramic radiographs, with distinct diagnostic profiles between the evaluated models. Although the findings suggest potential future applications of LLMs as auxiliary tools for dental image interpretation, their current performance does not yet support consistent clinical applicability for pulp stone detection.

BMC Oral Health
Openalex Percentile: Top 9%
Dental Radiography and Imaging
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.