Diagnostic performance of multimodal large language models in detecting pulp stones on panoramic radiographs
Abstract Background Pulp stones are calcified structures within the dental pulp that may increase the complexity of endodontic treatment and are commonly identified through radiographic examinations. Despite recent advances in artificial intelligence applied to dental imaging, the diagnostic capability of publicly available multimodal large language models (LLMs) for detecting pulp stones under zero-shot conditions remains unclear. Therefore, this study evaluated the diagnostic performance and response consistency of multimodal LLMs in detecting pulp stones on panoramic radiographs of posterior teeth. Methods This comparative and exploratory study evaluated the GPT-5 and Gemini 2.5 Flash multimodal LLMs using panoramic radiograph image crops containing posterior teeth. A total of 64 image crops, comprising 192 posterior teeth, were analyzed. The models were tested under standardized zero-shot conditions using independent sessions and a standardized English-language prompt. Each image was submitted six times to both models to assess response consistency. Diagnostic performance was evaluated using accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1-score with 95% confidence intervals. Intra-model consistency was assessed using majority agreement, complete agreement across repetitions, and Fleiss’ Kappa coefficient. Results The models demonstrated similar overall accuracy for binary detection of pulp stones, with 56.5% for ChatGPT and 57.0% for Gemini. ChatGPT showed high specificity (96.2%) and greater intra-model stability, whereas Gemini demonstrated higher sensitivity (81.8%) and superior performance in identifying the affected tooth and dental group. However, Gemini also presented lower specificity and greater variability across repeated evaluations. Fleiss’ Kappa coefficients indicated low intra-model agreement beyond chance for both models, particularly for Gemini. Conclusion Multimodal LLMs demonstrated limited capability in detecting pulp stones on panoramic radiographs, with distinct diagnostic profiles between the evaluated models. Although the findings suggest potential future applications of LLMs as auxiliary tools for dental image interpretation, their current performance does not yet support consistent clinical applicability for pulp stone detection.
Authors
- Flares Baratto‐Filho (ORCID: https://orcid.org/0000-0002-5649-7234)
- Michelle Nascimento Meger (ORCID: https://orcid.org/0000-0002-1776-2373)
- Luíz Fernando Fariniuk (ORCID: https://orcid.org/0000-0003-0731-9893)
- Érika Calvano Küchler (ORCID: https://orcid.org/0000-0001-5351-2526)
- Rafaela Scariot (ORCID: https://orcid.org/0000-0002-4911-6413)
- Christian Kirschneck (ORCID: https://orcid.org/0000-0001-9473-8724)
- Bianca Marques de Mattos de Araujo (ORCID: https://orcid.org/0000-0002-7507-4667)
- Allan Abuabara (ORCID: https://orcid.org/0000-0003-2454-3360)
- Maria Eduarda Cavassin (ORCID: https://orcid.org/0009-0003-5923-1380)
- Maria Luisa Giacon
Publication Details
- Journal
- BMC Oral Health
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1186/s12903-026-10119-6
- Primary Topic
- Dental Radiography and Imaging
- Type
- article
- Field-Weighted Citation Impact
- 0.00