Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study
Background/Objectives: Large language model (LLM) chatbots are increasingly used for dental information and decision support, yet their accuracy and short-term reproducibility in endodontics remain insufficiently established. This study compared five chatbots using open-ended questions derived from established AAE and ESE endodontic guidelines. Methods: Twenty-six guideline-based questions were content-validated by five endodontists using Lawshe’s Content Validity Ratio. ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, ScholarGPT (academic version built on OpenAI’s GPT-4 architecture), and MedGebra GPT-4 answered each question across three days and three sessions per day, yielding 1170 responses. Two blinded endodontists scored responses on a 5-point guideline-concordance scale. Brunner–Langer LD-F2 analyses assessed model and temporal effects, while weighted kappa evaluated response consistency. Results: The overall model effect was significant (p < 0.001). ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, and ScholarGPT showed comparable performance, whereas MedGebra GPT-4 performed significantly lower after Bonferroni correction. The day effect was not significant (p = 0.054), while the session effect (p = 0.048) and model × session interaction (p = 0.030) were significant. Weighted kappa values varied across models and assessment days, ranging from 0.689–0.730 for Gemini 2.5 Pro, 0.606–0.662 for ChatGPT-4o, 0.520–0.645 for DeepSeek-V3-0324, 0.458–0.592 for ScholarGPT, and 0.240–0.739 for MedGebra GPT-4. Conclusions: Guideline-aligned performance and short-term reproducibility differed across the evaluated chatbots, showing that accuracy and consistency represent distinct aspects of performance. Repeated assessment captured variation missed by single-session testing. These findings support guideline-based evaluation and clinician verification when generative AI chatbots are used to provide endodontic information relevant to decision support or dental education.
Authors
- Ilgın Akçay (ORCID: https://orcid.org/0000-0001-7546-2048)
- Timur Köse
- Ezgi Avcı
Institutions
- Ege University (TR)
Publication Details
- Journal
- Dentistry Journal
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/dj14090598
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00