Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study

Background/Objectives: Large language model (LLM) chatbots are increasingly used for dental information and decision support, yet their accuracy and short-term reproducibility in endodontics remain insufficiently established. This study compared five chatbots using open-ended questions derived from established AAE and ESE endodontic guidelines. Methods: Twenty-six guideline-based questions were content-validated by five endodontists using Lawshe’s Content Validity Ratio. ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, ScholarGPT (academic version built on OpenAI’s GPT-4 architecture), and MedGebra GPT-4 answered each question across three days and three sessions per day, yielding 1170 responses. Two blinded endodontists scored responses on a 5-point guideline-concordance scale. Brunner–Langer LD-F2 analyses assessed model and temporal effects, while weighted kappa evaluated response consistency. Results: The overall model effect was significant (p < 0.001). ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, and ScholarGPT showed comparable performance, whereas MedGebra GPT-4 performed significantly lower after Bonferroni correction. The day effect was not significant (p = 0.054), while the session effect (p = 0.048) and model × session interaction (p = 0.030) were significant. Weighted kappa values varied across models and assessment days, ranging from 0.689–0.730 for Gemini 2.5 Pro, 0.606–0.662 for ChatGPT-4o, 0.520–0.645 for DeepSeek-V3-0324, 0.458–0.592 for ScholarGPT, and 0.240–0.739 for MedGebra GPT-4. Conclusions: Guideline-aligned performance and short-term reproducibility differed across the evaluated chatbots, showing that accuracy and consistency represent distinct aspects of performance. Repeated assessment captured variation missed by single-session testing. These findings support guideline-based evaluation and clinician verification when generative AI chatbots are used to provide endodontic information relevant to decision support or dental education.

Authors

Institutions

Publication Details

Journal
Dentistry Journal
Published
2026-09-16
DOI
https://doi.org/10.3390/dj14090598
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study

Ilgın Akçay, Timur Köse, Ezgi Avcı
Dentistry Journal
Artificial Intelligence in Healthcare and Education
article

Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study

Ilgın Akçay, Timur Köse, Ezgi Avcı
article en

Abstract

Background/Objectives: Large language model (LLM) chatbots are increasingly used for dental information and decision support, yet their accuracy and short-term reproducibility in endodontics remain insufficiently established. This study compared five chatbots using open-ended questions derived from established AAE and ESE endodontic guidelines. Methods: Twenty-six guideline-based questions were content-validated by five endodontists using Lawshe’s Content Validity Ratio. ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, ScholarGPT (academic version built on OpenAI’s GPT-4 architecture), and MedGebra GPT-4 answered each question across three days and three sessions per day, yielding 1170 responses. Two blinded endodontists scored responses on a 5-point guideline-concordance scale. Brunner–Langer LD-F2 analyses assessed model and temporal effects, while weighted kappa evaluated response consistency. Results: The overall model effect was significant (p < 0.001). ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, and ScholarGPT showed comparable performance, whereas MedGebra GPT-4 performed significantly lower after Bonferroni correction. The day effect was not significant (p = 0.054), while the session effect (p = 0.048) and model × session interaction (p = 0.030) were significant. Weighted kappa values varied across models and assessment days, ranging from 0.689–0.730 for Gemini 2.5 Pro, 0.606–0.662 for ChatGPT-4o, 0.520–0.645 for DeepSeek-V3-0324, 0.458–0.592 for ScholarGPT, and 0.240–0.739 for MedGebra GPT-4. Conclusions: Guideline-aligned performance and short-term reproducibility differed across the evaluated chatbots, showing that accuracy and consistency represent distinct aspects of performance. Repeated assessment captured variation missed by single-session testing. These findings support guideline-based evaluation and clinician verification when generative AI chatbots are used to provide endodontic information relevant to decision support or dental education.

Dentistry JournalVol. 14(9)
Ege University (TR)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.