Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements

Objectives: This study aimed to compare the performance of DeepSeek-V3, ChatGPT 4.0, ChatGPT 4.5, Gemini Advanced 2.0, and Grok in answering single-correct answer multiple-choice (SCQ), multiple-correct answer multiple-choice (MCQ), binary (true/false) (BQ), and open-ended questions (OEQ) based on the European Society of Endodontology (ESE) position statements.Materials and Methods: A total of 70 questions (24 SCQ/MCQ, 29 BQ, and 17 OEQ) were posed to five LLMs. Responses were scored on a 0–10 scale for comprehensiveness, scientific accuracy, clarity, and relevance. SCQ, MCQ, and BQ formats were additionally evaluated for correctness (1 = correct, 0 = incorrect). Data were analyzed using Pearson’s chi-square test and the Kruskal-Wallis H test (p < 0.05).Results: In the SCQ/MCQ format, Gemini Advanced 2.0 scored significantly higher than Grok. In the BQ format, ChatGPT 4.5 and Gemini Advanced 2.0 achieved significantly higher scores than ChatGPT 4.0 and Grok. In the OEQ format, ChatGPT 4.5 obtained significantly higher scores than Grok (p < 0.05). Accuracy rates were 67.9% for ChatGPT 4.5, 66.0% for Gemini Advanced 2.0, 58.5% for DeepSeek-V3, 52.8% for ChatGPT 4.0, and 49.1% for Grok (p > 0.05). In all models, scientific accuracy scores were significantly lower than relevance scores (p < 0.05).Conclusions: ChatGPT 4.5 and Gemini Advanced 2.0 demonstrated stronger performance than the other models, whereas Grok showed lower response-content performance. However, since LLMs may appear convincing even when generating incorrect clinical information, they should be used only as supplementary tools in endodontic education and preliminary information retrieval. Clinically relevant outputs should be verified by an endodontist or an appropriately qualified clinician.

Authors

Institutions

Publication Details

Journal
Cumhuriyet Dental Journal
Published
2026-09-30
DOI
https://doi.org/10.7126/cumudj.1953387
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements

Merve Yeniçeri Özata, Ayşenur Çam, Merve Defişet
Cumhuriyet Dental Journal
Artificial Intelligence in Healthcare and Education
article

Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements

Merve Yeniçeri Özata, Ayşenur Çam, Merve Defişet
article en

Abstract

Objectives: This study aimed to compare the performance of DeepSeek-V3, ChatGPT 4.0, ChatGPT 4.5, Gemini Advanced 2.0, and Grok in answering single-correct answer multiple-choice (SCQ), multiple-correct answer multiple-choice (MCQ), binary (true/false) (BQ), and open-ended questions (OEQ) based on the European Society of Endodontology (ESE) position statements.Materials and Methods: A total of 70 questions (24 SCQ/MCQ, 29 BQ, and 17 OEQ) were posed to five LLMs. Responses were scored on a 0–10 scale for comprehensiveness, scientific accuracy, clarity, and relevance. SCQ, MCQ, and BQ formats were additionally evaluated for correctness (1 = correct, 0 = incorrect). Data were analyzed using Pearson’s chi-square test and the Kruskal-Wallis H test (p < 0.05).Results: In the SCQ/MCQ format, Gemini Advanced 2.0 scored significantly higher than Grok. In the BQ format, ChatGPT 4.5 and Gemini Advanced 2.0 achieved significantly higher scores than ChatGPT 4.0 and Grok. In the OEQ format, ChatGPT 4.5 obtained significantly higher scores than Grok (p < 0.05). Accuracy rates were 67.9% for ChatGPT 4.5, 66.0% for Gemini Advanced 2.0, 58.5% for DeepSeek-V3, 52.8% for ChatGPT 4.0, and 49.1% for Grok (p > 0.05). In all models, scientific accuracy scores were significantly lower than relevance scores (p < 0.05).Conclusions: ChatGPT 4.5 and Gemini Advanced 2.0 demonstrated stronger performance than the other models, whereas Grok showed lower response-content performance. However, since LLMs may appear convincing even when generating incorrect clinical information, they should be used only as supplementary tools in endodontic education and preliminary information retrieval. Clinically relevant outputs should be verified by an endodontist or an appropriately qualified clinician.

Cumhuriyet Dental JournalVol. 29(3)
Dicle University (TR)
Quality Education
Openalex Percentile: Top 16%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.