Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study

Large language models (LLMs) are increasingly being explored for clinical and educational applications in dentistry. However, their ability to generate comprehensive fixed prosthodontic treatment plans has not been adequately compared under standardized conditions. This study aimed to compare the quality of prosthodontic treatment plans generated by Qwen 3.7 Plus (Qwen), Claude Sonnet 5 (Claude), and Mistral Medium 3.5 (Mistral). Thirty standardized prosthodontic clinical scenarios were presented to each LLM, yielding 90 treatment plans. Two blinded prosthodontic evaluators independently assessed each plan with a seven-domain, five-point Likert-scale rubric covering diagnostic accuracy, treatment selection, treatment sequencing, completeness, reference adherence, patient safety, and overall clinical appropriateness. Evaluator scores were averaged for the primary analysis. Differences among models were assessed with the Friedman test, followed by paired Wilcoxon signed-rank tests with Holm adjustment. Inter-rater agreement was assessed with quadratic-weighted Cohen’s κ and a two-way mixed-effects ICC (single- and average-measure). A sensitivity analysis evaluated whether model rankings remained consistent when each evaluator was analyzed separately. In the primary analysis of averaged evaluator scores, Qwen had the highest mean total score (33.82 ± 1.74), followed by Claude (32.52 ± 2.30) and Mistral (29.68 ± 2.66) (Friedman χ² (2) = 35.05, p < 0.001; Kendall’s W = 0.584). All three pairwise comparisons were significant after Holm adjustment. Across individual domains, Qwen scored highest for completeness and reference adherence, while Claude achieved the highest patient safety scores. However, inter-rater agreement was poor, with a single-measure ICC of 0.378 and an average-measure ICC of 0.549. Sensitivity analysis revealed that the model ranking was not consistent across evaluators: Evaluator 1 ranked Claude marginally higher than Qwen, whereas Evaluator 2 ranked Qwen clearly higher than Claude. This discrepancy was associated with a pronounced ceiling effect in Evaluator 1’s ratings, who assigned the maximum score to the majority of Qwen and Claude cases. Qwen achieved the highest mean score in the primary analysis, with particular strengths in completeness and reference adherence, while Claude showed the highest patient safety scores. However, the model ranking was not consistent across evaluators, and Mistral had the lowest mean score across all seven domains. Ceiling effects and limited inter-rater agreement were observed, and these findings do not support universal superiority of any single model. From a dental education perspective, LLMs may have potential as adjunctive tools for case-based learning and critical appraisal of treatment-planning decisions, but educational effectiveness was not directly assessed in this study.

Authors

Institutions

Publication Details

Journal
BMC Medical Education
Published
2026-10-09
DOI
https://doi.org/10.1186/s12909-026-10560-9
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study

Heba Wageh Abozaed, Mohammad Alokla, Shahad Albader, Ahmed Algohar et al.
BMC Medical Education
Artificial Intelligence in Healthcare and Education
article

Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study

Heba Wageh Abozaed, Mohammad Alokla, Shahad Albader, Ahmed Algohar, Mohammed S. Murayshed, Rafif Alshenaiber
article en

Abstract

Large language models (LLMs) are increasingly being explored for clinical and educational applications in dentistry. However, their ability to generate comprehensive fixed prosthodontic treatment plans has not been adequately compared under standardized conditions. This study aimed to compare the quality of prosthodontic treatment plans generated by Qwen 3.7 Plus (Qwen), Claude Sonnet 5 (Claude), and Mistral Medium 3.5 (Mistral). Thirty standardized prosthodontic clinical scenarios were presented to each LLM, yielding 90 treatment plans. Two blinded prosthodontic evaluators independently assessed each plan with a seven-domain, five-point Likert-scale rubric covering diagnostic accuracy, treatment selection, treatment sequencing, completeness, reference adherence, patient safety, and overall clinical appropriateness. Evaluator scores were averaged for the primary analysis. Differences among models were assessed with the Friedman test, followed by paired Wilcoxon signed-rank tests with Holm adjustment. Inter-rater agreement was assessed with quadratic-weighted Cohen’s κ and a two-way mixed-effects ICC (single- and average-measure). A sensitivity analysis evaluated whether model rankings remained consistent when each evaluator was analyzed separately. In the primary analysis of averaged evaluator scores, Qwen had the highest mean total score (33.82 ± 1.74), followed by Claude (32.52 ± 2.30) and Mistral (29.68 ± 2.66) (Friedman χ² (2) = 35.05, p < 0.001; Kendall’s W = 0.584). All three pairwise comparisons were significant after Holm adjustment. Across individual domains, Qwen scored highest for completeness and reference adherence, while Claude achieved the highest patient safety scores. However, inter-rater agreement was poor, with a single-measure ICC of 0.378 and an average-measure ICC of 0.549. Sensitivity analysis revealed that the model ranking was not consistent across evaluators: Evaluator 1 ranked Claude marginally higher than Qwen, whereas Evaluator 2 ranked Qwen clearly higher than Claude. This discrepancy was associated with a pronounced ceiling effect in Evaluator 1’s ratings, who assigned the maximum score to the majority of Qwen and Claude cases. Qwen achieved the highest mean score in the primary analysis, with particular strengths in completeness and reference adherence, while Claude showed the highest patient safety scores. However, the model ranking was not consistent across evaluators, and Mistral had the lowest mean score across all seven domains. Ceiling effects and limited inter-rater agreement were observed, and these findings do not support universal superiority of any single model. From a dental education perspective, LLMs may have potential as adjunctive tools for case-based learning and critical appraisal of treatment-planning decisions, but educational effectiveness was not directly assessed in this study.

BMC Medical Education
Prince Sattam Bin Abdulaziz University (SA), Mansoura University (EG)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.