Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study
Large language models (LLMs) are increasingly being explored for clinical and educational applications in dentistry. However, their ability to generate comprehensive fixed prosthodontic treatment plans has not been adequately compared under standardized conditions. This study aimed to compare the quality of prosthodontic treatment plans generated by Qwen 3.7 Plus (Qwen), Claude Sonnet 5 (Claude), and Mistral Medium 3.5 (Mistral). Thirty standardized prosthodontic clinical scenarios were presented to each LLM, yielding 90 treatment plans. Two blinded prosthodontic evaluators independently assessed each plan with a seven-domain, five-point Likert-scale rubric covering diagnostic accuracy, treatment selection, treatment sequencing, completeness, reference adherence, patient safety, and overall clinical appropriateness. Evaluator scores were averaged for the primary analysis. Differences among models were assessed with the Friedman test, followed by paired Wilcoxon signed-rank tests with Holm adjustment. Inter-rater agreement was assessed with quadratic-weighted Cohen’s κ and a two-way mixed-effects ICC (single- and average-measure). A sensitivity analysis evaluated whether model rankings remained consistent when each evaluator was analyzed separately. In the primary analysis of averaged evaluator scores, Qwen had the highest mean total score (33.82 ± 1.74), followed by Claude (32.52 ± 2.30) and Mistral (29.68 ± 2.66) (Friedman χ² (2) = 35.05, p < 0.001; Kendall’s W = 0.584). All three pairwise comparisons were significant after Holm adjustment. Across individual domains, Qwen scored highest for completeness and reference adherence, while Claude achieved the highest patient safety scores. However, inter-rater agreement was poor, with a single-measure ICC of 0.378 and an average-measure ICC of 0.549. Sensitivity analysis revealed that the model ranking was not consistent across evaluators: Evaluator 1 ranked Claude marginally higher than Qwen, whereas Evaluator 2 ranked Qwen clearly higher than Claude. This discrepancy was associated with a pronounced ceiling effect in Evaluator 1’s ratings, who assigned the maximum score to the majority of Qwen and Claude cases. Qwen achieved the highest mean score in the primary analysis, with particular strengths in completeness and reference adherence, while Claude showed the highest patient safety scores. However, the model ranking was not consistent across evaluators, and Mistral had the lowest mean score across all seven domains. Ceiling effects and limited inter-rater agreement were observed, and these findings do not support universal superiority of any single model. From a dental education perspective, LLMs may have potential as adjunctive tools for case-based learning and critical appraisal of treatment-planning decisions, but educational effectiveness was not directly assessed in this study.
Authors
- Heba Wageh Abozaed (ORCID: https://orcid.org/0000-0002-3216-7361)
- Mohammad Alokla
- Shahad Albader (ORCID: https://orcid.org/0009-0008-1588-6330)
- Ahmed Algohar
- Mohammed S. Murayshed (ORCID: https://orcid.org/0000-0003-3843-2936)
- Rafif Alshenaiber
Institutions
- Prince Sattam Bin Abdulaziz University (SA)
- Mansoura University (EG)
Publication Details
- Journal
- BMC Medical Education
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1186/s12909-026-10560-9
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00