Comparative evaluation of large language models for patient-facing information in clear aligner therapy

Objectives: Patients considering clear aligner therapy should understand treatment requirements, limitations, and the importance of professional supervision. As many patients now seek orthodontic information from artificial intelligence (AI) chatbots, this study evaluated how effectively these systems communicate patient-facing information related to clear aligner therapy. Material and Methods: Twenty common patient-facing questions about clear aligner therapy were developed using conversational language. Three large language models (LLM), ChatGPT 5.2, Claude Sonnet 4.5, and Gemini Flash 3, generated responses using default settings. Three orthodontic specialists, blinded to model identity, rated responses for quality and empathy using 5-point Likert scales. Readability was assessed using the Flesch–Kincaid Grade Level. Inter-rater reliability was evaluated using intraclass correlation coefficients (2,1). Model performance was compared using repeated-measures analysis of variance with Bonferroni-adjusted post hoc testing. Results: Inter-rater reliability ranged from fair to good. Significant differences were observed among models for quality (F[2,38] = 11.43; p < 0.001) and empathy (F[2,38] = 3.92; p = 0.028), while readability did not differ significantly ( p = 0.321). Gemini Flash 3 achieved the highest mean scores for both quality (4.57 ± 0.42) and empathy (4.28 ± 0.51). All models consistently encouraged professional orthodontic consultation and avoided recommending unsupervised treatment. Conclusion: Large language models can provide appropriate patient-facing information about clear aligner therapy, though meaningful differences in communication quality and empathy exist. Clinical oversight remains essential when integrating AI tools into orthodontic patient education.

Authors

Institutions

Publication Details

Journal
APOS Trends in Orthodontics
Published
2026-09-15
DOI
https://doi.org/10.25259/apos_146_2026
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative evaluation of large language models for patient-facing information in clear aligner therapy

Prasun Mukhopadhyay, Raju Biswas, Alka Mukhopadhyay, Santanu Mukhopadhyay et al.
APOS Trends in Orthodontics
Artificial Intelligence in Healthcare and Education
article

Comparative evaluation of large language models for patient-facing information in clear aligner therapy

Prasun Mukhopadhyay, Raju Biswas, Alka Mukhopadhyay, Santanu Mukhopadhyay, Sayani Adhikari
article en

Abstract

Objectives: Patients considering clear aligner therapy should understand treatment requirements, limitations, and the importance of professional supervision. As many patients now seek orthodontic information from artificial intelligence (AI) chatbots, this study evaluated how effectively these systems communicate patient-facing information related to clear aligner therapy. Material and Methods: Twenty common patient-facing questions about clear aligner therapy were developed using conversational language. Three large language models (LLM), ChatGPT 5.2, Claude Sonnet 4.5, and Gemini Flash 3, generated responses using default settings. Three orthodontic specialists, blinded to model identity, rated responses for quality and empathy using 5-point Likert scales. Readability was assessed using the Flesch–Kincaid Grade Level. Inter-rater reliability was evaluated using intraclass correlation coefficients (2,1). Model performance was compared using repeated-measures analysis of variance with Bonferroni-adjusted post hoc testing. Results: Inter-rater reliability ranged from fair to good. Significant differences were observed among models for quality (F[2,38] = 11.43; p < 0.001) and empathy (F[2,38] = 3.92; p = 0.028), while readability did not differ significantly ( p = 0.321). Gemini Flash 3 achieved the highest mean scores for both quality (4.57 ± 0.42) and empathy (4.28 ± 0.51). All models consistently encouraged professional orthodontic consultation and avoided recommending unsupervised treatment. Conclusion: Large language models can provide appropriate patient-facing information about clear aligner therapy, though meaningful differences in communication quality and empathy exist. Clinical oversight remains essential when integrating AI tools into orthodontic patient education.

APOS Trends in OrthodonticsVol. 0
Amazon (United States) (US), Dr. R. Ahmed Dental College and Hospital (IN), Malda Medical College and Hospital (IN), Southern Command Hospital (IN)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.