Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety

Abstract This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categories were submitted once to each model, producing 100 first-pass responses. C-PEMAT and GQS assessed educational quality, seven English readability formulas were applied to consensus translations of the Chinese outputs, and two preliminary clinician-developed measures assessed clinical key-point coverage and safety-related caution. C-PEMAT and GQS differed significantly among models (both P < 0.001), with GPT-5 obtaining the highest scores, followed by DeepSeek and Doubao. Quality scores did not differ significantly across content categories, whereas all readability indices did. C-PEMAT and GQS were moderately correlated ( r = 0.68), but quality measures were weakly associated with most readability indices. GPT-5 and DeepSeek showed relatively favorable Clinical Intent Alignment scores, while GPT-5 had the highest descriptive Clinical Safety Score. These findings are specific to the tested access date, platforms, prompts, and model versions. LLM-generated materials may serve as auxiliary drafts, but clinician review remains necessary for medical accuracy, safety, readability, and patient-specific appropriateness.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-05
DOI
https://doi.org/10.1038/s41598-026-66627-6
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety

Huanan Li, Jiachen Liang, Wanying Gui
Scientific Reports
Artificial Intelligence in Healthcare and Education
article

Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety

Huanan Li, Jiachen Liang, Wanying Gui
article en

Abstract

Abstract This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categories were submitted once to each model, producing 100 first-pass responses. C-PEMAT and GQS assessed educational quality, seven English readability formulas were applied to consensus translations of the Chinese outputs, and two preliminary clinician-developed measures assessed clinical key-point coverage and safety-related caution. C-PEMAT and GQS differed significantly among models (both P < 0.001), with GPT-5 obtaining the highest scores, followed by DeepSeek and Doubao. Quality scores did not differ significantly across content categories, whereas all readability indices did. C-PEMAT and GQS were moderately correlated ( r = 0.68), but quality measures were weakly associated with most readability indices. GPT-5 and DeepSeek showed relatively favorable Clinical Intent Alignment scores, while GPT-5 had the highest descriptive Clinical Safety Score. These findings are specific to the tested access date, platforms, prompts, and model versions. LLM-generated materials may serve as auxiliary drafts, but clinician review remains necessary for medical accuracy, safety, readability, and patient-specific appropriateness.

Scientific ReportsVol. 16(1)
Tianjin University of Traditional Chinese Medicine (CN), First Teaching Hospital of Tianjin University of Traditional Chinese Medicine (CN)
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.