Reliability, quality, and readability of large language models in carpal tunnel syndrome: A comparative cross-linguistic, cross-sectional observational study

Background AI-based chatbots are increasingly used for health information, while cross-linguistic studies suggest that response performance may vary by language. Given the high prevalence of carpal tunnel syndrome (CTS), this study evaluated the reliability, quality, patient education suitability, and readability of AI-generated responses to CTS questions in English and Turkish. Materials and methods Seventeen patient-centered CTS questions were directed to ChatGPT ® (GPT-5 Mini), Gemini ® 3 Flash, and Claude ® 4.5 Sonnet, using zero-shot prompting across 102 independent sessions in English and Turkish. Responses were evaluated for reliability, quality, and patient education using mDISCERN (MD), Global Quality Score (GQS), and the Patient Education Materials Assessment Tool (PEMAT). Readability was assessed via Flesch Reading Ease (English) and Ateşman Index (Turkish). Friedman and Wilcoxon tests with Bonferroni correction were applied. Results Significant inter-model differences were observed in English across all metrics ( p < 0.05), with Gemini achieving higher MD and PEMAT scores than ChatGPT and Claude ( p < 0.01). Cross-language differences were observed in mDISCERN scores, while Gemini showing a significant difference in favor of English (r = 0.93; p < 0.001). None of the models met recommended readability levels in either language. Conclusion AI-generated CTS information varied by model and language. Differences in MD scores should be interpreted cautiously because this instrument partly reflects the provision of references and does not directly assess factual clinical accuracy. None of the evaluated models achieved recommended readability levels for patient education. These findings support the use of AIs as supplementary rather than standalone sources of CTS-related patient information.

Authors

Institutions

Publication Details

Journal
Journal of Back and Musculoskeletal Rehabilitation
Published
2026-10-08
DOI
https://doi.org/10.1177/10538127261492778
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Reliability, quality, and readability of large language models in carpal tunnel syndrome: A comparative cross-linguistic, cross-sectional observational study

Bilgehan Kolutek Ay, Cem Zafer Yıldır
Journal of Back and Musculoskeletal Rehabilitation
Artificial Intelligence in Healthcare and Education
article

Reliability, quality, and readability of large language models in carpal tunnel syndrome: A comparative cross-linguistic, cross-sectional observational study

Bilgehan Kolutek Ay, Cem Zafer Yıldır
article en

Abstract

Background AI-based chatbots are increasingly used for health information, while cross-linguistic studies suggest that response performance may vary by language. Given the high prevalence of carpal tunnel syndrome (CTS), this study evaluated the reliability, quality, patient education suitability, and readability of AI-generated responses to CTS questions in English and Turkish. Materials and methods Seventeen patient-centered CTS questions were directed to ChatGPT ® (GPT-5 Mini), Gemini ® 3 Flash, and Claude ® 4.5 Sonnet, using zero-shot prompting across 102 independent sessions in English and Turkish. Responses were evaluated for reliability, quality, and patient education using mDISCERN (MD), Global Quality Score (GQS), and the Patient Education Materials Assessment Tool (PEMAT). Readability was assessed via Flesch Reading Ease (English) and Ateşman Index (Turkish). Friedman and Wilcoxon tests with Bonferroni correction were applied. Results Significant inter-model differences were observed in English across all metrics ( p < 0.05), with Gemini achieving higher MD and PEMAT scores than ChatGPT and Claude ( p < 0.01). Cross-language differences were observed in mDISCERN scores, while Gemini showing a significant difference in favor of English (r = 0.93; p < 0.001). None of the models met recommended readability levels in either language. Conclusion AI-generated CTS information varied by model and language. Differences in MD scores should be interpreted cautiously because this instrument partly reflects the provision of references and does not directly assess factual clinical accuracy. None of the evaluated models achieved recommended readability levels for patient education. These findings support the use of AIs as supplementary rather than standalone sources of CTS-related patient information.

Journal of Back and Musculoskeletal Rehabilitation
Imam Mohammad ibn Saud Islamic University (SA)
Openalex Percentile: Top 20%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.