Evaluating the Reliability, Accuracy, and Readability of ChatGPT-3.5 and GPT-4 in Providing Patient Education Related to Glenohumeral Joint Osteoarthritis.

OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions. METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4. RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses. CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.

Authors

Institutions

Publication Details

Journal
PubMed
Published
2026-10-01
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating the Reliability, Accuracy, and Readability of ChatGPT-3.5 and GPT-4 in Providing Patient Education Related to Glenohumeral Joint Osteoarthritis.

Adam Z. Khan, Brian W. Hill, Mohammad Daher, Mohamad Y. Fares et al.
PubMed
Artificial Intelligence in Healthcare and Education
article

Evaluating the Reliability, Accuracy, and Readability of ChatGPT-3.5 and GPT-4 in Providing Patient Education Related to Glenohumeral Joint Osteoarthritis.

Adam Z. Khan, Brian W. Hill, Mohammad Daher, Mohamad Y. Fares, Peter Boufadel, Joseph Albert Abboud, John G. Horneff, Tarishi Parmar, Jonathan Berg
article en

Abstract

OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions. METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4. RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses. CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.

PubMedVol. 109(10)
Pennsylvania State University (US), Kaiser Permanente (US), Thomas Jefferson University (US), Brown University (US), Rothman Orthopaedics (US), Rothman Institute (US), University of Pennsylvania (US)
Quality Education
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.