Evaluating the Reliability, Accuracy, and Readability of ChatGPT-3.5 and GPT-4 in Providing Patient Education Related to Glenohumeral Joint Osteoarthritis.
OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions. METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4. RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses. CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.
Authors
- Adam Z. Khan (ORCID: https://orcid.org/0000-0002-7251-6503)
- Brian W. Hill (ORCID: https://orcid.org/0000-0001-8403-8265)
- Mohammad Daher (ORCID: https://orcid.org/0000-0002-9256-9952)
- Mohamad Y. Fares (ORCID: https://orcid.org/0000-0001-8228-3953)
- Peter Boufadel (ORCID: https://orcid.org/0000-0001-7078-9148)
- Joseph Albert Abboud (ORCID: https://orcid.org/0000-0002-3845-7220)
- John G. Horneff (ORCID: https://orcid.org/0000-0003-4802-9913)
- Tarishi Parmar (ORCID: https://orcid.org/0000-0002-5139-6282)
- Jonathan Berg
Institutions
- Pennsylvania State University (US)
- Kaiser Permanente (US)
- Thomas Jefferson University (US)
- Brown University (US)
- Rothman Orthopaedics (US)
- Rothman Institute (US)
- University of Pennsylvania (US)
Publication Details
- Journal
- PubMed
- Published
- 2026-10-01
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00