Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support.
Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate generative AI in THA/TKA across 3 clinical domains. Methods: A PubMed and Embase systematic literature review was performed on July 9, 2025, in accordance with the Preffered Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Included studies evaluated generative AI use in THA/TKA and addressed 1 of our 3 domains. Excluded studies used nongenerative AI, involved populations not undergoing THA/TKA, or were non-English or non-peer-reviewed. Quality metrics that were assessed included blinded clinician ratings, readability scores, DISCERN scores, and diagnostic accuracy measures. The heterogeneity of the included studies led to a narrative synthesis, and no formal risk of bias was conducted. Results: Of the 91 articles retrieved, 23 met the inclusion criteria. ChatGPT versions 3.5 or 4 were assessed across all studies, and 3 studies included Google Gemini, Claude 3 Opus, and DeepSeek. Among studies in patient communication and education (n = 19), blinded clinician ratings showed that ChatGPT-generated responses to frequently asked questions (FAQ's) were as accurate, clear, and complete as surgeon-written responses. Regarding documentation (n = 2), LLMs demonstrated 97.5% to 100% accuracy in identifying operative report information and created patient consent documents with improved readability scores (grade level 12.6 vs. 16.8) and completeness (2.4/3 vs. 1.8/3). Among studies assessing decision support (n = 2), ChatGPT demonstrated high sensitivity when predicting surgical candidacy but low specificity when predicting surgical outcomes. Conclusion: LLMs performed well in several aspects of patient communication and education with support from high clinician ratings, accuracy, and readability scores. However, evidence for documentation and decision support remained limited. Moreover, proper implementation that considers clinical workflows could increase clinician efficiency and positively affect patient experience.
Authors
- Charles M. Lawrie
- Ivan A. Garces (ORCID: https://orcid.org/0009-0005-4147-2674)
- Andrew McDaid
- Andres G. Wong
- Jonathan A. Brutti
- Stefano A Bini
Institutions
- University of Auckland (NZ)
- University of California, San Francisco (US)
- Florida International University (US)
- Auckland University of Technology (NZ)
- Baptist Health South Florida (US)
Publication Details
- Journal
- PubMed
- Published
- 2026-09-18
- DOI
- https://doi.org/10.2106/jbjs.oa.26.00077
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00