Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support.

Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate generative AI in THA/TKA across 3 clinical domains. Methods: A PubMed and Embase systematic literature review was performed on July 9, 2025, in accordance with the Preffered Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Included studies evaluated generative AI use in THA/TKA and addressed 1 of our 3 domains. Excluded studies used nongenerative AI, involved populations not undergoing THA/TKA, or were non-English or non-peer-reviewed. Quality metrics that were assessed included blinded clinician ratings, readability scores, DISCERN scores, and diagnostic accuracy measures. The heterogeneity of the included studies led to a narrative synthesis, and no formal risk of bias was conducted. Results: Of the 91 articles retrieved, 23 met the inclusion criteria. ChatGPT versions 3.5 or 4 were assessed across all studies, and 3 studies included Google Gemini, Claude 3 Opus, and DeepSeek. Among studies in patient communication and education (n = 19), blinded clinician ratings showed that ChatGPT-generated responses to frequently asked questions (FAQ's) were as accurate, clear, and complete as surgeon-written responses. Regarding documentation (n = 2), LLMs demonstrated 97.5% to 100% accuracy in identifying operative report information and created patient consent documents with improved readability scores (grade level 12.6 vs. 16.8) and completeness (2.4/3 vs. 1.8/3). Among studies assessing decision support (n = 2), ChatGPT demonstrated high sensitivity when predicting surgical candidacy but low specificity when predicting surgical outcomes. Conclusion: LLMs performed well in several aspects of patient communication and education with support from high clinician ratings, accuracy, and readability scores. However, evidence for documentation and decision support remained limited. Moreover, proper implementation that considers clinical workflows could increase clinician efficiency and positively affect patient experience.

Authors

Institutions

Publication Details

Journal
PubMed
Published
2026-09-18
DOI
https://doi.org/10.2106/jbjs.oa.26.00077
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support.

Charles M. Lawrie, Ivan A. Garces, Andrew McDaid, Andres G. Wong et al.
PubMed
Artificial Intelligence in Healthcare and Education
article

Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support.

Charles M. Lawrie, Ivan A. Garces, Andrew McDaid, Andres G. Wong, Jonathan A. Brutti, Stefano A Bini
article en

Abstract

Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate generative AI in THA/TKA across 3 clinical domains. Methods: A PubMed and Embase systematic literature review was performed on July 9, 2025, in accordance with the Preffered Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Included studies evaluated generative AI use in THA/TKA and addressed 1 of our 3 domains. Excluded studies used nongenerative AI, involved populations not undergoing THA/TKA, or were non-English or non-peer-reviewed. Quality metrics that were assessed included blinded clinician ratings, readability scores, DISCERN scores, and diagnostic accuracy measures. The heterogeneity of the included studies led to a narrative synthesis, and no formal risk of bias was conducted. Results: Of the 91 articles retrieved, 23 met the inclusion criteria. ChatGPT versions 3.5 or 4 were assessed across all studies, and 3 studies included Google Gemini, Claude 3 Opus, and DeepSeek. Among studies in patient communication and education (n = 19), blinded clinician ratings showed that ChatGPT-generated responses to frequently asked questions (FAQ's) were as accurate, clear, and complete as surgeon-written responses. Regarding documentation (n = 2), LLMs demonstrated 97.5% to 100% accuracy in identifying operative report information and created patient consent documents with improved readability scores (grade level 12.6 vs. 16.8) and completeness (2.4/3 vs. 1.8/3). Among studies assessing decision support (n = 2), ChatGPT demonstrated high sensitivity when predicting surgical candidacy but low specificity when predicting surgical outcomes. Conclusion: LLMs performed well in several aspects of patient communication and education with support from high clinician ratings, accuracy, and readability scores. However, evidence for documentation and decision support remained limited. Moreover, proper implementation that considers clinical workflows could increase clinician efficiency and positively affect patient experience.

PubMedVol. 11(3)
University of Auckland (NZ), University of California, San Francisco (US), Florida International University (US), Auckland University of Technology (NZ), Baptist Health South Florida (US)
Quality Education
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.