Patient-Perceived and Record-Based Concordance of AI-Generated Perioperative Information After Total Hip Arthroplasty: A Cross-Sectional Study with Multidisciplinary Review
Background and Objectives: Large language models, a form of artificial intelligence (AI), can answer common patient questions about total hip arthroplasty (THA), but most evaluations emphasize clinician-rated accuracy or patient preference rather than whether generalized recovery statements correspond to patients’ documented postoperative courses. We evaluated a fixed set of ChatGPT 5.1 Thinking responses to 12 investigator-selected questions about elective primary THA from patient, medical-record, and multidisciplinary clinical perspectives, with particular attention to complicated recovery. Materials and Methods: In this cross-sectional study with retrospective record review, 101 adults at least 6 months after elective primary THA reviewed a fixed Turkish-language set of ChatGPT 5.1 Thinking answers to 12 investigator-selected questions. Patients rated overall experiential concordance, retrospectively perceived preoperative usefulness, and clarity (0–10). Medical records provided operative duration, first mobilization day, length of stay, and 90-day complications. Three prespecified chart-verifiable timing statements formed a 0–3 timing concordance score. Six multidisciplinary clinicians independently rated all answers during separate 30 min digital assessment sessions. Results: Mean overall patient-perceived concordance was 8.44 ± 0.95. Fourteen patients (13.9%) had a complicated recovery and reported lower concordance than those with uncomplicated recovery (7.57 ± 1.22 vs. 8.57 ± 0.83; p = 0.0026). The stated ranges for operative duration, early mobilization, and length of stay matched the corresponding records in 93.1%, 94.1%, and 91.1% of patients, respectively; 85.1% met all three timing criteria. Complete three-criterion timing concordance was lower after complicated recovery (28.6% vs. 94.3%; p < 0.001). Patient-perceived concordance correlated with the timing concordance score (Spearman ρ = 0.395; p < 0.001). Mean multidisciplinary appropriateness across 72 ratings was 8.50 ± 0.50. Conclusions: The fixed AI-generated THA information set was rated favorably by patients and clinicians, and its three prespecified timing statements corresponded closely to routine documented recovery. Both experiential and timing concordance were lower after complicated courses. These findings support clinician-reviewed AI information as an adjunct while underscoring that generalized timelines are conditional ranges, not individualized predictions.
Authors
- Oktay Adanır (ORCID: https://orcid.org/0000-0003-3327-4705)
- Ozancan Biçer (ORCID: https://orcid.org/0000-0002-6080-177X)
- Mehmet Yağız Yenigün (ORCID: https://orcid.org/0000-0002-9073-0124)
- Cemre Aydın (ORCID: https://orcid.org/0000-0003-4169-7919)
- Ali Tarık Kutlu
- Sidar Güneş
Institutions
- Bağcılar Eğitim ve Araştırma Hastanesi (TR)
- Sağlık Bilimleri Üniversitesi (TR)
Publication Details
- Journal
- Medicina
- Published
- 2026-09-27
- DOI
- https://doi.org/10.3390/medicina62101873
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00