Patient-Perceived and Record-Based Concordance of AI-Generated Perioperative Information After Total Hip Arthroplasty: A Cross-Sectional Study with Multidisciplinary Review

Background and Objectives: Large language models, a form of artificial intelligence (AI), can answer common patient questions about total hip arthroplasty (THA), but most evaluations emphasize clinician-rated accuracy or patient preference rather than whether generalized recovery statements correspond to patients’ documented postoperative courses. We evaluated a fixed set of ChatGPT 5.1 Thinking responses to 12 investigator-selected questions about elective primary THA from patient, medical-record, and multidisciplinary clinical perspectives, with particular attention to complicated recovery. Materials and Methods: In this cross-sectional study with retrospective record review, 101 adults at least 6 months after elective primary THA reviewed a fixed Turkish-language set of ChatGPT 5.1 Thinking answers to 12 investigator-selected questions. Patients rated overall experiential concordance, retrospectively perceived preoperative usefulness, and clarity (0–10). Medical records provided operative duration, first mobilization day, length of stay, and 90-day complications. Three prespecified chart-verifiable timing statements formed a 0–3 timing concordance score. Six multidisciplinary clinicians independently rated all answers during separate 30 min digital assessment sessions. Results: Mean overall patient-perceived concordance was 8.44 ± 0.95. Fourteen patients (13.9%) had a complicated recovery and reported lower concordance than those with uncomplicated recovery (7.57 ± 1.22 vs. 8.57 ± 0.83; p = 0.0026). The stated ranges for operative duration, early mobilization, and length of stay matched the corresponding records in 93.1%, 94.1%, and 91.1% of patients, respectively; 85.1% met all three timing criteria. Complete three-criterion timing concordance was lower after complicated recovery (28.6% vs. 94.3%; p < 0.001). Patient-perceived concordance correlated with the timing concordance score (Spearman ρ = 0.395; p < 0.001). Mean multidisciplinary appropriateness across 72 ratings was 8.50 ± 0.50. Conclusions: The fixed AI-generated THA information set was rated favorably by patients and clinicians, and its three prespecified timing statements corresponded closely to routine documented recovery. Both experiential and timing concordance were lower after complicated courses. These findings support clinician-reviewed AI information as an adjunct while underscoring that generalized timelines are conditional ranges, not individualized predictions.

Authors

Institutions

Publication Details

Journal
Medicina
Published
2026-09-27
DOI
https://doi.org/10.3390/medicina62101873
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Patient-Perceived and Record-Based Concordance of AI-Generated Perioperative Information After Total Hip Arthroplasty: A Cross-Sectional Study with Multidisciplinary Review

Oktay Adanır, Ozancan Biçer, Mehmet Yağız Yenigün, Cemre Aydın et al.
Medicina
Artificial Intelligence in Healthcare and Education
article

Patient-Perceived and Record-Based Concordance of AI-Generated Perioperative Information After Total Hip Arthroplasty: A Cross-Sectional Study with Multidisciplinary Review

Oktay Adanır, Ozancan Biçer, Mehmet Yağız Yenigün, Cemre Aydın, Ali Tarık Kutlu, Sidar Güneş
article en

Abstract

Background and Objectives: Large language models, a form of artificial intelligence (AI), can answer common patient questions about total hip arthroplasty (THA), but most evaluations emphasize clinician-rated accuracy or patient preference rather than whether generalized recovery statements correspond to patients’ documented postoperative courses. We evaluated a fixed set of ChatGPT 5.1 Thinking responses to 12 investigator-selected questions about elective primary THA from patient, medical-record, and multidisciplinary clinical perspectives, with particular attention to complicated recovery. Materials and Methods: In this cross-sectional study with retrospective record review, 101 adults at least 6 months after elective primary THA reviewed a fixed Turkish-language set of ChatGPT 5.1 Thinking answers to 12 investigator-selected questions. Patients rated overall experiential concordance, retrospectively perceived preoperative usefulness, and clarity (0–10). Medical records provided operative duration, first mobilization day, length of stay, and 90-day complications. Three prespecified chart-verifiable timing statements formed a 0–3 timing concordance score. Six multidisciplinary clinicians independently rated all answers during separate 30 min digital assessment sessions. Results: Mean overall patient-perceived concordance was 8.44 ± 0.95. Fourteen patients (13.9%) had a complicated recovery and reported lower concordance than those with uncomplicated recovery (7.57 ± 1.22 vs. 8.57 ± 0.83; p = 0.0026). The stated ranges for operative duration, early mobilization, and length of stay matched the corresponding records in 93.1%, 94.1%, and 91.1% of patients, respectively; 85.1% met all three timing criteria. Complete three-criterion timing concordance was lower after complicated recovery (28.6% vs. 94.3%; p < 0.001). Patient-perceived concordance correlated with the timing concordance score (Spearman ρ = 0.395; p < 0.001). Mean multidisciplinary appropriateness across 72 ratings was 8.50 ± 0.50. Conclusions: The fixed AI-generated THA information set was rated favorably by patients and clinicians, and its three prespecified timing statements corresponded closely to routine documented recovery. Both experiential and timing concordance were lower after complicated courses. These findings support clinician-reviewed AI information as an adjunct while underscoring that generalized timelines are conditional ranges, not individualized predictions.

MedicinaVol. 62(10)
Bağcılar Eğitim ve Araştırma Hastanesi (TR), Sağlık Bilimleri Üniversitesi (TR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.