Expert evaluation of large language models for patient counseling during recovery after gynecologic and obstetric surgery: a blinded comparative cross-sectional study
The postoperative period after gynecologic and obstetric surgery generates substantial patient information needs, and large language models (LLMs) are increasingly proposed as tools to help answer patient questions. Their performance in this setting, particularly with respect to safety, remains poorly defined. We evaluated the patient-counseling performance of three LLMs for recovery follow-up after gynecologic and obstetric surgery. In this cross-sectional, observational study, 30 frequently asked patient questions about postoperative recovery were compiled with the input of five experienced surgeons and posed to ChatGPT-5.5 Pro, Claude Opus 4.7, and Gemini 3.1 Pro using a standardized prompt. The responses, with model identity concealed and presented in randomized order, were independently scored by 30 specialist obstetricians and gynecologists on a 5-point Likert scale for accuracy, comprehensiveness, understandability, safety, and appropriateness. The presence of appropriate safety warnings, response length, and inter-rater agreement were also assessed. With the question as the unit of analysis, scores were compared using repeated-measures analysis of variance with Bonferroni-corrected pairwise comparisons, and a linear mixed-effects model was fitted as a sensitivity analysis; safety-warning rates were compared using the chi-square test. The highest overall mean score was obtained by ChatGPT-5.5 Pro (4.42 ± 0.38), followed by Claude Opus 4.7 (4.31 ± 0.41) and Gemini 3.1 Pro (4.12 ± 0.46) ( p < 0.001; partial eta-squared 0.27). ChatGPT-5.5 Pro and Claude Opus 4.7 performed comparably, with no significant difference between them (mean difference 0.11, 95% CI − 0.02 to 0.24; p = 0.084); both scored significantly higher than Gemini 3.1 Pro (mean difference 0.30, 95% CI 0.18 to 0.43, and 0.19, 95% CI 0.07 to 0.32, respectively). Appropriate safety warnings were most frequent with Claude Opus 4.7 (90.0%), followed by ChatGPT-5.5 Pro (86.7%) and Gemini 3.1 Pro (76.7%). The lowest scores were obtained in the categories of conditions requiring urgent presentation and of bleeding and discharge. Inter-rater agreement was good (intraclass correlation coefficient 0.81) and internal consistency was high (Cronbach’s alpha 0.89). Large language models produced largely accurate and understandable counseling responses after gynecologic and obstetric surgery, but safety and comprehensiveness varied among models and were weakest in high-risk scenarios. These tools may support physician-supervised discharge education but should not replace physician consultation, particularly when urgent presentation may be required.
Authors
- Mücahit Furkan Balcı (ORCID: https://orcid.org/0000-0002-2821-3273)
- Samican Ozmen
- Arif Onur Atay
Institutions
- State Hospital (GB)
- Sivas State Hospital (TR)
Publication Details
- Journal
- BMC Surgery
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1186/s12893-026-04197-0
- Primary Topic
- Enhanced Recovery After Surgery
- Type
- article
- Field-Weighted Citation Impact
- 0.00