Expert evaluation of large language models for patient counseling during recovery after gynecologic and obstetric surgery: a blinded comparative cross-sectional study

The postoperative period after gynecologic and obstetric surgery generates substantial patient information needs, and large language models (LLMs) are increasingly proposed as tools to help answer patient questions. Their performance in this setting, particularly with respect to safety, remains poorly defined. We evaluated the patient-counseling performance of three LLMs for recovery follow-up after gynecologic and obstetric surgery. In this cross-sectional, observational study, 30 frequently asked patient questions about postoperative recovery were compiled with the input of five experienced surgeons and posed to ChatGPT-5.5 Pro, Claude Opus 4.7, and Gemini 3.1 Pro using a standardized prompt. The responses, with model identity concealed and presented in randomized order, were independently scored by 30 specialist obstetricians and gynecologists on a 5-point Likert scale for accuracy, comprehensiveness, understandability, safety, and appropriateness. The presence of appropriate safety warnings, response length, and inter-rater agreement were also assessed. With the question as the unit of analysis, scores were compared using repeated-measures analysis of variance with Bonferroni-corrected pairwise comparisons, and a linear mixed-effects model was fitted as a sensitivity analysis; safety-warning rates were compared using the chi-square test. The highest overall mean score was obtained by ChatGPT-5.5 Pro (4.42 ± 0.38), followed by Claude Opus 4.7 (4.31 ± 0.41) and Gemini 3.1 Pro (4.12 ± 0.46) ( p < 0.001; partial eta-squared 0.27). ChatGPT-5.5 Pro and Claude Opus 4.7 performed comparably, with no significant difference between them (mean difference 0.11, 95% CI − 0.02 to 0.24; p = 0.084); both scored significantly higher than Gemini 3.1 Pro (mean difference 0.30, 95% CI 0.18 to 0.43, and 0.19, 95% CI 0.07 to 0.32, respectively). Appropriate safety warnings were most frequent with Claude Opus 4.7 (90.0%), followed by ChatGPT-5.5 Pro (86.7%) and Gemini 3.1 Pro (76.7%). The lowest scores were obtained in the categories of conditions requiring urgent presentation and of bleeding and discharge. Inter-rater agreement was good (intraclass correlation coefficient 0.81) and internal consistency was high (Cronbach’s alpha 0.89). Large language models produced largely accurate and understandable counseling responses after gynecologic and obstetric surgery, but safety and comprehensiveness varied among models and were weakest in high-risk scenarios. These tools may support physician-supervised discharge education but should not replace physician consultation, particularly when urgent presentation may be required.

Authors

Institutions

Publication Details

Journal
BMC Surgery
Published
2026-09-15
DOI
https://doi.org/10.1186/s12893-026-04197-0
Primary Topic
Enhanced Recovery After Surgery
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Expert evaluation of large language models for patient counseling during recovery after gynecologic and obstetric surgery: a blinded comparative cross-sectional study

Mücahit Furkan Balcı, Samican Ozmen, Arif Onur Atay
BMC Surgery
Enhanced Recovery After Surgery
article

Expert evaluation of large language models for patient counseling during recovery after gynecologic and obstetric surgery: a blinded comparative cross-sectional study

Mücahit Furkan Balcı, Samican Ozmen, Arif Onur Atay
article en

Abstract

The postoperative period after gynecologic and obstetric surgery generates substantial patient information needs, and large language models (LLMs) are increasingly proposed as tools to help answer patient questions. Their performance in this setting, particularly with respect to safety, remains poorly defined. We evaluated the patient-counseling performance of three LLMs for recovery follow-up after gynecologic and obstetric surgery. In this cross-sectional, observational study, 30 frequently asked patient questions about postoperative recovery were compiled with the input of five experienced surgeons and posed to ChatGPT-5.5 Pro, Claude Opus 4.7, and Gemini 3.1 Pro using a standardized prompt. The responses, with model identity concealed and presented in randomized order, were independently scored by 30 specialist obstetricians and gynecologists on a 5-point Likert scale for accuracy, comprehensiveness, understandability, safety, and appropriateness. The presence of appropriate safety warnings, response length, and inter-rater agreement were also assessed. With the question as the unit of analysis, scores were compared using repeated-measures analysis of variance with Bonferroni-corrected pairwise comparisons, and a linear mixed-effects model was fitted as a sensitivity analysis; safety-warning rates were compared using the chi-square test. The highest overall mean score was obtained by ChatGPT-5.5 Pro (4.42 ± 0.38), followed by Claude Opus 4.7 (4.31 ± 0.41) and Gemini 3.1 Pro (4.12 ± 0.46) ( p < 0.001; partial eta-squared 0.27). ChatGPT-5.5 Pro and Claude Opus 4.7 performed comparably, with no significant difference between them (mean difference 0.11, 95% CI − 0.02 to 0.24; p = 0.084); both scored significantly higher than Gemini 3.1 Pro (mean difference 0.30, 95% CI 0.18 to 0.43, and 0.19, 95% CI 0.07 to 0.32, respectively). Appropriate safety warnings were most frequent with Claude Opus 4.7 (90.0%), followed by ChatGPT-5.5 Pro (86.7%) and Gemini 3.1 Pro (76.7%). The lowest scores were obtained in the categories of conditions requiring urgent presentation and of bleeding and discharge. Inter-rater agreement was good (intraclass correlation coefficient 0.81) and internal consistency was high (Cronbach’s alpha 0.89). Large language models produced largely accurate and understandable counseling responses after gynecologic and obstetric surgery, but safety and comprehensiveness varied among models and were weakest in high-risk scenarios. These tools may support physician-supervised discharge education but should not replace physician consultation, particularly when urgent presentation may be required.

BMC Surgery
State Hospital (GB), Sivas State Hospital (TR)
Openalex Percentile: Top 8%
Enhanced Recovery After Surgery
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.