Large Language Models vs Expert Response: A Comparative Analysis of Patient Education Quality After Skull Base Surgery
Objective: Patients undergoing endoscopic skull base surgery (ESBS) often seek postoperative guidance. Large language models (LLMs) are increasingly accessible resources, but their quality in this setting remains poorly studied. This study compared LLM-generated responses with expert responses for common postoperative ESBS questions. Methods: Eight common postoperative questions were answered by seven LLMs (ChatGPT-o1, Gemini, Claude, Meta, GROK, DeepSeek, Copilot) and expert surgeons. Eight blinded raters evaluated responses for understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT), while accuracy was graded on a 5-point Likert scale. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). ANOVA with Tukey HSD tested group differences. Results: Most LLM-generated responses showed accuracy comparable to expert responses (p>0.05), except ChatGPT-o1, which scored lower (p=0.006). Significant differences were found in understandability (F(7,56)=23.42, p<0.001, η²=0.75) and actionability (F(7,56)=21.06, p<0.001, η²=0.72). Gemini, GROK, Meta, and DeepSeek achieved the highest PEMAT scores. Expert responses scored lowest on PEMAT. Expert and Claude required college-level reading, whereas ChatGPT-o1 and Copilot produced grade-school–level content. Gemini, GROK, Meta, and DeepSeek produced high-school–level text. Conclusion: LLMs generated responses that were more understandable, actionable, and readable than expert responses while maintaining comparable accuracy. LLMs may improve postoperative counseling after skull base surgery, though expert oversight remains essential.
Authors
- Lirit Levi (ORCID: https://orcid.org/0000-0002-7075-8752)
- Axel Eluid Renteria (ORCID: https://orcid.org/0000-0002-3180-3047)
- Mahmood Manavi (ORCID: https://orcid.org/0000-0003-4114-1371)
- Michael T. Chang (ORCID: https://orcid.org/0000-0001-9519-5880)
- Zara M. Patel (ORCID: https://orcid.org/0000-0003-2072-982X)
- Juan Carlos Fernández Miranda
- David Liu
- Maxime Fieux
- Arifeen Rahman
- Jacquelyn Callender
- Jayakar Nayak
- Matei Banu
- Peter Hwang
- Noel Ayoub
Institutions
- Hospices Civils de Lyon (FR)
- Stanford Medicine (US)
- Université de Montréal (CA)
- Medical University of Vienna (AT)
- Stanford University (US)
Publication Details
- Journal
- Journal of Neurological Surgery Part B Skull Base
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1055/a-2976-6204
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00