Large Language Models vs Expert Response: A Comparative Analysis of Patient Education Quality After Skull Base Surgery

Objective: Patients undergoing endoscopic skull base surgery (ESBS) often seek postoperative guidance. Large language models (LLMs) are increasingly accessible resources, but their quality in this setting remains poorly studied. This study compared LLM-generated responses with expert responses for common postoperative ESBS questions. Methods: Eight common postoperative questions were answered by seven LLMs (ChatGPT-o1, Gemini, Claude, Meta, GROK, DeepSeek, Copilot) and expert surgeons. Eight blinded raters evaluated responses for understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT), while accuracy was graded on a 5-point Likert scale. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). ANOVA with Tukey HSD tested group differences. Results: Most LLM-generated responses showed accuracy comparable to expert responses (p>0.05), except ChatGPT-o1, which scored lower (p=0.006). Significant differences were found in understandability (F(7,56)=23.42, p<0.001, η²=0.75) and actionability (F(7,56)=21.06, p<0.001, η²=0.72). Gemini, GROK, Meta, and DeepSeek achieved the highest PEMAT scores. Expert responses scored lowest on PEMAT. Expert and Claude required college-level reading, whereas ChatGPT-o1 and Copilot produced grade-school–level content. Gemini, GROK, Meta, and DeepSeek produced high-school–level text. Conclusion: LLMs generated responses that were more understandable, actionable, and readable than expert responses while maintaining comparable accuracy. LLMs may improve postoperative counseling after skull base surgery, though expert oversight remains essential.

Authors

Institutions

Publication Details

Journal
Journal of Neurological Surgery Part B Skull Base
Published
2026-10-07
DOI
https://doi.org/10.1055/a-2976-6204
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Large Language Models vs Expert Response: A Comparative Analysis of Patient Education Quality After Skull Base Surgery

Lirit Levi, Axel Eluid Renteria, Mahmood Manavi, Michael T. Chang et al.
Journal of Neurological Surgery Part B Skull Base
Artificial Intelligence in Healthcare and Education
article

Large Language Models vs Expert Response: A Comparative Analysis of Patient Education Quality After Skull Base Surgery

Lirit Levi, Axel Eluid Renteria, Mahmood Manavi, Michael T. Chang, Zara M. Patel, Juan Carlos Fernández Miranda, David Liu, Maxime Fieux, Arifeen Rahman, Jacquelyn Callender, Jayakar Nayak, Matei Banu, Peter Hwang, Noel Ayoub
article en

Abstract

Objective: Patients undergoing endoscopic skull base surgery (ESBS) often seek postoperative guidance. Large language models (LLMs) are increasingly accessible resources, but their quality in this setting remains poorly studied. This study compared LLM-generated responses with expert responses for common postoperative ESBS questions. Methods: Eight common postoperative questions were answered by seven LLMs (ChatGPT-o1, Gemini, Claude, Meta, GROK, DeepSeek, Copilot) and expert surgeons. Eight blinded raters evaluated responses for understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT), while accuracy was graded on a 5-point Likert scale. Readability was assessed using Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES). ANOVA with Tukey HSD tested group differences. Results: Most LLM-generated responses showed accuracy comparable to expert responses (p>0.05), except ChatGPT-o1, which scored lower (p=0.006). Significant differences were found in understandability (F(7,56)=23.42, p<0.001, η²=0.75) and actionability (F(7,56)=21.06, p<0.001, η²=0.72). Gemini, GROK, Meta, and DeepSeek achieved the highest PEMAT scores. Expert responses scored lowest on PEMAT. Expert and Claude required college-level reading, whereas ChatGPT-o1 and Copilot produced grade-school–level content. Gemini, GROK, Meta, and DeepSeek produced high-school–level text. Conclusion: LLMs generated responses that were more understandable, actionable, and readable than expert responses while maintaining comparable accuracy. LLMs may improve postoperative counseling after skull base surgery, though expert oversight remains essential.

Journal of Neurological Surgery Part B Skull Base
Hospices Civils de Lyon (FR), Stanford Medicine (US), Université de Montréal (CA), Medical University of Vienna (AT), Stanford University (US)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.