Evaluating AI-Generated Tonsillectomy Guidance Against Clinical Practice Standards

Background/Objectives: Artificial intelligence (AI) tools are used in clinical education and patient communication. Concordance with American Academy of Otolaryngology–Head and Neck Surgery Foundation clinical practice guidelines (CPG) and overall readability remain poorly defined. Methods: Over 2 months, ChatGPT, Google Gemini, and Google Search AI were queried using a single LF prompt and a layered set of sub-questions. Responses were scored with a 12-point rubric from the CPG plain-language tonsillectomy guideline. Readability was assessed using Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Temporal trends and between-group differences were analyzed using regression and repeated-measures methods. Results: Layered prompts demonstrated greater alignment with guideline subtopics than LF (89.5% vs. 71.4%, p < 0.001). For LF queries, mean concordance was 74.5% (95% CI 72.0–77.0) for Google AI, 71.3% (68.9–73.7) for Gemini, and 68.4% (66.0–70.9) for ChatGPT. Overall, Google AI outperformed Gemini and ChatGPT on pairwise comparisons (p = 0.001). Google AI demonstrated improvement in concordance (slope = 0.03, p = 0.003), while Gemini showed no change (p > 0.05) and ChatGPT worsened over time (slope = −0.03, p < 0.001). Readability differed by strategy and model. Layered prompts produced more accessible text (FKGL 44.6, FRE 10.3) compared with LF outputs (FKGL 35.2, FRE 11.3) and the CPG (FKGL 10.6, FRE 41.7). IRR was high (ICC = 0.96). Conclusions: Model type and prompt structure shaped the quality of AI-generated tonsillectomy guidance. Google AI consistently outperformed ChatGPT, with Gemini intermediate, and uniquely showed temporal improvement. Layered prompting produced more CPG-aligned and readable responses across all models. Readability was consistently above a 7th grade level.

Authors

Institutions

Publication Details

Journal
Journal of Otorhinolaryngology Hearing and Balance Medicine
Published
2026-09-30
DOI
https://doi.org/10.3390/ohbm7020035
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating AI-Generated Tonsillectomy Guidance Against Clinical Practice Standards

Sudeepti Vedula, Brian Manzi, Hetal Lad, Shreeya Bahethi et al.
Journal of Otorhinolaryngology Hearing and Balance Medicine
Artificial Intelligence in Healthcare and Education
article

Evaluating AI-Generated Tonsillectomy Guidance Against Clinical Practice Standards

Sudeepti Vedula, Brian Manzi, Hetal Lad, Shreeya Bahethi, Shrey Shah
article en

Abstract

Background/Objectives: Artificial intelligence (AI) tools are used in clinical education and patient communication. Concordance with American Academy of Otolaryngology–Head and Neck Surgery Foundation clinical practice guidelines (CPG) and overall readability remain poorly defined. Methods: Over 2 months, ChatGPT, Google Gemini, and Google Search AI were queried using a single LF prompt and a layered set of sub-questions. Responses were scored with a 12-point rubric from the CPG plain-language tonsillectomy guideline. Readability was assessed using Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Temporal trends and between-group differences were analyzed using regression and repeated-measures methods. Results: Layered prompts demonstrated greater alignment with guideline subtopics than LF (89.5% vs. 71.4%, p < 0.001). For LF queries, mean concordance was 74.5% (95% CI 72.0–77.0) for Google AI, 71.3% (68.9–73.7) for Gemini, and 68.4% (66.0–70.9) for ChatGPT. Overall, Google AI outperformed Gemini and ChatGPT on pairwise comparisons (p = 0.001). Google AI demonstrated improvement in concordance (slope = 0.03, p = 0.003), while Gemini showed no change (p > 0.05) and ChatGPT worsened over time (slope = −0.03, p < 0.001). Readability differed by strategy and model. Layered prompts produced more accessible text (FKGL 44.6, FRE 10.3) compared with LF outputs (FKGL 35.2, FRE 11.3) and the CPG (FKGL 10.6, FRE 41.7). IRR was high (ICC = 0.96). Conclusions: Model type and prompt structure shaped the quality of AI-generated tonsillectomy guidance. Google AI consistently outperformed ChatGPT, with Gemini intermediate, and uniquely showed temporal improvement. Layered prompting produced more CPG-aligned and readable responses across all models. Readability was consistently above a 7th grade level.

Journal of Otorhinolaryngology Hearing and Balance MedicineVol. 7(2)
Hackensack Meridian Health (US), Center for Discovery (US), Rutgers New Jersey Medical School (US)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.