Guideline-grounded retrieval versus unconfigured generation in obstructive sleep apnea: a pilot comparison of notebookLM and ChatGPT for AASM v3.0 PSG scoring-rule questions
Large language models (LLMs) are increasingly used in sleep medicine, but their reliability for polysomnography (PSG) scoring-rule tasks that require up-to-date guideline knowledge remains uncertain. This pilot study compared response quality when identical AASM v3.0 PSG scoring-rule questions were answered by NotebookLM (now Gemini Notebook)—a retrieval-augmented generation (RAG) assistant configured with the AASM Scoring Manual v3.0 as its sole source—and ChatGPT (GPT-5.3), an unconfigured generative model without uploaded source documents. The task assessed guideline-based scoring knowledge, not interpretation of raw PSG signals or automated sleep staging. Thirty questions were developed across five PSG content domains: Technical ( n = 7), Sleep Staging ( n = 6), Respiratory Scoring ( n = 11), Movement ( n = 4), and Arousal ( n = 2). The AASM Manual v3.0 was uploaded to NotebookLM; no files were uploaded to ChatGPT. Identical question wording was used as the sole prompt (Appendix 1), without custom system prompts (prompting protocol in Appendix 3). Three board-certified sleep medicine physicians independently scored de-identified responses presented as Response A and Response B in randomized order using a binary rubric (0/1) across four evaluation domains—Accuracy, Evidence Reasoning, Additional Information, and Information Integration (720 ratings total). Accuracy was anchored to a pre-specified gold-standard answer key (Appendix 1). Analyses used the Wilcoxon signed-rank test, exact McNemar test with Holm correction, Fleiss’ kappa, and Cochran’s Q (Python 3.11; SciPy 1.13.1, statsmodels 0.14.2, and NumPy 1.26.4). NotebookLM received more favorable overall ratings (76.1% vs. 66.9%). Median total score per question was 9.5 versus 8.0 (mean matched difference + 1.10; Wilcoxon Z = 2.62, p = 0.0089; post-hoc power ≈ 83%). This difference derived entirely from Additional Information (46.7% vs. 23.3%; McNemar χ² = 12.60, Holm-adjusted p = 0.0020; odds ratio 4.0). No significant differences emerged in Accuracy (88.9%–90.0%), Evidence Reasoning, or Information Integration after correction. Inter-rater agreement was fair overall (Fleiss’ kappa 0.347 ChatGPT, 0.297 NotebookLM) but unreliable for Information Integration (kappa − 0.111 and − 0.023). Systematic rater effects appeared for ChatGPT (Cochran’s Q = 13.89, p = 0.0010) but not NotebookLM (Q = 1.61, p = 0.447). Under unequal resource allocation, guideline-grounded RAG provided clinically valuable supplementary PSG-scoring information more consistently than an unconfigured generative model. This likely reflects source grounding rather than intrinsic model superiority and does not establish RAG as universally preferable across all clinical AI use cases. AI may serve as a complementary educational and reference aid for PSG scoring rules, not as a substitute for expert clinical judgment or autonomous guideline surveillance in patient care.
Authors
- Volkan Tekin (ORCID: https://orcid.org/0000-0001-6605-0367)
- Mehmet Koçer (ORCID: https://orcid.org/0000-0001-9911-9260)
- Rukiye Oktay Tekin (ORCID: https://orcid.org/0009-0005-4863-9334)
Institutions
- University of Health Science (KH)
- Gülhane Askerî Tıp Akademisi (TR)
- Hacettepe University Hospital (TR)
- Sağlık Bilimleri Üniversitesi (TR)
- University of Health Sciences Antigua (AG)
Publication Details
- Journal
- Sleep Science and Practice
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1186/s41606-026-00203-9
- Primary Topic
- Obstructive Sleep Apnea Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00