Guideline-grounded retrieval versus unconfigured generation in obstructive sleep apnea: a pilot comparison of notebookLM and ChatGPT for AASM v3.0 PSG scoring-rule questions

Large language models (LLMs) are increasingly used in sleep medicine, but their reliability for polysomnography (PSG) scoring-rule tasks that require up-to-date guideline knowledge remains uncertain. This pilot study compared response quality when identical AASM v3.0 PSG scoring-rule questions were answered by NotebookLM (now Gemini Notebook)—a retrieval-augmented generation (RAG) assistant configured with the AASM Scoring Manual v3.0 as its sole source—and ChatGPT (GPT-5.3), an unconfigured generative model without uploaded source documents. The task assessed guideline-based scoring knowledge, not interpretation of raw PSG signals or automated sleep staging. Thirty questions were developed across five PSG content domains: Technical ( n = 7), Sleep Staging ( n = 6), Respiratory Scoring ( n = 11), Movement ( n = 4), and Arousal ( n = 2). The AASM Manual v3.0 was uploaded to NotebookLM; no files were uploaded to ChatGPT. Identical question wording was used as the sole prompt (Appendix 1), without custom system prompts (prompting protocol in Appendix 3). Three board-certified sleep medicine physicians independently scored de-identified responses presented as Response A and Response B in randomized order using a binary rubric (0/1) across four evaluation domains—Accuracy, Evidence Reasoning, Additional Information, and Information Integration (720 ratings total). Accuracy was anchored to a pre-specified gold-standard answer key (Appendix 1). Analyses used the Wilcoxon signed-rank test, exact McNemar test with Holm correction, Fleiss’ kappa, and Cochran’s Q (Python 3.11; SciPy 1.13.1, statsmodels 0.14.2, and NumPy 1.26.4). NotebookLM received more favorable overall ratings (76.1% vs. 66.9%). Median total score per question was 9.5 versus 8.0 (mean matched difference + 1.10; Wilcoxon Z = 2.62, p = 0.0089; post-hoc power ≈ 83%). This difference derived entirely from Additional Information (46.7% vs. 23.3%; McNemar χ² = 12.60, Holm-adjusted p = 0.0020; odds ratio 4.0). No significant differences emerged in Accuracy (88.9%–90.0%), Evidence Reasoning, or Information Integration after correction. Inter-rater agreement was fair overall (Fleiss’ kappa 0.347 ChatGPT, 0.297 NotebookLM) but unreliable for Information Integration (kappa − 0.111 and − 0.023). Systematic rater effects appeared for ChatGPT (Cochran’s Q = 13.89, p = 0.0010) but not NotebookLM (Q = 1.61, p = 0.447). Under unequal resource allocation, guideline-grounded RAG provided clinically valuable supplementary PSG-scoring information more consistently than an unconfigured generative model. This likely reflects source grounding rather than intrinsic model superiority and does not establish RAG as universally preferable across all clinical AI use cases. AI may serve as a complementary educational and reference aid for PSG scoring rules, not as a substitute for expert clinical judgment or autonomous guideline surveillance in patient care.

Authors

Institutions

Publication Details

Journal
Sleep Science and Practice
Published
2026-10-05
DOI
https://doi.org/10.1186/s41606-026-00203-9
Primary Topic
Obstructive Sleep Apnea Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Guideline-grounded retrieval versus unconfigured generation in obstructive sleep apnea: a pilot comparison of notebookLM and ChatGPT for AASM v3.0 PSG scoring-rule questions

Volkan Tekin, Mehmet Koçer, Rukiye Oktay Tekin
Sleep Science and Practice
Obstructive Sleep Apnea Research
article

Guideline-grounded retrieval versus unconfigured generation in obstructive sleep apnea: a pilot comparison of notebookLM and ChatGPT for AASM v3.0 PSG scoring-rule questions

Volkan Tekin, Mehmet Koçer, Rukiye Oktay Tekin
article en

Abstract

Large language models (LLMs) are increasingly used in sleep medicine, but their reliability for polysomnography (PSG) scoring-rule tasks that require up-to-date guideline knowledge remains uncertain. This pilot study compared response quality when identical AASM v3.0 PSG scoring-rule questions were answered by NotebookLM (now Gemini Notebook)—a retrieval-augmented generation (RAG) assistant configured with the AASM Scoring Manual v3.0 as its sole source—and ChatGPT (GPT-5.3), an unconfigured generative model without uploaded source documents. The task assessed guideline-based scoring knowledge, not interpretation of raw PSG signals or automated sleep staging. Thirty questions were developed across five PSG content domains: Technical ( n = 7), Sleep Staging ( n = 6), Respiratory Scoring ( n = 11), Movement ( n = 4), and Arousal ( n = 2). The AASM Manual v3.0 was uploaded to NotebookLM; no files were uploaded to ChatGPT. Identical question wording was used as the sole prompt (Appendix 1), without custom system prompts (prompting protocol in Appendix 3). Three board-certified sleep medicine physicians independently scored de-identified responses presented as Response A and Response B in randomized order using a binary rubric (0/1) across four evaluation domains—Accuracy, Evidence Reasoning, Additional Information, and Information Integration (720 ratings total). Accuracy was anchored to a pre-specified gold-standard answer key (Appendix 1). Analyses used the Wilcoxon signed-rank test, exact McNemar test with Holm correction, Fleiss’ kappa, and Cochran’s Q (Python 3.11; SciPy 1.13.1, statsmodels 0.14.2, and NumPy 1.26.4). NotebookLM received more favorable overall ratings (76.1% vs. 66.9%). Median total score per question was 9.5 versus 8.0 (mean matched difference + 1.10; Wilcoxon Z = 2.62, p = 0.0089; post-hoc power ≈ 83%). This difference derived entirely from Additional Information (46.7% vs. 23.3%; McNemar χ² = 12.60, Holm-adjusted p = 0.0020; odds ratio 4.0). No significant differences emerged in Accuracy (88.9%–90.0%), Evidence Reasoning, or Information Integration after correction. Inter-rater agreement was fair overall (Fleiss’ kappa 0.347 ChatGPT, 0.297 NotebookLM) but unreliable for Information Integration (kappa − 0.111 and − 0.023). Systematic rater effects appeared for ChatGPT (Cochran’s Q = 13.89, p = 0.0010) but not NotebookLM (Q = 1.61, p = 0.447). Under unequal resource allocation, guideline-grounded RAG provided clinically valuable supplementary PSG-scoring information more consistently than an unconfigured generative model. This likely reflects source grounding rather than intrinsic model superiority and does not establish RAG as universally preferable across all clinical AI use cases. AI may serve as a complementary educational and reference aid for PSG scoring rules, not as a substitute for expert clinical judgment or autonomous guideline surveillance in patient care.

Sleep Science and PracticeVol. 10(1)
University of Health Science (KH), Gülhane Askerî Tıp Akademisi (TR), Hacettepe University Hospital (TR), Sağlık Bilimleri Üniversitesi (TR), University of Health Sciences Antigua (AG)
Good health and well-being
Openalex Percentile: Top 13%
Obstructive Sleep Apnea Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.