Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results

Language models can write research hypotheses much faster than anyone can test them. Earlier studies have judged such hypotheses by expert ratings, benchmarks with known answers or laboratory experiments. We report a small, fully documented case from a different angle: what happens to a handful of model-generated predictions when they are turned into prespecified numerical tests on public data. We tested six operationalized sleep-EEG predictions selected from archived GLM-5.3 outputs, which supplied no bibliographic citations for those predictions. Before this study downloaded its data we froze a pre-analysis protocol in version control, with pass and refutation thresholds, controls and a label hierarchy; it was not deposited in a public registry, and the author had seen hypnograms and some EEG of the main dataset in an earlier, unrelated analysis (section 6). The main data were Sleep-EDF Expanded (Sleep Cassette, 78 healthy people, night 1); the Dreem DOD-H dataset (25 people, five independent scorers) was planned for scorer-agreement controls. Two hypotheses met their prespecified refutation condition. For the claim that spindle-band (12 to 16 Hz) bursts become more frequent in the last ten minutes of wake before the first N2 epoch, the median burst-rate ratio in 25 people was 0.71, below the refutation threshold of 1.2; the decision uses the point median, as prespecified, and the 95% interval (0.19 to 1.35) still includes 1.2, so this is a verdict of the protocol's rule rather than a firm disproof. For the claim that spindle amplitude grows with the time since the previous spindle, the median amplitude ratio in 77 people was 0.99 (95% CI 0.97 to 1.01) and the median Spearman rho was -0.04; because amplitude also decided which spindles our threshold detector counted, this applies to the detected spindles only. Three hypotheses were not decided because the sample minimum was not met after event-level qualification (17 and 12 people against 20; 1 record against 100); the model's fixed -75 µV amplitude rule, applied to the negative peak of the bipolar Fpz-Cz derivation, may have contributed to this attrition, which we did not test. One could not be tested, because no dataset we audited provides bilateral leg EMG in enough healthy sleepers. No hypothesis passed. Most of the effort went into turning each hypothesis into a decidable rule, and on standard open sleep data four of the six ended undecided rather than true or false. Several of the underlying mechanisms have published antecedents, which we compare with each prediction. This study contributes a documented case of operationalizing and evaluating six model-generated sleep predictions, including negative results, eligibility attrition and dataset limitations. The event detectors were checked only on synthetic signals, so the ability of the tests to detect a real effect was not established. The study does not estimate the accuracy of LLM-generated hypotheses or establish that the underlying physiological ideas are new. It is not a clinical study. Files. Kulma_2026_D1_Sleep_LLM_Hypotheses_v0.7.pdf is the report. The ZIP archive contains the report in Markdown, the pre-analysis protocol with its full history and deviations ledger, the analysis code with unit tests, result files, aggregation scripts, provenance of the generated hypotheses, every AI review round of the protocol, code and results, a manifest linking each claim to its supporting file, licences and data attribution, a short guide to reproduction (README.md) and SHA-256 checksums. Seventeen files are labelled public copies in which a few operational details were replaced; the manifest gives the hash of each original and of its copy. No recordings or hypnograms are redistributed; download scripts for both datasets are included. Some supporting notes are in Polish. Use of AI tools. The hypotheses were generated by GLM-5.3; the analysis code and drafts of the text were written with Claude Code (Anthropic) under the author's direction; OpenAI Codex and Google Antigravity were used for review and pre-release checks. These automated reviews do not replace review by a sleep electrophysiologist; no sleep electrophysiologist has read this report yet. AI tools are not authors; the author takes responsibility for all content (report, section 8). Preprint, not peer reviewed. Text and results: CC BY 4.0; raw model output: CC0 1.0; code: MIT. Datasets keep their own licences (see LICENSE.md).

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23227116
Primary Topic
Sleep and related disorders
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results

Mariusz Kulma
Zenodo (CERN European Organization for Nuclear Research)
Sleep and related disorders
preprint

Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results

Mariusz Kulma
preprint en

Abstract

Language models can write research hypotheses much faster than anyone can test them. Earlier studies have judged such hypotheses by expert ratings, benchmarks with known answers or laboratory experiments. We report a small, fully documented case from a different angle: what happens to a handful of model-generated predictions when they are turned into prespecified numerical tests on public data. We tested six operationalized sleep-EEG predictions selected from archived GLM-5.3 outputs, which supplied no bibliographic citations for those predictions. Before this study downloaded its data we froze a pre-analysis protocol in version control, with pass and refutation thresholds, controls and a label hierarchy; it was not deposited in a public registry, and the author had seen hypnograms and some EEG of the main dataset in an earlier, unrelated analysis (section 6). The main data were Sleep-EDF Expanded (Sleep Cassette, 78 healthy people, night 1); the Dreem DOD-H dataset (25 people, five independent scorers) was planned for scorer-agreement controls. Two hypotheses met their prespecified refutation condition. For the claim that spindle-band (12 to 16 Hz) bursts become more frequent in the last ten minutes of wake before the first N2 epoch, the median burst-rate ratio in 25 people was 0.71, below the refutation threshold of 1.2; the decision uses the point median, as prespecified, and the 95% interval (0.19 to 1.35) still includes 1.2, so this is a verdict of the protocol's rule rather than a firm disproof. For the claim that spindle amplitude grows with the time since the previous spindle, the median amplitude ratio in 77 people was 0.99 (95% CI 0.97 to 1.01) and the median Spearman rho was -0.04; because amplitude also decided which spindles our threshold detector counted, this applies to the detected spindles only. Three hypotheses were not decided because the sample minimum was not met after event-level qualification (17 and 12 people against 20; 1 record against 100); the model's fixed -75 µV amplitude rule, applied to the negative peak of the bipolar Fpz-Cz derivation, may have contributed to this attrition, which we did not test. One could not be tested, because no dataset we audited provides bilateral leg EMG in enough healthy sleepers. No hypothesis passed. Most of the effort went into turning each hypothesis into a decidable rule, and on standard open sleep data four of the six ended undecided rather than true or false. Several of the underlying mechanisms have published antecedents, which we compare with each prediction. This study contributes a documented case of operationalizing and evaluating six model-generated sleep predictions, including negative results, eligibility attrition and dataset limitations. The event detectors were checked only on synthetic signals, so the ability of the tests to detect a real effect was not established. The study does not estimate the accuracy of LLM-generated hypotheses or establish that the underlying physiological ideas are new. It is not a clinical study. Files. Kulma_2026_D1_Sleep_LLM_Hypotheses_v0.7.pdf is the report. The ZIP archive contains the report in Markdown, the pre-analysis protocol with its full history and deviations ledger, the analysis code with unit tests, result files, aggregation scripts, provenance of the generated hypotheses, every AI review round of the protocol, code and results, a manifest linking each claim to its supporting file, licences and data attribution, a short guide to reproduction (README.md) and SHA-256 checksums. Seventeen files are labelled public copies in which a few operational details were replaced; the manifest gives the hash of each original and of its copy. No recordings or hypnograms are redistributed; download scripts for both datasets are included. Some supporting notes are in Polish. Use of AI tools. The hypotheses were generated by GLM-5.3; the analysis code and drafts of the text were written with Claude Code (Anthropic) under the author's direction; OpenAI Codex and Google Antigravity were used for review and pre-release checks. These automated reviews do not replace review by a sleep electrophysiologist; no sleep electrophysiologist has read this report yet. AI tools are not authors; the author takes responsibility for all content (report, section 8). Preprint, not peer reviewed. Text and results: CC BY 4.0; raw model output: CC0 1.0; code: MIT. Datasets keep their own licences (see LICENSE.md).

Zenodo (CERN European Organization for Nuclear Research)
Sleep and related disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.