Testing six LLM-generated sleep-EEG predictions on open polysomnography: a case study with a pre-frozen protocol and negative results
Language models can write research hypotheses much faster than anyone can test them. Earlier studies have judged such hypotheses by expert ratings, benchmarks with known answers or laboratory experiments. We report a small, fully documented case from a different angle: what happens to a handful of model-generated predictions when they are turned into prespecified numerical tests on public data. We tested six operationalized sleep-EEG predictions selected from archived GLM-5.3 outputs, which supplied no bibliographic citations for those predictions. Before this study downloaded its data we froze a pre-analysis protocol in version control, with pass and refutation thresholds, controls and a label hierarchy; it was not deposited in a public registry, and the author had seen hypnograms and some EEG of the main dataset in an earlier, unrelated analysis (section 6). The main data were Sleep-EDF Expanded (Sleep Cassette, 78 healthy people, night 1); the Dreem DOD-H dataset (25 people, five independent scorers) was planned for scorer-agreement controls. Two hypotheses met their prespecified refutation condition. For the claim that spindle-band (12 to 16 Hz) bursts become more frequent in the last ten minutes of wake before the first N2 epoch, the median burst-rate ratio in 25 people was 0.71, below the refutation threshold of 1.2; the decision uses the point median, as prespecified, and the 95% interval (0.19 to 1.35) still includes 1.2, so this is a verdict of the protocol's rule rather than a firm disproof. For the claim that spindle amplitude grows with the time since the previous spindle, the median amplitude ratio in 77 people was 0.99 (95% CI 0.97 to 1.01) and the median Spearman rho was -0.04; because amplitude also decided which spindles our threshold detector counted, this applies to the detected spindles only. Three hypotheses were not decided because the sample minimum was not met after event-level qualification (17 and 12 people against 20; 1 record against 100); the model's fixed -75 µV amplitude rule, applied to the negative peak of the bipolar Fpz-Cz derivation, may have contributed to this attrition, which we did not test. One could not be tested, because no dataset we audited provides bilateral leg EMG in enough healthy sleepers. No hypothesis passed. Most of the effort went into turning each hypothesis into a decidable rule, and on standard open sleep data four of the six ended undecided rather than true or false. Several of the underlying mechanisms have published antecedents, which we compare with each prediction. This study contributes a documented case of operationalizing and evaluating six model-generated sleep predictions, including negative results, eligibility attrition and dataset limitations. The event detectors were checked only on synthetic signals, so the ability of the tests to detect a real effect was not established. The study does not estimate the accuracy of LLM-generated hypotheses or establish that the underlying physiological ideas are new. It is not a clinical study. Files. Kulma_2026_D1_Sleep_LLM_Hypotheses_v0.7.pdf is the report. The ZIP archive contains the report in Markdown, the pre-analysis protocol with its full history and deviations ledger, the analysis code with unit tests, result files, aggregation scripts, provenance of the generated hypotheses, every AI review round of the protocol, code and results, a manifest linking each claim to its supporting file, licences and data attribution, a short guide to reproduction (README.md) and SHA-256 checksums. Seventeen files are labelled public copies in which a few operational details were replaced; the manifest gives the hash of each original and of its copy. No recordings or hypnograms are redistributed; download scripts for both datasets are included. Some supporting notes are in Polish. Use of AI tools. The hypotheses were generated by GLM-5.3; the analysis code and drafts of the text were written with Claude Code (Anthropic) under the author's direction; OpenAI Codex and Google Antigravity were used for review and pre-release checks. These automated reviews do not replace review by a sleep electrophysiologist; no sleep electrophysiologist has read this report yet. AI tools are not authors; the author takes responsibility for all content (report, section 8). Preprint, not peer reviewed. Text and results: CC BY 4.0; raw model output: CC0 1.0; code: MIT. Datasets keep their own licences (see LICENSE.md).
Authors
- Mariusz Kulma (ORCID: https://orcid.org/0009-0000-5550-8723)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23227116
- Primary Topic
- Sleep and related disorders
- Type
- preprint