E100 pre-registration: does clearing non-body content from the post-H1 snippet window change whether ChatGPT cites a page, when the page body is held constant?
E100 is a pre-registered matched-pair experiment testing whether clearing non-body content from the window immediately after a page's H1 changes whether ChatGPT cites the page, with the page body held constant. Two pages on thegeolab.net share an identical title, H1 and body below the fold and differ only in what sits between the H1 and the opening sentence. Arm A has nothing there. Arm B retains five elements in that window: a category label, a publish date, a table of contents, an image with alt text and a widget. Capture is the ChatGPT API with web_search_preview enabled (model gpt-5.6-luna, max_output_tokens 3000), not the consumer instant-mode UI, and all claims are scoped to that path. Eight frozen queries are re-asked until 80 browsed observations per arm are reached, capped at 160 calls. Calls that return no search results are dropped from the denominator rather than scored as zero, and top-up renders cycle the frozen query order. The primary outcome is the arm's URL appearing in sources[]. The registered prediction is a directional lift for arm A, read as null if the A minus B difference is below 22 percentage points or its 95% confidence interval includes zero. The origin hypothesis is the RESONEO and Abondance report (July 2026) that ChatGPT selects on the title plus a stored snippet of roughly 200 characters anchored on the H1; that mechanism is what is under test, not an established description of the system. This record is the frozen protocol. T0 is the date of this deposit and no capture runs before T0. Any change made during capture or analysis will be reported as a deviation. The sibling experiment E097 (URL slug) was withdrawn from a joint deposit before registration after Google folded its two byte-identical arms to a single canonical, and is registered as a separate record. Any comparison between E097 and E100 is descriptive only. Site state at T0. Test pages: https://thegeolab.net/guides/paywalls-and-ai-search-a/ (arm A, clean window) and https://thegeolab.net/guides/paywalls-and-ai-search-b/ (arm B, clutter retained). Pages published 11 September 2026 08:00 UTC and frozen at commit baa06e9, tag e097-e100-pages-frozen, in the repository arturseo-geo/geo-lab-experiments. Both arms verified indexed by Google on two independent reads (14 and 16 September 2026) with each arm's self-canonical honoured, and re-confirmed unchanged on 19 September 2026 immediately before deposit. Phase 0 served-snippet read passed: the post-H1 window md5 is identical whether fetched as a browser or as Googlebot (arm A 80c2783a376bed47037ebf778c14cc7e, arm B 3eaf4fa976276888144a01cadd099f9b); arm A reaches its opening sentence at tag-stripped offset 0, arm B at offset 728. Protocol source frozen at tag e100-prereg-v1.3, commit 9bd3311, in the same repository.
Authors
- Artur Ferreira
- Marwa Saleh (ORCID: https://orcid.org/0009-0008-8525-2747)
Institutions
- Learning Through an Expanded Arts Program (US)
- Grantmakers for Effective Organizations (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22849119
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint