Understanding through Perturbation: lowering classifier-free guidance as a dose-response assay for two text-to-image latent models

I probe a generative model by lowering one control, pre-registered, and scoring what breaks with a rubric borrowed from Klüver. I generated 860 images on two latent text-to-image checkpoints, SDXL and SD 3.5: 6 prompts, 7 guidance values, 10 seeds, scored blind by two vision-language judges on criteria I committed to git before the run; two humans rated a subset. Lower classifier-free guidance was associated with greater rubric-scored visual breakdown, and the association remained after adjustment for two no-reference quality proxies, robustly to prompt resampling on SDXL but not on SD 3.5. SDXL met every pre-registered gate under the post-data amended judge panel. SD 3.5 missed the slope gate: its slope of -0.182 sits above the -0.20 line, and its interval straddles that line, so an effect of the registered size is neither shown nor ruled out. Looking afterwards, almost all of SD 3.5's drop happens between g = 1 and g = 2. A second pre-registered rubric, adapted from Suzuki's axes, replicated on SDXL and failed on SD 3.5, where veridicality moved opposite to the registered direction. Qwen2.5-VL-7B, the original second judge, failed silently, returning well-formed zeros on visibly broken images, and was caught only by a human subset after the run; Qwen3-VL-8B failed the same way in screening. Replacing the 7B with Qwen3-VL-32B, a post-data amendment, raised composite kappa from 0.290 to 0.562 (SDXL) and 0.158 to 0.440 (SD 3.5; 0.394 without the empty-prompt baselines); the two humans agreed at 0.567, and with the best judge at 0.337 and 0.426. An earlier experiment found no Klüver form constants, but its judge also flagged them on 45% of ordinary images, above a registered ceiling of 20%, so that null is uninterpretable under its own rule. The psychedelic analogy supplies the rubric and is not tested; code, pre-registrations, images and every judge reply are public. Code, pre-registrations, analyses and every judge reply: https://github.com/youssefhassan/operating-system-hypothesis-public (tag preprint-I-v1). Images and judgements: Hugging Face datasets youssefhassan13/exp01-guidance-sweep, youssefhassan13/exp02-form-constant-generator, youssefhassan13/exp03-l23-hardening.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23165988
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Understanding through Perturbation: lowering classifier-free guidance as a dose-response assay for two text-to-image latent models

Youssef Hassan
Zenodo (CERN European Organization for Nuclear Research)
Generative Adversarial Networks and Image Synthesis
preprint

Understanding through Perturbation: lowering classifier-free guidance as a dose-response assay for two text-to-image latent models

Youssef Hassan
preprint en

Abstract

I probe a generative model by lowering one control, pre-registered, and scoring what breaks with a rubric borrowed from Klüver. I generated 860 images on two latent text-to-image checkpoints, SDXL and SD 3.5: 6 prompts, 7 guidance values, 10 seeds, scored blind by two vision-language judges on criteria I committed to git before the run; two humans rated a subset. Lower classifier-free guidance was associated with greater rubric-scored visual breakdown, and the association remained after adjustment for two no-reference quality proxies, robustly to prompt resampling on SDXL but not on SD 3.5. SDXL met every pre-registered gate under the post-data amended judge panel. SD 3.5 missed the slope gate: its slope of -0.182 sits above the -0.20 line, and its interval straddles that line, so an effect of the registered size is neither shown nor ruled out. Looking afterwards, almost all of SD 3.5's drop happens between g = 1 and g = 2. A second pre-registered rubric, adapted from Suzuki's axes, replicated on SDXL and failed on SD 3.5, where veridicality moved opposite to the registered direction. Qwen2.5-VL-7B, the original second judge, failed silently, returning well-formed zeros on visibly broken images, and was caught only by a human subset after the run; Qwen3-VL-8B failed the same way in screening. Replacing the 7B with Qwen3-VL-32B, a post-data amendment, raised composite kappa from 0.290 to 0.562 (SDXL) and 0.158 to 0.440 (SD 3.5; 0.394 without the empty-prompt baselines); the two humans agreed at 0.567, and with the best judge at 0.337 and 0.426. An earlier experiment found no Klüver form constants, but its judge also flagged them on 45% of ordinary images, above a registered ceiling of 20%, so that null is uninterpretable under its own rule. The psychedelic analogy supplies the rubric and is not tested; code, pre-registrations, images and every judge reply are public. Code, pre-registrations, analyses and every judge reply: https://github.com/youssefhassan/operating-system-hypothesis-public (tag preprint-I-v1). Images and judgements: Hugging Face datasets youssefhassan13/exp01-guidance-sweep, youssefhassan13/exp02-form-constant-generator, youssefhassan13/exp03-l23-hardening.

Zenodo (CERN European Organization for Nuclear Research)
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Understanding through Perturbation: lowering classifier-free guidance as a dose-response assay for two text-to-image latent models — Youssef Hassan · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS