What Transfers and What Collapses: A Cross-Generator Study of AI-Generated Image Detection with Corrected Evaluation

A cross-generator study of AI-generated image detection on a 14-generator benchmark of modern systems (Stable Diffusion 1.3/1.4/2/XL, SD3, FLUX.1-dev/schnell, DALL-E 2/3, Midjourney v5, Imagen 3, GLIDE, Adobe Firefly), under a leak-free leave-generators-out protocol with corrected metrics and bootstrap confidence intervals. Findings. In-distribution accuracy does not predict cross-generator accuracy: a 2-D spectral detector scores 0.795 in-distribution but 0.523 (chance) on unseen generators, and a fine-tuned CNN drops to chance on its worst unseen generator; a hand-crafted physics detector goes confidently below chance (0.258, fingerprint inversion). Only frozen foundation features generalize (frozen-CLIP probe, 0.917), and a one-class real-manifold model structurally cannot sign-invert (worst-case floor 0.611). The collapse is governed by training-generator diversity (leave-one-out recovers the CNN to 0.876 and the CLIP probe to 0.960). A cross-architecture probe shows detectors do better on an older GAN than on unseen diffusion. Improvement. Feature-space extensions (DINOv2 ensembling, generator-direction removal, one-class fusion, reconstruction fusion) do not beat the frozen-CLIP probe, but simple 5-crop test-time aggregation does, by a paired-bootstrap-significant margin: 0.914 -> 0.949 (three-seen, 95% CI [0.028, 0.042]) and 0.961 -> 0.972 (leave-one-out), with the largest gains on the hardest generators. We also document an AUC-orientation evaluation bug that returns 1-AUC and silently inverts results. Code and benchmark: https://github.com/theFinex/cross-generator-ai-detection

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-06-27
DOI
https://doi.org/10.5281/zenodo.20952261
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

What Transfers and What Collapses: A Cross-Generator Study of AI-Generated Image Detection with Corrected Evaluation

Mohamed Alaya
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

What Transfers and What Collapses: A Cross-Generator Study of AI-Generated Image Detection with Corrected Evaluation

Mohamed Alaya
preprint en

Abstract

A cross-generator study of AI-generated image detection on a 14-generator benchmark of modern systems (Stable Diffusion 1.3/1.4/2/XL, SD3, FLUX.1-dev/schnell, DALL-E 2/3, Midjourney v5, Imagen 3, GLIDE, Adobe Firefly), under a leak-free leave-generators-out protocol with corrected metrics and bootstrap confidence intervals. Findings. In-distribution accuracy does not predict cross-generator accuracy: a 2-D spectral detector scores 0.795 in-distribution but 0.523 (chance) on unseen generators, and a fine-tuned CNN drops to chance on its worst unseen generator; a hand-crafted physics detector goes confidently below chance (0.258, fingerprint inversion). Only frozen foundation features generalize (frozen-CLIP probe, 0.917), and a one-class real-manifold model structurally cannot sign-invert (worst-case floor 0.611). The collapse is governed by training-generator diversity (leave-one-out recovers the CNN to 0.876 and the CLIP probe to 0.960). A cross-architecture probe shows detectors do better on an older GAN than on unseen diffusion. Improvement. Feature-space extensions (DINOv2 ensembling, generator-direction removal, one-class fusion, reconstruction fusion) do not beat the frozen-CLIP probe, but simple 5-crop test-time aggregation does, by a paired-bootstrap-significant margin: 0.914 -> 0.949 (three-seen, 95% CI [0.028, 0.042]) and 0.961 -> 0.972 (leave-one-out), with the largest gains on the hardest generators. We also document an AUC-orientation evaluation bug that returns 1-AUC and silently inverts results. Code and benchmark: https://github.com/theFinex/cross-generator-ai-detection

Zenodo (CERN European Organization for Nuclear Research)
Institut de Santé et de Sécurité au Travail (TN)
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What Transfers and What Collapses: A Cross-Generator Study of AI-Generated Image Detection with Corrected Evaluation — Mohamed Alaya · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS