Evasion attacks on generative safeguards: target-concept reappearance under structural control in grey-box settings
Abstract Concept erasure has emerged as a practical safeguard for suppressing unsafe or unwanted concepts in text-to-image diffusion models. However, existing methods are primarily designed for text-triggered generation and are usually evaluated under text-only prompting. In this paper, we identify a practical robustness gap by showing that concepts suppressed under text-only evaluation can reappear after a concept-erased backbone is composed with a compatible structural controller and supplied with target-consistent, information-bearing structural conditions. We propose Structure-Guided Concept Reappearance (SGCR), an inference-time attack that combines a benign prompt with a structural condition, such as a Canny edge map, depth map, or segmentation map, to induce target-related outputs without retraining, gradient-based optimization, or modification of the defended backbone. Extensive experiments on explicit content, copyrighted artistic styles, and generic objects demonstrate that SGCR remains effective across multiple representative erasure defenses and structural modalities. These findings show that suppression established under text-only evaluation may not persist in controller-augmented diffusion pipelines. The demonstrated threat applies to local modular systems and services exposing compatible structural-control interfaces, rather than closed text-only APIs.
Authors
- Qiqi Bao (ORCID: https://orcid.org/0000-0001-9599-1844)
- Yaguan Qian (ORCID: https://orcid.org/0000-0003-4056-9755)
- Zhaoquan Gu (ORCID: https://orcid.org/0000-0001-7546-852X)
- Farid Naït‐Abdesselam (ORCID: https://orcid.org/0000-0002-5042-5387)
- Yunxin Zhang (ORCID: https://orcid.org/0009-0003-7914-6159)
- Bin Wang (ORCID: https://orcid.org/0000-0003-1118-3618)
- Jiaoling Li
- Shouling Ji
Publication Details
- Journal
- Cybersecurity
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1186/s42400-026-00664-6
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00