Adversarial Attack via Latent Feature Attribution of Diffusion Models
Traditional adversarial attacks are constrained by \({l_p}\) norm limits, resulting in adversarial examples with high-frequency noise that compromise their imperceptibility. In contrast, attacks using diffusion models without \({l_p}\) norm constraints can generate visually more covert adversarial examples; however, existing methods generally suffer from overfitting of the surrogate model and rely on equal-step-size perturbation updates that treat all spatial features equally, leading to poor transferability. To address these issues, we propose an adversarial attack method based on latent feature attribution using diffusion models (LFA-Diff). First, we apply random masking and inject Gaussian noise into the latent space. By aggregating gradients after multiple transformations, we generate a normalized weight map and perform feature attribution based on it, focusing perturbations on important features that are common across models, thereby improving the transferability of adversarial examples; Second, we generate perturbations constrained by the reconstruction error of a latent diffusion model VAE and use a momentum-based iterative algorithm to adaptively update the perturbations, ensuring that the latent variables retain visual naturalness after being restored to the pixel space. Experimental results demonstrate that the proposed method achieves significantly higher cross-model average attack success rates while maintaining superior image imperceptibility. Source code: https://github.com/shichiale/LFA-Diff .
Authors
- Jiale Shi
- Yanan Wang
- Chengxian Ge
- Tianpeng Li
- Yafei Song
Institutions
- PLA Information Engineering University (CN)
- China Electronics Technology Group Corporation (CN)
- Air Force Engineering University (CN)
Publication Details
- Journal
- International Journal of Computational Intelligence Systems
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1007/s44196-026-01629-w
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00