Adversarial Attack via Latent Feature Attribution of Diffusion Models

Traditional adversarial attacks are constrained by \({l_p}\) norm limits, resulting in adversarial examples with high-frequency noise that compromise their imperceptibility. In contrast, attacks using diffusion models without \({l_p}\) norm constraints can generate visually more covert adversarial examples; however, existing methods generally suffer from overfitting of the surrogate model and rely on equal-step-size perturbation updates that treat all spatial features equally, leading to poor transferability. To address these issues, we propose an adversarial attack method based on latent feature attribution using diffusion models (LFA-Diff). First, we apply random masking and inject Gaussian noise into the latent space. By aggregating gradients after multiple transformations, we generate a normalized weight map and perform feature attribution based on it, focusing perturbations on important features that are common across models, thereby improving the transferability of adversarial examples; Second, we generate perturbations constrained by the reconstruction error of a latent diffusion model VAE and use a momentum-based iterative algorithm to adaptively update the perturbations, ensuring that the latent variables retain visual naturalness after being restored to the pixel space. Experimental results demonstrate that the proposed method achieves significantly higher cross-model average attack success rates while maintaining superior image imperceptibility. Source code: https://github.com/shichiale/LFA-Diff .

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence Systems
Published
2026-10-05
DOI
https://doi.org/10.1007/s44196-026-01629-w
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Adversarial Attack via Latent Feature Attribution of Diffusion Models

Jiale Shi, Yanan Wang, Chengxian Ge, Tianpeng Li et al.
International Journal of Computational Intelligence Systems
Adversarial Robustness in Machine Learning
article

Adversarial Attack via Latent Feature Attribution of Diffusion Models

Jiale Shi, Yanan Wang, Chengxian Ge, Tianpeng Li, Yafei Song
article en

Abstract

Traditional adversarial attacks are constrained by \({l_p}\) norm limits, resulting in adversarial examples with high-frequency noise that compromise their imperceptibility. In contrast, attacks using diffusion models without \({l_p}\) norm constraints can generate visually more covert adversarial examples; however, existing methods generally suffer from overfitting of the surrogate model and rely on equal-step-size perturbation updates that treat all spatial features equally, leading to poor transferability. To address these issues, we propose an adversarial attack method based on latent feature attribution using diffusion models (LFA-Diff). First, we apply random masking and inject Gaussian noise into the latent space. By aggregating gradients after multiple transformations, we generate a normalized weight map and perform feature attribution based on it, focusing perturbations on important features that are common across models, thereby improving the transferability of adversarial examples; Second, we generate perturbations constrained by the reconstruction error of a latent diffusion model VAE and use a momentum-based iterative algorithm to adaptively update the perturbations, ensuring that the latent variables retain visual naturalness after being restored to the pixel space. Experimental results demonstrate that the proposed method achieves significantly higher cross-model average attack success rates while maintaining superior image imperceptibility. Source code: https://github.com/shichiale/LFA-Diff .

International Journal of Computational Intelligence Systems
PLA Information Engineering University (CN), China Electronics Technology Group Corporation (CN), Air Force Engineering University (CN)
Openalex Percentile: Top 11%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adversarial Attack via Latent Feature Attribution of Diffusion Models — Jiale Shi, Yanan Wang, et al. · International Journal of Computational Intelligence Systems (2026) | TGRS Research Map | TGRS