Adversarial Attack and Detection in Token‐Pruned Large Vision‐Language Models

ABSTRACT Large vision‐language models (LVLMs) incur high computational costs, and visual token pruning is commonly used to remove redundant tokens and enhance efficiency. However, from a security perspective, such acceleration mechanisms may have a non‐monotonic impact on adversarial robustness: They may either improve robustness by removing perturbation‐related evidence or reduce robustness by discarding semantically important visual tokens. In this paper, we first analyze the setting‐dependent robustness effects of dynamically pruned LVLMs. Building on this analysis, we propose an attention‐flip attack that manipulates pruning‐related attention redistribution by suppressing high‐importance visual tokens while promoting low‐importance regions, thereby increasing the likelihood that discriminative visual evidence is discarded during inference. To counter this threat, we further propose a micro‐feature‐based adversarial detector. Rather than relying on global attention patterns, the detector captures local distortions around the pruning boundary using a multidimensional feature representation, including threshold‐neighborhood density, truncated residual energy ratio, mask spatial dispersion, and statistical descriptors. Experimental results show that, under the same perturbation budget and iterative setting, the proposed attack consistently achieves higher attack success rates (ASRs) than standard first‐order attack baselines on pruned Qwen2‐VL and LLaVA‐OneVision models. For adversarial examples generated by the proposed Attention‐Flip attack, the detector achieves accuracies of 94.79% on Qwen2‐VL and 93.58% on LLaVA‐OneVision. It further reaches 90.37% on standard PGD examples generated from an unpruned Qwen2‐VL model.

Authors

Institutions

Publication Details

Journal
Concurrency and Computation Practice and Experience
Published
2026-09-18
DOI
https://doi.org/10.1002/cpe.70951
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adversarial Attack and Detection in Token‐Pruned Large Vision‐Language Models

Hui Sun, Lefeng Zhang, Su Shi
Concurrency and Computation Practice and Experience
Adversarial Robustness in Machine Learning
article

Adversarial Attack and Detection in Token‐Pruned Large Vision‐Language Models

Hui Sun, Lefeng Zhang, Su Shi
article en

Abstract

ABSTRACT Large vision‐language models (LVLMs) incur high computational costs, and visual token pruning is commonly used to remove redundant tokens and enhance efficiency. However, from a security perspective, such acceleration mechanisms may have a non‐monotonic impact on adversarial robustness: They may either improve robustness by removing perturbation‐related evidence or reduce robustness by discarding semantically important visual tokens. In this paper, we first analyze the setting‐dependent robustness effects of dynamically pruned LVLMs. Building on this analysis, we propose an attention‐flip attack that manipulates pruning‐related attention redistribution by suppressing high‐importance visual tokens while promoting low‐importance regions, thereby increasing the likelihood that discriminative visual evidence is discarded during inference. To counter this threat, we further propose a micro‐feature‐based adversarial detector. Rather than relying on global attention patterns, the detector captures local distortions around the pruning boundary using a multidimensional feature representation, including threshold‐neighborhood density, truncated residual energy ratio, mask spatial dispersion, and statistical descriptors. Experimental results show that, under the same perturbation budget and iterative setting, the proposed attack consistently achieves higher attack success rates (ASRs) than standard first‐order attack baselines on pruned Qwen2‐VL and LLaVA‐OneVision models. For adversarial examples generated by the proposed Attention‐Flip attack, the detector achieves accuracies of 94.79% on Qwen2‐VL and 93.58% on LLaVA‐OneVision. It further reaches 90.37% on standard PGD examples generated from an unpruned Qwen2‐VL model.

Concurrency and Computation Practice and ExperienceVol. 38(19)
City University of Macau (MO)
Reduced inequalities
Openalex Percentile: Top 9%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adversarial Attack and Detection in Token‐Pruned Large Vision‐Language Models — Hui Sun, Lefeng Zhang, et al. · Concurrency and Computation Practice and Experience (2026) | TGRS Research Map | TGRS