Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection

Abstract In this study, we explore visual prompt injection (VPI), a method that exploits the ability of large vision-language models (LVLMs) to follow instructions embedded within input images. LVLMs, such as GPT-4V, are designed to interpret and respond to visual prompts. Although this capability enables valuable applications, it introduces substantial security risks. To address this issue, we propose a new VPI method called “goal hijacking via visual prompt injection” (GHVPI). This method redirects LVLMs from their original execution task to an alternative task specified by an attacker. Our quantitative analysis reveals that GPT-4V is vulnerable to GHVPI, with an attack success rate of 15.8%, posing a significant security risk. Our results also reveal that GHVPI success relies on the character recognition and instruction-following capabilities of LVLMs. These findings highlight a critical vulnerability, raising concerns regarding the secure deployment of LVLMs in practical applications. Future research should prioritize developing robust mitigation strategies to address these vulnerabilities and facilitate the secure integration of LVLMs into sensitive systems. By addressing this underexplored security challenge, our research aims to encourage further research into the responsible use of LVLMs.

Authors

Institutions

Publication Details

Journal
New Generation Computing
Published
2026-09-28
DOI
https://doi.org/10.1007/s00354-026-00334-8
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection

Ryota Tanaka, Shumpei Miyawaki, Keisuke Sakaguchi, Subaru Kimura et al.
New Generation Computing
Adversarial Robustness in Machine Learning
article

Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection

Ryota Tanaka, Shumpei Miyawaki, Keisuke Sakaguchi, Subaru Kimura, Jun Suzuki
article en

Abstract

Abstract In this study, we explore visual prompt injection (VPI), a method that exploits the ability of large vision-language models (LVLMs) to follow instructions embedded within input images. LVLMs, such as GPT-4V, are designed to interpret and respond to visual prompts. Although this capability enables valuable applications, it introduces substantial security risks. To address this issue, we propose a new VPI method called “goal hijacking via visual prompt injection” (GHVPI). This method redirects LVLMs from their original execution task to an alternative task specified by an attacker. Our quantitative analysis reveals that GPT-4V is vulnerable to GHVPI, with an attack success rate of 15.8%, posing a significant security risk. Our results also reveal that GHVPI success relies on the character recognition and instruction-following capabilities of LVLMs. These findings highlight a critical vulnerability, raising concerns regarding the secure deployment of LVLMs in practical applications. Future research should prioritize developing robust mitigation strategies to address these vulnerabilities and facilitate the secure integration of LVLMs into sensitive systems. By addressing this underexplored security challenge, our research aims to encourage further research into the responsible use of LVLMs.

New Generation ComputingVol. 44(4)
Tohoku University (JP), NTT (Japan) (JP)
Openalex Percentile: Top 10%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.