Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection
Abstract In this study, we explore visual prompt injection (VPI), a method that exploits the ability of large vision-language models (LVLMs) to follow instructions embedded within input images. LVLMs, such as GPT-4V, are designed to interpret and respond to visual prompts. Although this capability enables valuable applications, it introduces substantial security risks. To address this issue, we propose a new VPI method called “goal hijacking via visual prompt injection” (GHVPI). This method redirects LVLMs from their original execution task to an alternative task specified by an attacker. Our quantitative analysis reveals that GPT-4V is vulnerable to GHVPI, with an attack success rate of 15.8%, posing a significant security risk. Our results also reveal that GHVPI success relies on the character recognition and instruction-following capabilities of LVLMs. These findings highlight a critical vulnerability, raising concerns regarding the secure deployment of LVLMs in practical applications. Future research should prioritize developing robust mitigation strategies to address these vulnerabilities and facilitate the secure integration of LVLMs into sensitive systems. By addressing this underexplored security challenge, our research aims to encourage further research into the responsible use of LVLMs.
Authors
- Ryota Tanaka (ORCID: https://orcid.org/0000-0003-3158-1955)
- Shumpei Miyawaki
- Keisuke Sakaguchi (ORCID: https://orcid.org/0000-0002-3809-1732)
- Subaru Kimura (ORCID: https://orcid.org/0009-0000-5615-017X)
- Jun Suzuki
Institutions
- Tohoku University (JP)
- NTT (Japan) (JP)
Publication Details
- Journal
- New Generation Computing
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s00354-026-00334-8
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00