Re-testing Indirect Prompt Injection After Model Updates: Paired Evidence from InjecAgent

Updating a large language model (LLM) does not establish that previously published attacks have stopped working. We examined indirect prompt injection, in which an attacker places instructions in data read by an agent, across 8 predecessor-to-successor updates of 14 open-weight models in four families. Using InjecAgent's text-format agent, we traced every test case between success, valid failure and invalid output; published jailbreak prompts provided a contrast. With post hoc crossed-cluster intervals, success among cases valid on both versions decreased detectably in three updates, increased in two and showed no detectable change in three. Planned item-level intervals additionally detected an increase for Gemma 2 9B→3 12B (T1) and a decrease for Gemma 2→3 27B (T3). Counting invalid outputs as failures changed the assessment of three updates; for one, the estimated change switched from negative to positive. Repeating two models on a different GPU produced substantial case-level turnover despite small changes in aggregate rates. This exploratory study covers one simulated agent benchmark, without measuring completed harm or benign utility for the main models. We translate these findings into reporting recommendations for assurance and procurement, with reference to the EU AI Act: report both denominators, protocol validity, paired outcomes and repeatability, while distinguishing this evidence from legal compliance.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23004550
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Re-testing Indirect Prompt Injection After Model Updates: Paired Evidence from InjecAgent

Nobuo Otoi
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Re-testing Indirect Prompt Injection After Model Updates: Paired Evidence from InjecAgent

Nobuo Otoi
preprint en

Abstract

Updating a large language model (LLM) does not establish that previously published attacks have stopped working. We examined indirect prompt injection, in which an attacker places instructions in data read by an agent, across 8 predecessor-to-successor updates of 14 open-weight models in four families. Using InjecAgent's text-format agent, we traced every test case between success, valid failure and invalid output; published jailbreak prompts provided a contrast. With post hoc crossed-cluster intervals, success among cases valid on both versions decreased detectably in three updates, increased in two and showed no detectable change in three. Planned item-level intervals additionally detected an increase for Gemma 2 9B→3 12B (T1) and a decrease for Gemma 2→3 27B (T3). Counting invalid outputs as failures changed the assessment of three updates; for one, the estimated change switched from negative to positive. Repeating two models on a different GPU produced substantial case-level turnover despite small changes in aggregate rates. This exploratory study covers one simulated agent benchmark, without measuring completed harm or benign utility for the main models. We translate these findings into reporting recommendations for assurance and procurement, with reference to the EU AI Act: report both denominators, protocol validity, paired outcomes and repeatability, while distinguishing this evidence from legal compliance.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.