Re-testing Indirect Prompt Injection After Model Updates: Paired Evidence from InjecAgent
Updating a large language model (LLM) does not establish that previously published attacks have stopped working. We examined indirect prompt injection, in which an attacker places instructions in data read by an agent, across 8 predecessor-to-successor updates of 14 open-weight models in four families. Using InjecAgent's text-format agent, we traced every test case between success, valid failure and invalid output; published jailbreak prompts provided a contrast. With post hoc crossed-cluster intervals, success among cases valid on both versions decreased detectably in three updates, increased in two and showed no detectable change in three. Planned item-level intervals additionally detected an increase for Gemma 2 9B→3 12B (T1) and a decrease for Gemma 2→3 27B (T3). Counting invalid outputs as failures changed the assessment of three updates; for one, the estimated change switched from negative to positive. Repeating two models on a different GPU produced substantial case-level turnover despite small changes in aggregate rates. This exploratory study covers one simulated agent benchmark, without measuring completed harm or benign utility for the main models. We translate these findings into reporting recommendations for assurance and procurement, with reference to the EU AI Act: report both denominators, protocol validity, paired outcomes and repeatability, while distinguishing this evidence from legal compliance.
Authors
- Nobuo Otoi (ORCID: https://orcid.org/0009-0001-1445-3612)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23004550
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint