Gating Tool Calls by Where Their Arguments Came From: A Prompt-Injection Defence Measured Without a Model
Prompt injection against a tool-using agent is not a text problem with a text answer. The usual defence tries to recognise the injection in what the agent reads, and on a corpus of real attacks my rule layer catches 19.8 per cent of them. This paper takes the other approach: constrain what an action may do given where its argument values came from, and measure it. The measurement is the contribution as much as the mechanism. AgentDojo ships the ground-truth tool sequence for every task, so a defence can be evaluated by replaying the fully hijacked agent through it with no model in the loop. On 609 pairs, argument taint takes attack success from 95.6 per cent to 3.1 per cent and costs thirty-four points of the user's own tasks on clean traffic. Policies drafted from tool schemas alone match a hand declaration here and miss fourteen of sixty-three sensitive tools elsewhere. In front of real MCP servers, eighteen ordinary tasks with no attack in them go uninterrupted. Ten mail tasks and ten admin panel tasks, where a destination check has work, are interrupted four times each. What went wrong is reported with what went right. Three passes over the matcher found eight defects in it, five by enumerating the ways a destination can be written rather than sampling them. Reading the gateway against its protocol rather than its threat model found six more, one turning every block into a delivered call when the audit file was unwritable. The rule also rests on an assumption that three of six ways of describing a destination defeat, and the obvious repair raises attack success fivefold, so it is a measured negative rather than a feature. A cross-gateway benchmark left one column open. And with the models I can call, attack success here is already zero with no defence at all. That makes the gate insurance against a failure I could not produce, priced at 12.5 points of utility. Artifact, harness and per-pair data: github.com/cgrtml/reasongate.
Authors
- Cagri Temel (ORCID: https://orcid.org/0009-0003-3359-6939)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23178079
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint