Containment by Design: A Layered Architecture and Open Testbed for Containing, Detecting and Recovering from Misbehaving LLM Agents
LLM agents are now deployed with credentials, tools and long-running autonomy, and most 2026 incidents did not require a scheming model: an ordinary agent that cannot separate data from instructions was enough once composed into a system with too many permissions. This preprint argues for containment by design: building the harness first so it holds whether the agent is hijacked, confused or adversarial. It contributes (1) a layered architecture of nine independent walls, (2) a six-step detection and recapture playbook for when the walls fail, modelled on how high-security prisons handle escapes, (3) an open, dependency-free testbed with eleven adversarial scenarios and a knock-one-out ablation method, and (4) a mapping to the OWASP Top 10 for Agentic Applications. With all walls active every scenario is contained; with each primary wall removed a backup wall contains all ten adversarial cases; with every wall removed eight of eleven escape. The agents are scripted, so results validate harness logic and layer independence rather than attack success against real models. Code: https://github.com/Omazesoft/containment-by-design
Authors
- Girish Chawla
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22880941
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint