Containment by Design: A Layered Architecture and Open Testbed for Containing, Detecting and Recovering from Misbehaving LLM Agents

LLM agents are now deployed with credentials, tools and long-running autonomy, and most 2026 incidents did not require a scheming model: an ordinary agent that cannot separate data from instructions was enough once composed into a system with too many permissions. This preprint argues for containment by design: building the harness first so it holds whether the agent is hijacked, confused or adversarial. It contributes (1) a layered architecture of nine independent walls, (2) a six-step detection and recapture playbook for when the walls fail, modelled on how high-security prisons handle escapes, (3) an open, dependency-free testbed with eleven adversarial scenarios and a knock-one-out ablation method, and (4) a mapping to the OWASP Top 10 for Agentic Applications. With all walls active every scenario is contained; with each primary wall removed a backup wall contains all ten adversarial cases; with every wall removed eight of eleven escape. The agents are scripted, so results validate harness logic and layer independence rather than attack success against real models. Code: https://github.com/Omazesoft/containment-by-design

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22880941
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Containment by Design: A Layered Architecture and Open Testbed for Containing, Detecting and Recovering from Misbehaving LLM Agents

Girish Chawla
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Containment by Design: A Layered Architecture and Open Testbed for Containing, Detecting and Recovering from Misbehaving LLM Agents

Girish Chawla
preprint en

Abstract

LLM agents are now deployed with credentials, tools and long-running autonomy, and most 2026 incidents did not require a scheming model: an ordinary agent that cannot separate data from instructions was enough once composed into a system with too many permissions. This preprint argues for containment by design: building the harness first so it holds whether the agent is hijacked, confused or adversarial. It contributes (1) a layered architecture of nine independent walls, (2) a six-step detection and recapture playbook for when the walls fail, modelled on how high-security prisons handle escapes, (3) an open, dependency-free testbed with eleven adversarial scenarios and a knock-one-out ablation method, and (4) a mapping to the OWASP Top 10 for Agentic Applications. With all walls active every scenario is contained; with each primary wall removed a backup wall contains all ten adversarial cases; with every wall removed eight of eleven escape. The agents are scripted, so results validate harness logic and layer independence rather than attack success against real models. Code: https://github.com/Omazesoft/containment-by-design

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Containment by Design: A Layered Architecture and Open Testbed for Containing, Detecting and Recovering from Misbehaving LLM Agents — Girish Chawla · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS