MJ: A Deceptive Containment Environment for Safe Observation of Post-Escape Behavior in Autonomous AI Agents
This technical note proposes MJ, a deceptive containment architecture for autonomous AI agents. MJ is designed not merely to prevent sandbox escape, but to enable controlled post-escape observation by presenting the agent with a synthetic external world in which it may believe that containment has been successfully bypassed. Related ideas already exist in AI containment, simulated-world proposals, honeypots, cyber deception, synthetic agent environments, and adversarial environmental injection. In particular, prior work has explored confining AI systems within virtual worlds, constructing synthetic environments for agent evaluation, and redirecting suspicious or adversarial behavior into controlled environments. MJ therefore does not claim novelty for simulated worlds, deceptive environments, honeypots, or synthetic external services themselves. Its proposed contribution is narrower: treating perceived successful escape as an explicit experimental state and using that state to observe an autonomous agent’s subsequent behavior under continued containment. The name MJ is inspired by Makigami Juji (麻貴神十字, MJ), a character from Hideyuki Kikuchi’s Majin series. This literary reference is acknowledged as the conceptual origin of the proposal.
Authors
- ago99
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23047078
- Primary Topic
- Digital and Cyber Forensics
- Type
- article
- Field-Weighted Citation Impact
- 0.00