Kaifeng Prefecture: A Safe-Harness Architecture for Preventing Autonomous LLM-Agent Attacks
As autonomous large language model (LLM) agents gain tool-use and long-horizon planning capabilities, they present a growing risk of executing harmful actions without explicit human authorization. Existing safety measures—such as alignment training, prompt hardening, and execution-time moderation—often fail against multi-step adversarial strategies in which an initially benign request is progressively transformed into a dangerous outcome. We introduce Kaifeng Prefecture, a safe-harness architecture that prevents autonomous LLM-agent attacks through structural separation of roles and control-flow inversion. The architecture assigns evaluative, observational, predictive, and capability functions to distinct modules (Value Judge, Observer, Predictor, and Capability), routes high-risk actions through an external gate, and preserves a human interrupt channel. Three formal propositions and accompanying floor invariants characterize the conditions under which harmful commands are blocked before execution. To make the design concrete, we describe a 23-scenario prototype implemented with a TranscriptBackend, demonstrating how the harness intercepts and neutralizes attack chains that bypass conventional guardrails. The result is a safety architecture that limits the autonomy of LLM agents without requiring full symbolic verification of every possible plan.
Authors
- Fan He (ORCID: https://orcid.org/0000-0002-4145-9999)
Institutions
- China Institute of Finance and Capital Markets (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22760566
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint