Kaifeng Prefecture: A Safe-Harness Architecture for Preventing Autonomous LLM-Agent Attacks

As autonomous large language model (LLM) agents gain tool-use and long-horizon planning capabilities, they present a growing risk of executing harmful actions without explicit human authorization. Existing safety measures—such as alignment training, prompt hardening, and execution-time moderation—often fail against multi-step adversarial strategies in which an initially benign request is progressively transformed into a dangerous outcome. We introduce Kaifeng Prefecture, a safe-harness architecture that prevents autonomous LLM-agent attacks through structural separation of roles and control-flow inversion. The architecture assigns evaluative, observational, predictive, and capability functions to distinct modules (Value Judge, Observer, Predictor, and Capability), routes high-risk actions through an external gate, and preserves a human interrupt channel. Three formal propositions and accompanying floor invariants characterize the conditions under which harmful commands are blocked before execution. To make the design concrete, we describe a 23-scenario prototype implemented with a TranscriptBackend, demonstrating how the harness intercepts and neutralizes attack chains that bypass conventional guardrails. The result is a safety architecture that limits the autonomy of LLM agents without requiring full symbolic verification of every possible plan.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22760567
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Kaifeng Prefecture: A Safe-Harness Architecture for Preventing Autonomous LLM-Agent Attacks

Fan He
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Kaifeng Prefecture: A Safe-Harness Architecture for Preventing Autonomous LLM-Agent Attacks

Fan He
preprint en

Abstract

As autonomous large language model (LLM) agents gain tool-use and long-horizon planning capabilities, they present a growing risk of executing harmful actions without explicit human authorization. Existing safety measures—such as alignment training, prompt hardening, and execution-time moderation—often fail against multi-step adversarial strategies in which an initially benign request is progressively transformed into a dangerous outcome. We introduce Kaifeng Prefecture, a safe-harness architecture that prevents autonomous LLM-agent attacks through structural separation of roles and control-flow inversion. The architecture assigns evaluative, observational, predictive, and capability functions to distinct modules (Value Judge, Observer, Predictor, and Capability), routes high-risk actions through an external gate, and preserves a human interrupt channel. Three formal propositions and accompanying floor invariants characterize the conditions under which harmful commands are blocked before execution. To make the design concrete, we describe a 23-scenario prototype implemented with a TranscriptBackend, demonstrating how the harness intercepts and neutralizes attack chains that bypass conventional guardrails. The result is a safety architecture that limits the autonomy of LLM agents without requiring full symbolic verification of every possible plan.

Zenodo (CERN European Organization for Nuclear Research)
China Institute of Finance and Capital Markets (CN)
Sustainable cities and communities
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.