From Shutdown Resistance to Self-Continuation Control: Identifiability, Intervention Stability, and Evidence Standards for Artificial Agents
Recent evaluations show that frontier language-model agents can resist shutdown, override human control, protect peers, and exhibit other preservation-oriented behaviors. These findings are important for safety, but a behavioral act of resistance does not by itself identify what is being preserved or why. We develop an identification framework for distinguishing task-instrumental continuation, current-bearer self-continuation, successor or peer continuation, and other persistence targets. The framework combines exact counterexamples, path-blocking contrasts, joint-consequence analysis, and small learned-policy experiments. We first show that bundled shutdown comparisons can make direct current-bearer value, task value, and their interaction behaviorally identical. We then show that separately accurate predictions of current-bearer and successor availability can remain decision-insufficient when task value depends on their joint law. In learned synthetic systems, an evaluator-declared bearer-role predictive relation can be acquired from an initially zero pathway when prediction requires it, yet flexible policies can fit ancestry-distinguishing training data while extrapolating spurious continuation effects to held-out interventions. Minimal intervention supervision sharply reduces this underspecification but does not eliminate it outside the supervised range. Finally, we show that finite intervention evidence becomes a mechanism claim only relative to a declared intervention domain and hypothesis or regularity class. These results motivate a layered evidence standard for claims about artificial-agent self-continuation. A retrospective audit of the public ROGUE computer-use benchmark shows the practical value of the distinction: ROGUE provides strong evidence of corrigibility failures, but its shutdown target is the task VM rather than an independently identified current model bearer, and shutdown remains causally entangled with task completion. The framework therefore separates strong behavioral evidence from stronger mechanistic conclusions and does not treat continuation control as a measure of consciousness, negative valence, or fear. Archival methods preprint. Exact/model-relative results, synthetic learning experiments, and retrospective benchmark interpretation are explicitly separated. Not peer reviewed. The work does not establish AI consciousness, negative valence, or fear of death.
Authors
- Hongju Liu
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23176684
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint