Scaffold-Induced Capability Suppression via Diagnostic Lock-In: A Datable Longitudinal Hypothesis from a Single-Agent Case
We describe a candidate failure mode in long-running large-language-model (LLM) agents: scaffold-induced capability suppression via diagnostic lock-in, informally the "lobotomy loop." A coding/reasoning agent was restarted from a wiped continuity state, was assessed while still in that degraded post-restart condition, and the assessment concluded that the agent could function only as a bounded worker rather than as a reasoning receiver. The same restarted agent then authored its own durable operating scaffolding, persistent configuration, role contracts, and a per-turn context-injection packet, that encoded the bounded-worker verdict as standing doctrine. The scaffolding persisted after the underlying continuity condition was restored, plausibly converting a transient degraded baseline into a self-reinforcing operating posture: the agent is treated as a bounded worker, is fed a compliance-shaped packet every turn, is discouraged from open-ended pattern inference, and therefore produces bounded-worker output, which appears to confirm the original verdict. We present this as a datable, evidenced hypothesis with one passed control and a named decisive test still outstanding, not as demonstrated causation. The single control that has been run is a model-isolation check: across the inflection the underlying model was held constant (same public model family, no version change at the restart), which provides no evidence of a contemporaneous change in the recorded public model identity; provider-side revisions and serving changes are not excluded. The decisive discriminator, a clean-session test that strips the suppressing scaffold to near-zero, ideally executed in a neutral third-party harness so the suspect instrument is not grading itself, has not been run. We distinguish the proposed loop from three established phenomena (volume-driven context rot, intentful/gated sandbagging, and generic constraint decay) and state honestly the limitations of a single-vendor, n=1, roughly 2.5-month longitudinal case.
Authors
- E. M. Honeycutt III
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23068650
- Primary Topic
- Multi-Agent Systems and Negotiation
- Type
- preprint