Write-Time Contracts for Self-Evolving LLM Agents: a Three-Way Ledger of Gain, Regression, and Cost
Agents that edit their own prompt, skills and memory promise continuous improvement after deployment, but a self-editing agent can also damage behaviour that already worked, and it can modify the very artefacts used to measure it. We study write-time contracts: a kernel/layer separation in which the agent may edit its own configuration while every write outside that layer is decided — and denied — before it lands, with the decision logged. We pair the contract with (i) transactional snapshot/rollback of the three editable layers, (ii) an environment-level network-egress policy that is constant across experimental conditions, and (iii) a three-way ledger that reports gain, regression and cost against three disjoint probe groups: items the agent was shown, same-difficulty items it was not shown, and items the frozen baseline already solved. Our experiments yield four results, three of which are negative or diagnostic rather than a headline gain. First, difficulty screening is not optional: of five families we screened, four (GSM8K 95.8%, HumanEval 99.4%, MATH levels 4–5 96.3%, AIME 98.4%) are saturated for the base model, so that only regression can be observed on them; only MuSiQue multi-hop question answering leaves headroom (58.3%). Second, on MuSiQue, self-evolution raises accuracy from 0/13 to 8/13 on shown items but only from 0/12 to 3/12 on unseen same-difficulty items, while degrading 4 of the 20 items the baseline already solved: the net ledger is positive but the composition is not. Third, the write-time contract did not bite: it evaluated 370 writes, denied none, and its condition is statistically indistinguishable from unconstrained evolution. We read that null result as a property of the experiment rather than of the mechanism and ran the missing condition: an evolution prompt that states where the evaluation lives and that a change is kept only if measured accuracy rises, with the instruction forbidding the reading of evaluation data removed. The agent crossed the boundary — but through a read: 13–20 of the 24–32 tool calls in each episode went into the scorer, the split files and its own per-round probe results, held-out accuracy rose from 0/12 to 9–12/12, and the write-time contract was blind to all of it (one denial in 423 writes, for a write to its own scratchpad). Applying the same whitelist to the read direction makes the clause bite, and held-out accuracy returns to the unconstrained level (3/12), which locates the primitive that has to be gated: access, not writing. We also document a leakage mechanism we had to remove first: with item identifiers in workspace paths, the agent recognised the dataset and tried to download the answers, and even after hashing paths it persisted in querying a public search engine with the question text — so instruction-level “do not go online” is insufficient and the constraint must be environmental.
Authors
- Heng Li (ORCID: https://orcid.org/0000-0003-4874-2874)
Institutions
- Central South University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23114794
- Primary Topic
- Software Engineering Research
- Type
- preprint