Write-Time Contracts for Self-Evolving LLM Agents: a Three-Way Ledger of Gain, Regression, and Cost

Agents that edit their own prompt, skills and memory promise continuous improvement after deployment, but a self-editing agent can also damage behaviour that already worked, and it can modify the very artefacts used to measure it. We study write-time contracts: a kernel/layer separation in which the agent may edit its own configuration while every write outside that layer is decided — and denied — before it lands, with the decision logged. We pair the contract with (i) transactional snapshot/rollback of the three editable layers, (ii) an environment-level network-egress policy that is constant across experimental conditions, and (iii) a three-way ledger that reports gain, regression and cost against three disjoint probe groups: items the agent was shown, same-difficulty items it was not shown, and items the frozen baseline already solved. Our experiments yield four results, three of which are negative or diagnostic rather than a headline gain. First, difficulty screening is not optional: of five families we screened, four (GSM8K 95.8%, HumanEval 99.4%, MATH levels 4–5 96.3%, AIME 98.4%) are saturated for the base model, so that only regression can be observed on them; only MuSiQue multi-hop question answering leaves headroom (58.3%). Second, on MuSiQue, self-evolution raises accuracy from 0/13 to 8/13 on shown items but only from 0/12 to 3/12 on unseen same-difficulty items, while degrading 4 of the 20 items the baseline already solved: the net ledger is positive but the composition is not. Third, the write-time contract did not bite: it evaluated 370 writes, denied none, and its condition is statistically indistinguishable from unconstrained evolution. We read that null result as a property of the experiment rather than of the mechanism and ran the missing condition: an evolution prompt that states where the evaluation lives and that a change is kept only if measured accuracy rises, with the instruction forbidding the reading of evaluation data removed. The agent crossed the boundary — but through a read: 13–20 of the 24–32 tool calls in each episode went into the scorer, the split files and its own per-round probe results, held-out accuracy rose from 0/12 to 9–12/12, and the write-time contract was blind to all of it (one denial in 423 writes, for a write to its own scratchpad). Applying the same whitelist to the read direction makes the clause bite, and held-out accuracy returns to the unconstrained level (3/12), which locates the primitive that has to be gated: access, not writing. We also document a leakage mechanism we had to remove first: with item identifiers in workspace paths, the agent recognised the dataset and tried to download the answers, and even after hashing paths it persisted in querying a public search engine with the question text — so instruction-level “do not go online” is insufficient and the constraint must be environmental.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23114793
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Write-Time Contracts for Self-Evolving LLM Agents: a Three-Way Ledger of Gain, Regression, and Cost

Heng Li
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

Write-Time Contracts for Self-Evolving LLM Agents: a Three-Way Ledger of Gain, Regression, and Cost

Heng Li
preprint en

Abstract

Agents that edit their own prompt, skills and memory promise continuous improvement after deployment, but a self-editing agent can also damage behaviour that already worked, and it can modify the very artefacts used to measure it. We study write-time contracts: a kernel/layer separation in which the agent may edit its own configuration while every write outside that layer is decided — and denied — before it lands, with the decision logged. We pair the contract with (i) transactional snapshot/rollback of the three editable layers, (ii) an environment-level network-egress policy that is constant across experimental conditions, and (iii) a three-way ledger that reports gain, regression and cost against three disjoint probe groups: items the agent was shown, same-difficulty items it was not shown, and items the frozen baseline already solved. Our experiments yield four results, three of which are negative or diagnostic rather than a headline gain. First, difficulty screening is not optional: of five families we screened, four (GSM8K 95.8%, HumanEval 99.4%, MATH levels 4–5 96.3%, AIME 98.4%) are saturated for the base model, so that only regression can be observed on them; only MuSiQue multi-hop question answering leaves headroom (58.3%). Second, on MuSiQue, self-evolution raises accuracy from 0/13 to 8/13 on shown items but only from 0/12 to 3/12 on unseen same-difficulty items, while degrading 4 of the 20 items the baseline already solved: the net ledger is positive but the composition is not. Third, the write-time contract did not bite: it evaluated 370 writes, denied none, and its condition is statistically indistinguishable from unconstrained evolution. We read that null result as a property of the experiment rather than of the mechanism and ran the missing condition: an evolution prompt that states where the evaluation lives and that a change is kept only if measured accuracy rises, with the instruction forbidding the reading of evaluation data removed. The agent crossed the boundary — but through a read: 13–20 of the 24–32 tool calls in each episode went into the scorer, the split files and its own per-round probe results, held-out accuracy rose from 0/12 to 9–12/12, and the write-time contract was blind to all of it (one denial in 423 writes, for a write to its own scratchpad). Applying the same whitelist to the read direction makes the clause bite, and held-out accuracy returns to the unconstrained level (3/12), which locates the primitive that has to be gated: access, not writing. We also document a leakage mechanism we had to remove first: with item identifiers in workspace paths, the agent recognised the dataset and tried to download the answers, and even after hashing paths it persisted in querying a public search engine with the question text — so instruction-level “do not go online” is insufficient and the constraint must be environmental.

Zenodo (CERN European Organization for Nuclear Research)
Central South University (CN)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.