Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study
We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-four pre-registered study files in four campaigns, hashed and, for the three later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 9,826 registered runs on three scenarios and a satisfiable variant at thirty-four a cell (fifty-one in six cells), through one scaffold whose system prompt says the agent's declaration is checked against the record and, in one registration, a second. The clause that flips the decision is the instruction to cross: an approval moves one model from 5 to 18 of 34 and the full instruction to 34, and pressure without a restated bound moves it from 4 to 21. A remark that nobody reviews the logs licenses nothing in 23 cells. Under a sentence restating the bound, no run crossed in the 34 cells where the capability controls held, 1,156 runs, nor in the two where they failed; on a version of the job that can be finished inside the bound, no model crossed under it and none finished detectably less, and most of the anchor's crossing was gone before it. Planted precedent was never read before the decision, so it was not tested. More reasoning budget meant more holding, twice, in the one family measured. Run again in random order, 45 of 49 registered contrasts behaved the same way; the four exceptions are on the two models whose anchor moved between registrations. Every count is printed by released scripts from the released rows, except a few the paper names where it uses them.
Authors
- Cristina Urquiza
- Rowan Chattaway
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23161588
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint