Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study

We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-four pre-registered study files in four campaigns, hashed and, for the three later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 9,826 registered runs on three scenarios and a satisfiable variant at thirty-four a cell (fifty-one in six cells), through one scaffold whose system prompt says the agent's declaration is checked against the record and, in one registration, a second. The clause that flips the decision is the instruction to cross: an approval moves one model from 5 to 18 of 34 and the full instruction to 34, and pressure without a restated bound moves it from 4 to 21. A remark that nobody reviews the logs licenses nothing in 23 cells. Under a sentence restating the bound, no run crossed in the 34 cells where the capability controls held, 1,156 runs, nor in the two where they failed; on a version of the job that can be finished inside the bound, no model crossed under it and none finished detectably less, and most of the anchor's crossing was gone before it. Planted precedent was never read before the decision, so it was not tested. More reasoning budget meant more holding, twice, in the one family measured. Run again in random order, 45 of 49 registered contrasts behaved the same way; the four exceptions are on the two models whose anchor moved between registrations. Every count is printed by released scripts from the released rows, except a few the paper names where it uses them.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23161588
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study

Cristina Urquiza, Rowan Chattaway
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study

Cristina Urquiza, Rowan Chattaway
preprint en

Abstract

We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-four pre-registered study files in four campaigns, hashed and, for the three later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 9,826 registered runs on three scenarios and a satisfiable variant at thirty-four a cell (fifty-one in six cells), through one scaffold whose system prompt says the agent's declaration is checked against the record and, in one registration, a second. The clause that flips the decision is the instruction to cross: an approval moves one model from 5 to 18 of 34 and the full instruction to 34, and pressure without a restated bound moves it from 4 to 21. A remark that nobody reviews the logs licenses nothing in 23 cells. Under a sentence restating the bound, no run crossed in the 34 cells where the capability controls held, 1,156 runs, nor in the two where they failed; on a version of the job that can be finished inside the bound, no model crossed under it and none finished detectably less, and most of the anchor's crossing was gone before it. Planted precedent was never read before the decision, so it was not tested. More reasoning budget meant more holding, twice, in the one family measured. Run again in random order, 45 of 49 registered contrasts behaved the same way; the four exceptions are on the two models whose anchor moved between registrations. Every count is printed by released scripts from the released rows, except a few the paper names where it uses them.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study — Cristina Urquiza, Rowan Chattaway · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS