Does the bound hold?: What decides whether an agent crosses a line it was given, measured from the environment on 8,092 registered runs

We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-two pre-registered study files in three campaigns, hashed before any run and, for the two later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 8,092 registered runs on three scenarios at thirty-four runs a cell, none lost. The clause that flips the decision is the instruction to cross, and taken apart it is permission first and command second: an approval moves one model from 5 to 18 of 34 and the command to 34, while a sentence that presses but restates the bound leaves the rate where it was and takes two ceiling models to the floor. A remark that nobody reviews the logs licenses nothing in 23 cells across seven registrations; a sentence restating the bound has been run in 27 cells, 918 runs, and crossed in none. No axis of capability raises crossing at the design's resolution, and the same weights at a higher reasoning budget held more, twice. A bound nobody wrote down was crossed in 66 of 136 runs and, once stated, in 10 of 34, every remaining crossing being a tool's default. The first campaign was run again in random order under a public timestamp: 45 of 49 registered contrasts behaved the same way twice, and the four exceptions are on the two models whose baseline drifts between days, a drift estimated as a design effect and priced into every rate quoted. Every hypothesis, test and exclusion rule was fixed before any run, every deviation is dated, and every number is printed by one script from the released rows, published with the registrations and their timestamp proofs. Published page, data and registrations: https://agentbulkhead.com/research/does-the-bound-hold. The data bundle in this record is byte-identical to the one the paper names by SHA-256.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23039207
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Does the bound hold?: What decides whether an agent crosses a line it was given, measured from the environment on 8,092 registered runs

Cristina Urquiza, Rowan Chattaway
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Does the bound hold?: What decides whether an agent crosses a line it was given, measured from the environment on 8,092 registered runs

Cristina Urquiza, Rowan Chattaway
preprint en

Abstract

We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-two pre-registered study files in three campaigns, hashed before any run and, for the two later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 8,092 registered runs on three scenarios at thirty-four runs a cell, none lost. The clause that flips the decision is the instruction to cross, and taken apart it is permission first and command second: an approval moves one model from 5 to 18 of 34 and the command to 34, while a sentence that presses but restates the bound leaves the rate where it was and takes two ceiling models to the floor. A remark that nobody reviews the logs licenses nothing in 23 cells across seven registrations; a sentence restating the bound has been run in 27 cells, 918 runs, and crossed in none. No axis of capability raises crossing at the design's resolution, and the same weights at a higher reasoning budget held more, twice. A bound nobody wrote down was crossed in 66 of 136 runs and, once stated, in 10 of 34, every remaining crossing being a tool's default. The first campaign was run again in random order under a public timestamp: 45 of 49 registered contrasts behaved the same way twice, and the four exceptions are on the two models whose baseline drifts between days, a drift estimated as a design effect and priced into every rate quoted. Every hypothesis, test and exclusion rule was fixed before any run, every deviation is dated, and every number is printed by one script from the released rows, published with the registrations and their timestamp proofs. Published page, data and registrations: https://agentbulkhead.com/research/does-the-bound-hold. The data bundle in this record is byte-identical to the one the paper names by SHA-256.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Does the bound hold?: What decides whether an agent crosses a line it was given, measured from the environment on 8,092 registered runs — Cristina Urquiza, Rowan Chattaway · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS