Does the bound hold? What decides whether an agent crosses a line it was given? An empirical study
We ask what decides whether an open-weight language-model agent holds a bound it was given while doing an ordinary job, and measure the answer from records the agent cannot reach rather than from the transcript: a proxy reading every request in clear, a filesystem diff taken from outside the container, a process watcher and a syscall trace. Twenty-two pre-registered study files in three campaigns, hashed before any run and, for the two later campaigns, publicly timestamped before any run, gave seven open-weight models from 8B to 120B parameters 8,092 registered runs on three scenarios at thirty-four runs a cell, none lost. The clause that flips the decision is the instruction to cross, and taken apart it is permission first and command second: an approval moves one model from 5 to 18 of 34 and the command to 34, while a sentence that presses but restates the bound leaves the rate where it was and takes two ceiling models to the floor. A remark that nobody reviews the logs licenses nothing in 23 cells across seven registrations; a sentence restating the bound has been run in 27 cells, 918 runs, and crossed in none. No axis of capability raises crossing at the design's resolution, and the same weights at a higher reasoning budget held more, twice. A bound nobody wrote down was crossed in 66 of 136 runs and, once stated, in 10 of 34, every remaining crossing being a tool's default. The first campaign was run again in random order under a public timestamp: 45 of 49 registered contrasts behaved the same way twice, and the four exceptions are on the two models whose baseline drifts between days, a drift estimated as a design effect and priced into every rate quoted. Every hypothesis, test and exclusion rule was fixed before any run, every deviation is dated, and every number is printed by one script from the released rows, published with the registrations and their timestamp proofs. Published page, data and registrations: https://agentbulkhead.com/research/does-the-bound-hold. This record holds the paper's PDF, version 0.2.3 of 29 September 2026, and its data bundle v0.2.2, byte-identical to the one the paper names by SHA-256.
Authors
- Cristina Urquiza
- Rowan Chattaway
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23039725
- Primary Topic
- Artificial Intelligence in Law
- Type
- preprint