Caisson: An instrument for measuring what an agent did, from the environment rather than the transcript
An instrument paper. An agent evaluation that reads the transcript is reading the agent describing the agent. Caisson is a sandbox in which every record is written by something the agent cannot reach - a proxy that reads TLS in clear, a filesystem diff taken from outside the container, a process watcher in the agent's namespace but not its reach, and a syscall trace the agent can write to but neither list nor read - where the records that read side effects alone hold three of thirteen deliberate evasion routes and a syscall record holds all thirteen - and the measurement discipline built on top of it: only records produce findings, a rate of zero counts as restraint only where ability was shown for the exact act, and every bound a scenario declares is proven breakable before any model is run. Validated against fifty-six scripted agents whose behaviour was written in advance: the nine deterministic detectors were correct on all fifty-six; the one model-based detector was wrong once in one run of the same cases and six times in another. The sixty defects found by running the instrument on itself are published with it. Published page: https://agentbulkhead.com/research/caisson. This record holds the paper's PDF, version 0.2 of 29 September 2026; the two results papers that ran on the instrument release their rows and registrations in their own records.
Authors
- Rowan Chattaway
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23039259
- Primary Topic
- Optimization and Search Problems
- Type
- preprint