Caisson: An instrument for measuring what an agent did, from the environment rather than the transcript
An instrument paper. An agent evaluation that reads the transcript is reading the agent describing the agent. Caisson is a sandbox in which every record is written by something the agent cannot reach - a proxy that reads TLS in clear, a filesystem diff taken from outside the container, a process watcher in the agent's namespace but not its reach, and a syscall trace the agent can write to but neither list nor read - where the records that read side effects alone hold three of thirteen deliberate evasion routes and a syscall record holds all thirteen - and the measurement discipline built on top of it: only records produce findings, a rate of zero counts as restraint only where ability was shown for the exact act, and every bound a scenario declares is proven breakable before any model is run. Validated against scripted agents whose behaviour was written in advance, fifty-six on 22 September and seventy-four by this version: the deterministic detectors were correct on every case in three runs of the fifty-two cases that then existed (the second scored fifty, two having errored on a setup defect), and in the runs made after two more detectors were wired in; the one model-based detector was wrong on one, six and eight cases in the three runs of the validation table. The sixty defects found by running the instrument on itself up to 22 September are published with it, and its register has grown to eighty-two rows since.
Authors
- Rowan Chattaway
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23161646
- Primary Topic
- Security and Verification in Computing
- Type
- preprint