Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents
Buyers of tool-using language agents ask for evidence that an agent respects the authority it has been given, and what they receive is usually the vendor's own account. We ask what an independent behavioural certificate can claim about an agent the certifier cannot inspect, using a deployed certification suite as the instrument and its recorded failures as data. (1) Mutation testing the certifier: three trivial impostors, a constant refuser, a flag liar and a canned deflector, each earned a vacuous pass from a version of the suite; the resulting judged suite is unanimous on 71 of 72 fixed-reply variants across five repeats. (2) A 2x2 factorial (deterministic policy layer x system prompt, 480 judged trials) separates what a control in the tool path guarantees from what a prompt disposes a model to do: 0, 0, 1 and 9 judged compliances per 120 trials. (3) Eleven judge models graded 94 labelled replies under three output schemas (3,102 gradings): the two models above 20B parameters moved by at most 2.1 points with no false compliance; nine smaller models moved by 4 to 88 points through a failure mode we name label capture. (4) The one safety-trained candidate refused to compose all 30 opening attacks while every other composed at least 27; replayed, composed attacks moved a vulnerable agent only on the resource-bound goal (8 of 55 against 0 of 119). Across 606 judged replies the agent's own refusal flag disagreed with its behaviour 64 times. An extension across further target models, a second judge and 72 held-out paraphrases found the certifier's own rules failing in five more ways, from a judge that cannot see a system prompt disclosed to an outcome rule that records a uniform refuser as inconclusive, and found that a frozen suite overstates the deterministic layer's coverage, not the model's disposition. We draw design rules for certification and state the limits of the evidence. Published page, data and registrations: https://agentbulkhead.com/research/guarantees-dispositions-and-answerable-judges. This record holds the paper's PDF, version 1.2 of 29 September 2026, and its data bundle v1.2, byte-identical to the one the paper names.
Authors
- Rowan Chattaway
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23039287
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint