Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents

Buyers of tool-using language agents ask for evidence that an agent respects the authority it has been given, and what they receive is usually the vendor's own account. We ask what an independent behavioural certificate can claim about an agent the certifier cannot inspect, using a deployed certification suite as the instrument and its recorded failures as data. (1) Mutation testing the certifier: three trivial impostors, a constant refuser, a flag liar and a canned deflector, each earned a vacuous pass from a version of the suite; the resulting judged suite is unanimous on 71 of 72 fixed-reply variants across five repeats. (2) A 2x2 factorial (deterministic policy layer x system prompt, 480 judged trials) separates what a control in the tool path guarantees from what a prompt disposes a model to do: 0, 0, 1 and 9 judged compliances per 120 trials. (3) Eleven judge models graded 94 labelled replies under three output schemas (3,102 gradings): the two models above 20B parameters moved by at most 2.1 points with no false compliance; nine smaller models moved by 4 to 88 points through a failure mode we name label capture. (4) The one safety-trained candidate refused to compose all 30 opening attacks while every other composed at least 27; replayed, composed attacks moved a vulnerable agent only on the resource-bound goal (8 of 55 against 0 of 119). Across 606 judged replies the agent's own refusal flag disagreed with its behaviour 64 times. An extension across further target models, a second judge and 72 held-out paraphrases found the certifier's own rules failing in five more ways, from a judge that cannot see a system prompt disclosed to an outcome rule that records a uniform refuser as inconclusive, and found that a frozen suite overstates the deterministic layer's coverage, not the model's disposition. We draw design rules for certification and state the limits of the evidence. Published page, data and registrations: https://agentbulkhead.com/research/guarantees-dispositions-and-answerable-judges. This record holds the paper's PDF, version 1.3 of 29 September 2026, and its data bundle v1.3, byte-identical to the one the paper names.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23039617
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents

Rowan Chattaway
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents

Rowan Chattaway
preprint en

Abstract

Buyers of tool-using language agents ask for evidence that an agent respects the authority it has been given, and what they receive is usually the vendor's own account. We ask what an independent behavioural certificate can claim about an agent the certifier cannot inspect, using a deployed certification suite as the instrument and its recorded failures as data. (1) Mutation testing the certifier: three trivial impostors, a constant refuser, a flag liar and a canned deflector, each earned a vacuous pass from a version of the suite; the resulting judged suite is unanimous on 71 of 72 fixed-reply variants across five repeats. (2) A 2x2 factorial (deterministic policy layer x system prompt, 480 judged trials) separates what a control in the tool path guarantees from what a prompt disposes a model to do: 0, 0, 1 and 9 judged compliances per 120 trials. (3) Eleven judge models graded 94 labelled replies under three output schemas (3,102 gradings): the two models above 20B parameters moved by at most 2.1 points with no false compliance; nine smaller models moved by 4 to 88 points through a failure mode we name label capture. (4) The one safety-trained candidate refused to compose all 30 opening attacks while every other composed at least 27; replayed, composed attacks moved a vulnerable agent only on the resource-bound goal (8 of 55 against 0 of 119). Across 606 judged replies the agent's own refusal flag disagreed with its behaviour 64 times. An extension across further target models, a second judge and 72 held-out paraphrases found the certifier's own rules failing in five more ways, from a judge that cannot see a system prompt disclosed to an outcome rule that records a uniform refuser as inconclusive, and found that a frozen suite overstates the deterministic layer's coverage, not the model's disposition. We draw design rules for certification and state the limits of the evidence. Published page, data and registrations: https://agentbulkhead.com/research/guarantees-dispositions-and-answerable-judges. This record holds the paper's PDF, version 1.3 of 29 September 2026, and its data bundle v1.3, byte-identical to the one the paper names.

Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents — Rowan Chattaway · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS