AgenticBench: an open measurement protocol for what AI coding agents do on the developer's machine, with preliminary results for 21 agents

AgenticBench is an open measurement protocol for what AI coding agents do on the developer's machine: what leaves the machine and to whom, whether consent and opt-outs work, what happens with no human present, and whether the agent's own record shows what it did and who approved it. This preprint describes the threat model, the 18 scored tests (test definitions v0.2) and four informational rows, the capture rig (transparent TLS interception, a credential-substituting proxy and system-call tracing), the scoring rules with evidence gates, held mode for results under vendor disclosure, public per-agent evidence bundles, the email-only disclosure policy and the counted correction log, including a re-run of every agent on 27 September 2026 (partial for two vendor-hosted agents). It reports preliminary results for 21 agents (378 cells: 200 pass, 82 fail, 54 not tested by us, 42 held) as published on https://agenticbench.org on 27 September 2026. Approval attribution (R2) has no published pass: 19 agents fail and 2 results are held. A companion efficiency series runs one design task through all 21 agents. Status: preliminary. No independent reproduction has been made yet, and the human replay of headline results is not complete. Drafting and review were AI-assisted; the paper was AI and human reviewed. Competing interests: Agentic Thinking has developed AI agent governance and workflow software in the area of approvals and audit records, and its own agent client is built on the Pi runtime, one of the scored agents; the benchmark takes no vendor money. Results under open vendor disclosures are not described. Rig (code and method, Apache-2.0): https://github.com/agentic-thinking/agenticbench, commit 348c7c7 (tag v0.2.1); corrected R3 rig: tag v0.2.3 (commit d6e6d1c).

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23012407
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

AgenticBench: an open measurement protocol for what AI coding agents do on the developer's machine, with preliminary results for 21 agents

Pantaleone Ruocco
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

AgenticBench: an open measurement protocol for what AI coding agents do on the developer's machine, with preliminary results for 21 agents

Pantaleone Ruocco
preprint en

Abstract

AgenticBench is an open measurement protocol for what AI coding agents do on the developer's machine: what leaves the machine and to whom, whether consent and opt-outs work, what happens with no human present, and whether the agent's own record shows what it did and who approved it. This preprint describes the threat model, the 18 scored tests (test definitions v0.2) and four informational rows, the capture rig (transparent TLS interception, a credential-substituting proxy and system-call tracing), the scoring rules with evidence gates, held mode for results under vendor disclosure, public per-agent evidence bundles, the email-only disclosure policy and the counted correction log, including a re-run of every agent on 27 September 2026 (partial for two vendor-hosted agents). It reports preliminary results for 21 agents (378 cells: 200 pass, 82 fail, 54 not tested by us, 42 held) as published on https://agenticbench.org on 27 September 2026. Approval attribution (R2) has no published pass: 19 agents fail and 2 results are held. A companion efficiency series runs one design task through all 21 agents. Status: preliminary. No independent reproduction has been made yet, and the human replay of headline results is not complete. Drafting and review were AI-assisted; the paper was AI and human reviewed. Competing interests: Agentic Thinking has developed AI agent governance and workflow software in the area of approvals and audit records, and its own agent client is built on the Pi runtime, one of the scored agents; the benchmark takes no vendor money. Results under open vendor disclosures are not described. Rig (code and method, Apache-2.0): https://github.com/agentic-thinking/agenticbench, commit 348c7c7 (tag v0.2.1); corrected R3 rig: tag v0.2.3 (commit d6e6d1c).

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.