Auditing Agent Actions with Customer-Retained Evidence: Report Integrity, Outcome Verification, and Economic Design

AI-agent auditing requires distinguishing report integrity from evidence that a requested action occurred. We implement two customer-controlled HTTP components: a gateway that retains tool requests and responses, and an outcome service that binds a registered action to native Git or SQLite records. In the first study, twelve model-agent workflows generated 27 tool calls matching an independent simulated-backend ledger. A comparison of four evidence interfaces produced 176 eligible model-reviewer judgments. Verified gateway evidence identified 26 of 26 planted alterations, compared with 12 of 26 for submitted logs and 24 of 26 for independently retained ordinary logs. No interface falsely rejected an authentic control in eighteen control judgments per interface. In the second study, 26 authored backend/fault scenarios tested lost responses, retries, overwrites, stale versions, rollback and missing causal records. The outcome service distinguished action application from the current postcondition, returning eighteen correct execution decisions, eight unresolved decisions and no incorrect decisions. Retained native logs with the same reconciliation algorithm produced identical verdicts. These studies support a bounded implementation and characterize its failure behavior; they do not establish a general detection advantage for signatures. A prospective economic design separates USDG customer payments from operator capital and shared protocol control. It specifies bounded obligations, attributable violations and comparisons with conventional contracts and stablecoin collateral. Customer demand, investigation savings and native-token benefits remain unmeasured. The replication archive includes both implementations, raw observations, fault protocols and offline verification code. Working paper, version 0.6, revised September 18, 2026. This deposit includes the manuscript PDF and its reproduction archive. Project website: https://naysay.xyz/research.html

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22830017
Primary Topic
Software System Performance and Reliability
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Auditing Agent Actions with Customer-Retained Evidence: Report Integrity, Outcome Verification, and Economic Design

Howie Xu
Zenodo (CERN European Organization for Nuclear Research)
Software System Performance and Reliability
preprint

Auditing Agent Actions with Customer-Retained Evidence: Report Integrity, Outcome Verification, and Economic Design

Howie Xu
preprint en

Abstract

AI-agent auditing requires distinguishing report integrity from evidence that a requested action occurred. We implement two customer-controlled HTTP components: a gateway that retains tool requests and responses, and an outcome service that binds a registered action to native Git or SQLite records. In the first study, twelve model-agent workflows generated 27 tool calls matching an independent simulated-backend ledger. A comparison of four evidence interfaces produced 176 eligible model-reviewer judgments. Verified gateway evidence identified 26 of 26 planted alterations, compared with 12 of 26 for submitted logs and 24 of 26 for independently retained ordinary logs. No interface falsely rejected an authentic control in eighteen control judgments per interface. In the second study, 26 authored backend/fault scenarios tested lost responses, retries, overwrites, stale versions, rollback and missing causal records. The outcome service distinguished action application from the current postcondition, returning eighteen correct execution decisions, eight unresolved decisions and no incorrect decisions. Retained native logs with the same reconciliation algorithm produced identical verdicts. These studies support a bounded implementation and characterize its failure behavior; they do not establish a general detection advantage for signatures. A prospective economic design separates USDG customer payments from operator capital and shared protocol control. It specifies bounded obligations, attributable violations and comparisons with conventional contracts and stablecoin collateral. Customer demand, investigation savings and native-token benefits remain unmeasured. The replication archive includes both implementations, raw observations, fault protocols and offline verification code. Working paper, version 0.6, revised September 18, 2026. This deposit includes the manuscript PDF and its reproduction archive. Project website: https://naysay.xyz/research.html

Zenodo (CERN European Organization for Nuclear Research)
Nabsys (United States) (US)
Peace, Justice and strong institutions
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Auditing Agent Actions with Customer-Retained Evidence: Report Integrity, Outcome Verification, and Economic Design — Howie Xu · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS