Agentic Execution Assurance: Observation-Bounded Certification, Runtime Enforcement, and Tamper-Evident Evidence
Tool-using AI agents execute multi-step interactions with external systems, making final task success insufficient to characterize whether an execution can be certified, safely controlled, or independently audited. This work develops a bounded execution-assurance framework connecting four layers: specification, observability, runtime enforcement, and evidence integrity. We formalize observation-induced equivalence over execution traces and characterize when an execution property admits sound-and-complete Boolean certification from a given observation boundary. A conservative audit semantics distinguishes Certified, Violated, Unknown, and out-of-model observations. We then specialize established safety-enforcement reasoning to typed committed executions under complete mediation, pre-effect verification, rejection-as-stuttering, and explicit state-consistency assumptions. The paper introduces a claim-dependent assurance profile rather than a universal scalar reliability score and develops the Agentic Reliability Evidence Profile (AREF), which specifies evidence semantics that can be transported through existing telemetry systems such as OpenTelemetry. AREF separates evidentiary completeness, cryptographic integrity, policy binding, effect evidence, resource support, and trust assumptions. AegisRun v1.0.1 is examined as a pinned reference implementation. The source-level case study identifies both implemented assurance mechanisms and important boundaries, including signer trust anchoring, non-atomic resource admission, post-effect accounting, and crash-consistency limitations. Controlled experiments illustrate observation-dependent certifiability over a disclosed finite execution universe, cryptographic corruption detection, the distinction between integrity and semantic completeness, and support-aware gating metrics. The empirical results are explicitly bounded and are not presented as production effectiveness measurements. The contribution is the particular integration of observation-bounded certification, typed pre-effect enforcement, claim-dependent evidence semantics, and a source-audited agent gateway implementation, rather than a claim to originate the individual underlying mechanisms. Reproducibility artifact: https://doi.org/10.5281/zenodo.23118665 Paper DOI: https://doi.org/10.5281/zenodo.23119087 Version: 2.1
Authors
- Stamatis-Christos Saridakis (ORCID: https://orcid.org/0009-0002-1699-2043)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23119086
- Primary Topic
- Safety Systems Engineering in Autonomy
- Type
- preprint