Agentic Execution Assurance: Observation-Bounded Certification, Runtime Enforcement, and Tamper-Evident Evidence

Tool-using AI agents execute multi-step interactions with external systems, making final task success insufficient to characterize whether an execution can be certified, safely controlled, or independently audited. This work develops a bounded execution-assurance framework connecting four layers: specification, observability, runtime enforcement, and evidence integrity. We formalize observation-induced equivalence over execution traces and characterize when an execution property admits sound-and-complete Boolean certification from a given observation boundary. A conservative audit semantics distinguishes Certified, Violated, Unknown, and out-of-model observations. We then specialize established safety-enforcement reasoning to typed committed executions under complete mediation, pre-effect verification, rejection-as-stuttering, and explicit state-consistency assumptions. The paper introduces a claim-dependent assurance profile rather than a universal scalar reliability score and develops the Agentic Reliability Evidence Profile (AREF), which specifies evidence semantics that can be transported through existing telemetry systems such as OpenTelemetry. AREF separates evidentiary completeness, cryptographic integrity, policy binding, effect evidence, resource support, and trust assumptions. AegisRun v1.0.1 is examined as a pinned reference implementation. The source-level case study identifies both implemented assurance mechanisms and important boundaries, including signer trust anchoring, non-atomic resource admission, post-effect accounting, and crash-consistency limitations. Controlled experiments illustrate observation-dependent certifiability over a disclosed finite execution universe, cryptographic corruption detection, the distinction between integrity and semantic completeness, and support-aware gating metrics. The empirical results are explicitly bounded and are not presented as production effectiveness measurements. The contribution is the particular integration of observation-bounded certification, typed pre-effect enforcement, claim-dependent evidence semantics, and a source-audited agent gateway implementation, rather than a claim to originate the individual underlying mechanisms. Reproducibility artifact: https://doi.org/10.5281/zenodo.23118665 Paper DOI: https://doi.org/10.5281/zenodo.23119087 Version: 2.1

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23119086
Primary Topic
Safety Systems Engineering in Autonomy
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Agentic Execution Assurance: Observation-Bounded Certification, Runtime Enforcement, and Tamper-Evident Evidence

Stamatis-Christos Saridakis
Zenodo (CERN European Organization for Nuclear Research)
Safety Systems Engineering in Autonomy
preprint

Agentic Execution Assurance: Observation-Bounded Certification, Runtime Enforcement, and Tamper-Evident Evidence

Stamatis-Christos Saridakis
preprint en

Abstract

Tool-using AI agents execute multi-step interactions with external systems, making final task success insufficient to characterize whether an execution can be certified, safely controlled, or independently audited. This work develops a bounded execution-assurance framework connecting four layers: specification, observability, runtime enforcement, and evidence integrity. We formalize observation-induced equivalence over execution traces and characterize when an execution property admits sound-and-complete Boolean certification from a given observation boundary. A conservative audit semantics distinguishes Certified, Violated, Unknown, and out-of-model observations. We then specialize established safety-enforcement reasoning to typed committed executions under complete mediation, pre-effect verification, rejection-as-stuttering, and explicit state-consistency assumptions. The paper introduces a claim-dependent assurance profile rather than a universal scalar reliability score and develops the Agentic Reliability Evidence Profile (AREF), which specifies evidence semantics that can be transported through existing telemetry systems such as OpenTelemetry. AREF separates evidentiary completeness, cryptographic integrity, policy binding, effect evidence, resource support, and trust assumptions. AegisRun v1.0.1 is examined as a pinned reference implementation. The source-level case study identifies both implemented assurance mechanisms and important boundaries, including signer trust anchoring, non-atomic resource admission, post-effect accounting, and crash-consistency limitations. Controlled experiments illustrate observation-dependent certifiability over a disclosed finite execution universe, cryptographic corruption detection, the distinction between integrity and semantic completeness, and support-aware gating metrics. The empirical results are explicitly bounded and are not presented as production effectiveness measurements. The contribution is the particular integration of observation-bounded certification, typed pre-effect enforcement, claim-dependent evidence semantics, and a source-audited agent gateway implementation, rather than a claim to originate the individual underlying mechanisms. Reproducibility artifact: https://doi.org/10.5281/zenodo.23118665 Paper DOI: https://doi.org/10.5281/zenodo.23119087 Version: 2.1

Zenodo (CERN European Organization for Nuclear Research)
Safety Systems Engineering in Autonomy
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.