Hierarchical Temporal Runtime Assurance for Controlled Agentic AI Systems: Safety Shielding, Auditable Action Repair, and Bounded Recovery

Agentic artificial intelligence requires safeguards that remain effective across trajectories rather than only at individual decisions. This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems. A compact, auditable fragment of Signal Temporal Logic (STL) provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements. The final architecture uses invariant hard-safety formulas to define feasibility, while recovery/liveness is monitored separately and may trigger escalation. Across a 5750-episode main campaign, MARIS-TRA achieved 100.00% realized hard-safety window satisfaction, 100.00% hard-safety episode satisfaction (450/450; Wilson 95% CI 99.15–100.00%), zero collision episodes, and 2.55 ms mean latency in the core comparison. Under the primary strict numerical semantics, an independent CBF-QP baseline achieved 88.89% hard-safety episode satisfaction; all 50 strict failures were very small arena-boundary overshoots in the boundary-stress scenario, with no collision or separation failures, and the post hoc tolerance sensitivity reached 100% at epsilon = 10−4 normalized simulator units. An additional 8640-episode targeted validation examined recovery hysteresis, model mismatch, AHO control flow, and safety–recovery conflicts. Under confirmatory high-density testing, safety-only shielding preserved hard safety in 810/810 episodes, whereas joint-hard enforcement produced 28/810 separation-safety failures (3.46%; Wilson 95% CI 2.40–4.95%) without collisions. Actuation-noise and delay experiments further show that the formal result is a conditional predicted-trace certification rather than a disturbance-robust guarantee on future receding-horizon execution. The empirical claims are, therefore, limited to the evaluated continuous-action multi-agent setting, while the architecture remains policy-separable and auditable.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-09-21
DOI
https://doi.org/10.3390/make8090293
Primary Topic
Formal Methods in Verification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hierarchical Temporal Runtime Assurance for Controlled Agentic AI Systems: Safety Shielding, Auditable Action Repair, and Bounded Recovery

Tymoteusz I. Miller
Machine Learning and Knowledge Extraction
Formal Methods in Verification
article

Hierarchical Temporal Runtime Assurance for Controlled Agentic AI Systems: Safety Shielding, Auditable Action Repair, and Bounded Recovery

Tymoteusz I. Miller
article en

Abstract

Agentic artificial intelligence requires safeguards that remain effective across trajectories rather than only at individual decisions. This study introduces MARIS-TRA, a hierarchical temporal runtime assurance extension of Controlled Agentic AI Systems. A compact, auditable fragment of Signal Temporal Logic (STL) provides quantitative robustness for speed, arena, pairwise separation, restricted-zone, and bounded-recovery requirements. The final architecture uses invariant hard-safety formulas to define feasibility, while recovery/liveness is monitored separately and may trigger escalation. Across a 5750-episode main campaign, MARIS-TRA achieved 100.00% realized hard-safety window satisfaction, 100.00% hard-safety episode satisfaction (450/450; Wilson 95% CI 99.15–100.00%), zero collision episodes, and 2.55 ms mean latency in the core comparison. Under the primary strict numerical semantics, an independent CBF-QP baseline achieved 88.89% hard-safety episode satisfaction; all 50 strict failures were very small arena-boundary overshoots in the boundary-stress scenario, with no collision or separation failures, and the post hoc tolerance sensitivity reached 100% at epsilon = 10−4 normalized simulator units. An additional 8640-episode targeted validation examined recovery hysteresis, model mismatch, AHO control flow, and safety–recovery conflicts. Under confirmatory high-density testing, safety-only shielding preserved hard safety in 810/810 episodes, whereas joint-hard enforcement produced 28/810 separation-safety failures (3.46%; Wilson 95% CI 2.40–4.95%) without collisions. Actuation-noise and delay experiments further show that the formal result is a conditional predicted-trace certification rather than a disturbance-robust guarantee on future receding-horizon execution. The empirical claims are, therefore, limited to the evaluated continuous-action multi-agent setting, while the architecture remains policy-separable and auditable.

Machine Learning and Knowledge ExtractionVol. 8(9)
University of Szczecin (PL), INTI International University (MY)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Formal Methods in Verification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.