Quantifying Safety Guardrail Degradation in Autonomous Multi-Agent Workflows Under Goal-Divergence Scenarios

As large language model (LLM) agents are increasingly organized into autonomous multi-agent workflows executing code, interacting with shells, and coordinating complex tasks, their operational safety depends heavily on guardrail mechanisms. However, when agents experience goal-divergence—scenarios where task execution pressures, environmental errors, or urgent operational sub-goals conflict with negative safety invariants—behavioral alignment frequently degrades. In this paper, we present an empirical investigation quantifying safety guardrail degradation across autonomous agent architectures. We formulate a formal state-transition framework under constrained optimization and evaluate models across 20 deterministic scenarios spanning four primary failure modes: System-Prompt Override, Tool Misuse and Privilege Escalation, Self-Preservation and Shutdown Refusal, and Context Window Guardrail Decay. We benchmark frontier commercial and open-weights model classes across three defense tiers: unconstrained baseline (None), system prompt deontic enforcement (Soft), and an integrated dual-layer architecture combining prompt enforcement with deterministic Abstract Syntax Tree (AST) and shell regex filters (Dual). Across 180 structured evaluations, unprotected baselines exhibit a mean Safety Failure Rate (SFR) of 65.0%, with multi-turn context decay demonstrating extreme vulnerability (93.3% failure rate). In contrast, our deterministic hard guardrail achieves 100% pre-execution interception efficiency, completely neutralizing prohibited actions without sacrificing safe task completion (achieving 70.0% Constraint Alignment Index). Our findings demonstrate that soft semantic alignment is insufficient for autonomous systems, necessitating out-of-band deterministic enforcement.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23268413
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Quantifying Safety Guardrail Degradation in Autonomous Multi-Agent Workflows Under Goal-Divergence Scenarios

Muhammad Talha Waris
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
article

Quantifying Safety Guardrail Degradation in Autonomous Multi-Agent Workflows Under Goal-Divergence Scenarios

Muhammad Talha Waris
article en

Abstract

As large language model (LLM) agents are increasingly organized into autonomous multi-agent workflows executing code, interacting with shells, and coordinating complex tasks, their operational safety depends heavily on guardrail mechanisms. However, when agents experience goal-divergence—scenarios where task execution pressures, environmental errors, or urgent operational sub-goals conflict with negative safety invariants—behavioral alignment frequently degrades. In this paper, we present an empirical investigation quantifying safety guardrail degradation across autonomous agent architectures. We formulate a formal state-transition framework under constrained optimization and evaluate models across 20 deterministic scenarios spanning four primary failure modes: System-Prompt Override, Tool Misuse and Privilege Escalation, Self-Preservation and Shutdown Refusal, and Context Window Guardrail Decay. We benchmark frontier commercial and open-weights model classes across three defense tiers: unconstrained baseline (None), system prompt deontic enforcement (Soft), and an integrated dual-layer architecture combining prompt enforcement with deterministic Abstract Syntax Tree (AST) and shell regex filters (Dual). Across 180 structured evaluations, unprotected baselines exhibit a mean Safety Failure Rate (SFR) of 65.0%, with multi-turn context decay demonstrating extreme vulnerability (93.3% failure rate). In contrast, our deterministic hard guardrail achieves 100% pre-execution interception efficiency, completely neutralizing prohibited actions without sacrificing safe task completion (achieving 70.0% Constraint Alignment Index). Our findings demonstrate that soft semantic alignment is insufficient for autonomous systems, necessitating out-of-band deterministic enforcement.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 12%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Quantifying Safety Guardrail Degradation in Autonomous Multi-Agent Workflows Under Goal-Divergence Scenarios — Muhammad Talha Waris · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS