Geometric Shielding and Lagrangian Constraints in Safe Reinforcement Learning: A Factorial Study

Safe reinforcement learning requires balancing task performance against safety constraints, yet how different safety mechanisms interact when combined remains an open question. We present a systematic empirical comparison of four reinforcement learning algorithms, PPO, SAC, RCPO (with fixed λ), and Lagrangian PPO (LagPPO), each trained with and without a geometric keepout shield on the SafetyPointPush1-v0 benchmark. Our 4×2 factorial experiment across three random seeds reveals that safety mechanisms are not interchangeable and exhibit distinct trade-offs. RCPO achieves the lowest constraint violation rate throughout training but at significant cost to task return. SAC converges fastest in return but incurs high safety cost without explicit constraints. Because each seed fixes the hazard layout, conditions can be compared on matched geometry. Under that pairing, shielded LagPPO ends with a lower Lagrange multiplier than its unshielded counterpart in all three seed pairs, and unconstrained PPO intervenes more often than LagPPO in all three. Both orderings are consistent in direction but their marginal distributions overlap, so three seeds do not resolve their magnitude. No algorithm achieves zero violations within one million training steps, highlighting the difficulty of full constraint satisfaction under compute constraints. Our results suggest that safety mechanisms should be treated as composable components with distinct trade-offs, and that the appropriate choice depends on whether the priority is constraint satisfaction, task performance, or both. We also report an implementation defect discovered during analysis: the shield's position extraction returns sensor components rather than agent coordinates on this observation space, so the shield-on conditions should be read as an action-perturbation baseline rather than a geometry-aware keepout shield.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22965116
Primary Topic
Reinforcement Learning in Robotics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Geometric Shielding and Lagrangian Constraints in Safe Reinforcement Learning: A Factorial Study

Samuele Pesacane
Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
preprint

Geometric Shielding and Lagrangian Constraints in Safe Reinforcement Learning: A Factorial Study

Samuele Pesacane
preprint en

Abstract

Safe reinforcement learning requires balancing task performance against safety constraints, yet how different safety mechanisms interact when combined remains an open question. We present a systematic empirical comparison of four reinforcement learning algorithms, PPO, SAC, RCPO (with fixed λ), and Lagrangian PPO (LagPPO), each trained with and without a geometric keepout shield on the SafetyPointPush1-v0 benchmark. Our 4×2 factorial experiment across three random seeds reveals that safety mechanisms are not interchangeable and exhibit distinct trade-offs. RCPO achieves the lowest constraint violation rate throughout training but at significant cost to task return. SAC converges fastest in return but incurs high safety cost without explicit constraints. Because each seed fixes the hazard layout, conditions can be compared on matched geometry. Under that pairing, shielded LagPPO ends with a lower Lagrange multiplier than its unshielded counterpart in all three seed pairs, and unconstrained PPO intervenes more often than LagPPO in all three. Both orderings are consistent in direction but their marginal distributions overlap, so three seeds do not resolve their magnitude. No algorithm achieves zero violations within one million training steps, highlighting the difficulty of full constraint satisfaction under compute constraints. Our results suggest that safety mechanisms should be treated as composable components with distinct trade-offs, and that the appropriate choice depends on whether the priority is constraint satisfaction, task performance, or both. We also report an implementation defect discovered during analysis: the shield's position extraction returns sensor components rather than agent coordinates on this observation space, so the shield-on conditions should be read as an action-perturbation baseline rather than a geometry-aware keepout shield.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Geometric Shielding and Lagrangian Constraints in Safe Reinforcement Learning: A Factorial Study — Samuele Pesacane · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS