Geometric Shielding and Lagrangian Constraints in Safe Reinforcement Learning: A Factorial Study
Safe reinforcement learning requires balancing task performance against safety constraints, yet how different safety mechanisms interact when combined remains an open question. We present a systematic empirical comparison of four reinforcement learning algorithms, PPO, SAC, RCPO (with fixed λ), and Lagrangian PPO (LagPPO), each trained with and without a geometric keepout shield on the SafetyPointPush1-v0 benchmark. Our 4×2 factorial experiment across three random seeds reveals that safety mechanisms are not interchangeable and exhibit distinct trade-offs. RCPO achieves the lowest constraint violation rate throughout training but at significant cost to task return. SAC converges fastest in return but incurs high safety cost without explicit constraints. Because each seed fixes the hazard layout, conditions can be compared on matched geometry. Under that pairing, shielded LagPPO ends with a lower Lagrange multiplier than its unshielded counterpart in all three seed pairs, and unconstrained PPO intervenes more often than LagPPO in all three. Both orderings are consistent in direction but their marginal distributions overlap, so three seeds do not resolve their magnitude. No algorithm achieves zero violations within one million training steps, highlighting the difficulty of full constraint satisfaction under compute constraints. Our results suggest that safety mechanisms should be treated as composable components with distinct trade-offs, and that the appropriate choice depends on whether the priority is constraint satisfaction, task performance, or both. We also report an implementation defect discovered during analysis: the shield's position extraction returns sensor components rather than agent coordinates on this observation space, so the shield-on conditions should be read as an action-perturbation baseline rather than a geometry-aware keepout shield.
Authors
- Samuele Pesacane (ORCID: https://orcid.org/0009-0009-1570-6712)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22965116
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- preprint