Explainable Stability Certification of Reinforcement Learning Control Policies Using Sparse Dynamics Identification
Reinforcement learning (RL) is a promising alternative to classical guidance and control methods; however, the black-box nature of deep neural network policies and the lack of interpretable stability evidence remain barriers to real-world aerospace adoption. This paper presents an a posteriori methodology using Sparse Identification of Nonlinear Dynamics (SINDy) to recover sparse analytical representations of the closed-loop dynamics induced by a trained RL controller, with the goal of certifying its stability. When the uncontrolled dynamics and input map are known, the identified model provides an explicit analytical approximation of the state-feedback control law, offering functional explainability through interpretable state couplings and nonlinear terms, as well as a lightweight surrogate for real-time deployment. In parallel, an analytical approximation of the positive cost to go is obtained as a candidate Lyapunov function. The Lyapunov conditions are evaluated for both the original RL policy and the reconstructed analytical controller over a bounded operating domain, while a Probably Approximately Correct (PAC) bound quantifies confidence in finite-sample verification. Demonstrations on a spring-mass oscillator, spacecraft attitude control, and asteroid hovering show that sparse analytical laws can reproduce and explain trained RL controllers, while PAC-supported Lyapunov analysis provides a principled framework for their systematic stability certification.
Authors
- Andrea D’Ambrosio (ORCID: https://orcid.org/0000-0002-8084-4101)
- Andrea Scorsoglio (ORCID: https://orcid.org/0000-0001-5875-3804)
- Roberto Furfaro (ORCID: https://orcid.org/0000-0001-6076-8992)
Institutions
- University of Arizona (US)
- University of South Florida (US)
Publication Details
- Journal
- Journal of Guidance Control and Dynamics
- Published
- 2026-10-07
- DOI
- https://doi.org/10.2514/1.g009769
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00