Exact variance of random return in distributional LQR and its application to mean–variance optimal control
The classical linear quadratic regulator (LQR) minimizes the expected cumulative return but fails to account for performance variability, rendering it inadequate for risk-aware applications. To address this, we introduce the variance of the cumulative return as a risk measure in LQR. We derive the first exact closed-form expression for the variance of the discounted in?finite horizon return within the discrete-time Distributional LQR framework, for i.i.d. disturbances with symmetric probability densities. For Gaussian disturbances, this expression elegantly simplifies to a form dependent only on the disturbance covariance. Leveraging these theoretical foundations, we formulate a mean-variance optimal control problem that explicitly manages the trade-off between expected return and performance variability. To address the resulting non-convex optimization problem, we propose a novel adjoint gradient descent algorithm for a penalized formulation of the original problem, and establish that all iterates remain stabilizing and converge to a stationary point of the penalized objective. The effectiveness of this framework and the inherent risk-performance trade-off? are demonstrated through numerical experiments.
Authors
- Ruyi Teng
- Yulong Gao
- Zifan Wang
Publication Details
- Journal
- Automatica
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1016/j.automatica.2026.113327
- Primary Topic
- Adaptive Dynamic Programming Control
- Type
- article
- Field-Weighted Citation Impact
- 0.00