Value-Aware Estimation of Diffusion Systems for Continuous-Time Reinforcement Learning: Surrogate Loss and Financial Applications
Abstract. We estimate diffusion systems modeled by SDEs to solve optimal control problems in a continuous-time model-based reinforcement learning framework. Instead of minimizing a statistical loss, we estimate models by reducing the mismatch between the model-based value process and empirical rewards under a given policy. Direct optimization of this loss is generally intractable due to the difficulty of computing the value function of the model. To overcome this challenge, we construct an emulator of the model-based value process, yielding a surrogate loss tractable for optimization even in high dimensions. We establish convergence of the surrogate minimizer to the original loss minimizer and quantify the error from time discretization. Applications to mean-variance portfolio selection and stock portfolio liquidation demonstrate that our approach substantially reduces decision bias and more accurately estimates weak signals from noisy environments than classical statistical methods. We also show that our model-based approach requires far fewer samples to reach comparable performance levels than model-free learning in an empirical study.
Authors
- Boyu Wang (ORCID: https://orcid.org/0000-0002-6108-3589)
- Lingfei Li (ORCID: https://orcid.org/0000-0002-1761-3458)
- Bo Wu
Institutions
- Chinese University of Hong Kong (HK)
Publication Details
- Journal
- SIAM Journal on Financial Mathematics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1137/25m1731885
- Primary Topic
- Advanced Bandit Algorithms Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00