Value-Aware Estimation of Diffusion Systems for Continuous-Time Reinforcement Learning: Surrogate Loss and Financial Applications

Abstract. We estimate diffusion systems modeled by SDEs to solve optimal control problems in a continuous-time model-based reinforcement learning framework. Instead of minimizing a statistical loss, we estimate models by reducing the mismatch between the model-based value process and empirical rewards under a given policy. Direct optimization of this loss is generally intractable due to the difficulty of computing the value function of the model. To overcome this challenge, we construct an emulator of the model-based value process, yielding a surrogate loss tractable for optimization even in high dimensions. We establish convergence of the surrogate minimizer to the original loss minimizer and quantify the error from time discretization. Applications to mean-variance portfolio selection and stock portfolio liquidation demonstrate that our approach substantially reduces decision bias and more accurately estimates weak signals from noisy environments than classical statistical methods. We also show that our model-based approach requires far fewer samples to reach comparable performance levels than model-free learning in an empirical study.

Authors

Institutions

Publication Details

Journal
SIAM Journal on Financial Mathematics
Published
2026-09-21
DOI
https://doi.org/10.1137/25m1731885
Primary Topic
Advanced Bandit Algorithms Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Value-Aware Estimation of Diffusion Systems for Continuous-Time Reinforcement Learning: Surrogate Loss and Financial Applications

Boyu Wang, Lingfei Li, Bo Wu
SIAM Journal on Financial Mathematics
Advanced Bandit Algorithms Research
article

Value-Aware Estimation of Diffusion Systems for Continuous-Time Reinforcement Learning: Surrogate Loss and Financial Applications

Boyu Wang, Lingfei Li, Bo Wu
article en

Abstract

Abstract. We estimate diffusion systems modeled by SDEs to solve optimal control problems in a continuous-time model-based reinforcement learning framework. Instead of minimizing a statistical loss, we estimate models by reducing the mismatch between the model-based value process and empirical rewards under a given policy. Direct optimization of this loss is generally intractable due to the difficulty of computing the value function of the model. To overcome this challenge, we construct an emulator of the model-based value process, yielding a surrogate loss tractable for optimization even in high dimensions. We establish convergence of the surrogate minimizer to the original loss minimizer and quantify the error from time discretization. Applications to mean-variance portfolio selection and stock portfolio liquidation demonstrate that our approach substantially reduces decision bias and more accurately estimates weak signals from noisy environments than classical statistical methods. We also show that our model-based approach requires far fewer samples to reach comparable performance levels than model-free learning in an empirical study.

SIAM Journal on Financial MathematicsVol. 17(3)
Chinese University of Hong Kong (HK)
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Advanced Bandit Algorithms Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Value-Aware Estimation of Diffusion Systems for Continuous-Time Reinforcement Learning: Surrogate Loss and Financial Applications — Boyu Wang, Lingfei Li, et al. · SIAM Journal on Financial Mathematics (2026) | TGRS Research Map | TGRS