Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management
Learning from Snapshots Is Not Enough: An Even-Driven Continuous-Time Reinforcement Learning Framework Revenue management systems evolve continuously, but reinforcement learning often requires dividing time into a fixed grid. Fine grids improve accuracy but increase computation; coarse grids can sacrifice performance. In “Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management,” Meng, Chen, and Gao address this tension with an event-driven continuous-time reinforcement learning framework. The key insight is that state jumps occur at customer arrival times and naturally partition each sample path, eliminating the need for up-front time discretization. Building on this structure, the authors extend policy evaluation to continuous time and develop event-driven actor-critic algorithms. A comprehensive numerical study shows their strong performance. In a bursty arrival environment, the proposed continuous-time approach achieves up to 16.64% higher revenue than a coarse-grid benchmark with comparable training time. The approach also handles a large-scale network revenue management problem with 100 resources and 200 products, and an extension to queue admission control further demonstrates the broad applicability of the framework.
Authors
- Xuefeng Gao (ORCID: https://orcid.org/0000-0003-2424-8257)
- Ningyuan Chen (ORCID: https://orcid.org/0000-0002-3948-1011)
- Huiling Meng (ORCID: https://orcid.org/0009-0002-2180-0974)
Institutions
- Chinese University of Hong Kong (HK)
- University of Toronto (CA)
Publication Details
- Journal
- Operations Research
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1287/opre.2024.1190
- Primary Topic
- Supply Chain and Inventory Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00