Offline and offline-to-online reinforcement learning for bus holding control
Abstract Bus bunching is a self-reinforcing phenomenon in which a delayed bus boards more passengers, incurs further delay, and is caught by its follower. Holding a bus at a stop to regulate headways is a widely studied corrective measure, yet most evaluations rely on a single shaped objective rather than independent passenger and service outcomes. This study asks whether a fixed operational log can support network-wide holding policies without simulator interaction, and whether offline-to-online fine-tuning yields sufficient improvement to justify its computational cost. We constructed a microscopic Simulation of Urban Mobility (SUMO) benchmark with 12 Changsha bus lines and 3.37 million logged holding-only transitions. We compared five transportation rules and three offline learners: behavior cloning (BC), conservative Q-learning (CQL), and robust ensemble Soft Actor-Critic (RE-SAC). We also compared online Soft Actor-Critic (SAC) with warm-start reinforcement learning (WSRL) and reinforcement learning with prior data (RLPD), two offline-to-online methods, under separate interaction budgets. Every reported checkpoint was evaluated across ten common 18 000-s SUMO episodes that ran to natural completion. WSRL achieved the highest selected-checkpoint return (−710 544 ± 17 162). The mean passenger waiting time was 301 ± 4 s and the mean total travel time was 1 266 ± 18 s. However, no single controller was best on every passenger and service outcome, underscoring the need for evaluation beyond shaped return. The study contributes a reproducible benchmark and a comparison design with explicitly separated interaction budgets. The findings are limited to the simulated setting, as observed operating trajectories were unavailable for external calibration.
Authors
- Yifan Zhang (ORCID: https://orcid.org/0009-0009-8132-7075)
- Qifan Zhang
- Liang Zheng
- Haoming Li
Institutions
- Monash University Malaysia (MY)
- Central South University (CN)
- Shanghai Ocean University (CN)
Publication Details
- Journal
- Transportation Safety and Environment
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1093/tse/tdag059
- Primary Topic
- Transportation Planning and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00