Transferable reinforcement learning for demand-responsive transit scheduling via offline pre-training and online fine-tuning

Demand-responsive transit (DRT) offers flexibility in adapting to dynamic travel demands and improving resource allocation efficiency. However, existing scheduling models suffer from limited generalization and the cold-start problem when deployed in new operational contexts. This study adapts reinforcement learning to cross-region DRT scheduling by combining offline imitation learning with online proximal policy optimization in a transferable framework (TR-PPO). The proposed approach follows a two-stage learning paradigm. In the first stage, behavior cloning is conducted on extensive expert datasets to learn transferable dispatching patterns and establish an initial policy prior. By using relative spatial representations rather than absolute coordinates, the framework reduces dependence on absolute coordinates and improves cross-region feature transfer. In the second stage, the pre-trained model is fine-tuned online in the target region using PPO to improve policy adaptation to local spatiotemporal dynamics under limited-data conditions. Case study findings from real-world datasets in China demonstrate that TR-PPO outperforms the tested scheduling baselines, achieving total system cost reductions of 11.6%–27.7%. Across the tested within-city, single-source cross-city, and multi-source cross-city transfer settings, the transfer-based policies also achieve competitive performance relative to the other baseline algorithms. TR-PPO accelerates convergence during online fine-tuning, reducing the number of training episodes by 57.8%. Moreover, TR-PPO demonstrates strong data efficiency during online fine-tuning, achieving competitive performance even with only 25% of the target-region training data. The results indicate that the proposed framework can improve the adaptation efficiency of DRT scheduling policies across the tested source-target regions with different demand patterns and network structures.

Authors

Institutions

Publication Details

Journal
Transportation Research Part C Emerging Technologies
Published
2026-09-15
DOI
https://doi.org/10.1016/j.trc.2026.106010
Primary Topic
Transportation and Mobility Innovations
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Transferable reinforcement learning for demand-responsive transit scheduling via offline pre-training and online fine-tuning

Jing Bian, Xiaolei Ma, Haoyang Yan, Hesham El-Sayed et al.
Transportation Research Part C Emerging Technologies
Transportation and Mobility Innovations
article

Transferable reinforcement learning for demand-responsive transit scheduling via offline pre-training and online fine-tuning

Jing Bian, Xiaolei Ma, Haoyang Yan, Hesham El-Sayed, Xiaohan Liu, Kun Gao
article en

Abstract

Demand-responsive transit (DRT) offers flexibility in adapting to dynamic travel demands and improving resource allocation efficiency. However, existing scheduling models suffer from limited generalization and the cold-start problem when deployed in new operational contexts. This study adapts reinforcement learning to cross-region DRT scheduling by combining offline imitation learning with online proximal policy optimization in a transferable framework (TR-PPO). The proposed approach follows a two-stage learning paradigm. In the first stage, behavior cloning is conducted on extensive expert datasets to learn transferable dispatching patterns and establish an initial policy prior. By using relative spatial representations rather than absolute coordinates, the framework reduces dependence on absolute coordinates and improves cross-region feature transfer. In the second stage, the pre-trained model is fine-tuned online in the target region using PPO to improve policy adaptation to local spatiotemporal dynamics under limited-data conditions. Case study findings from real-world datasets in China demonstrate that TR-PPO outperforms the tested scheduling baselines, achieving total system cost reductions of 11.6%–27.7%. Across the tested within-city, single-source cross-city, and multi-source cross-city transfer settings, the transfer-based policies also achieve competitive performance relative to the other baseline algorithms. TR-PPO accelerates convergence during online fine-tuning, reducing the number of training episodes by 57.8%. Moreover, TR-PPO demonstrates strong data efficiency during online fine-tuning, achieving competitive performance even with only 25% of the target-region training data. The results indicate that the proposed framework can improve the adaptation efficiency of DRT scheduling policies across the tested source-target regions with different demand patterns and network structures.

Transportation Research Part C Emerging TechnologiesVol. 194
United Arab Emirates University (AE), Chalmers University of Technology (SE), Beihang University (CN)
Decent work and economic growth
Openalex Percentile: Top 18%
Transportation and Mobility Innovations
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.