A forecast-guided reinforcement learning approach for trajectory planning of unmanned aerial base stations
Abstract The rapid growth of wireless devices and the emergence of dynamic traffic hotspots have increased the need for intelligent trajectory planning in unmanned aerial vehicle base stations (UAV-BSs) operating in complex urban environments, where conventional reactive methods relying only on current system states cannot anticipate near-future changes in demand, obstacle risk, and communication conditions, often resulting in inefficient coverage, higher energy consumption, and unsafe flight behavior. To address this limitation, this paper proposes Forecast-SAC, a forecasting-enhanced deep reinforcement learning framework for proactive 3D UAV-BS trajectory planning that integrates a CNN-LSTM forecasting module with a Soft Actor-Critic (SAC) controller, where the forecasting module learns spatiotemporal demand and obstacle-risk patterns from historical map sequences and the SAC agent uses both predicted and current states to generate continuous control actions under joint service, safety, and energy constraints. The framework is evaluated in a grid-based urban simulation environment with mobile Gaussian hotspot demand, spatial obstacle-risk fields, queue-based service accumulation, and an air-to-ground channel model dependent on elevation angle, while training is performed using sequence-based forecasting and environment interaction under collision-avoidance and energy constraints. Experimental results show stable convergence and demonstrate that Forecast-SAC achieves an average served traffic of 4439.26 ± 532.96 Mbit, energy efficiency of 0.0623 Mbit/J under a revised rotary-wing energy model, average risk of 2.85, a baseline fairness index of 0.229 (a limitation addressed by the proposed fairness-enhancement term, which raises the index to 0.318–0.401; Sect. 4.7), and battery retention of 96.23%, while maintaining low-latency inference (3.07 ± 0.21 ms) compatible with real-time control loops; deployment readiness beyond simulation would additionally require hardware-in-the-loop validation. Comparative results indicate that Forecast-SAC produces smoother and safer trajectories than reactive baselines while maintaining competitive throughput. Ablation studies further show that replacing SAC with heuristic forecast-driven rules increases risk by 57–157%, confirming the necessity of learned continuous control, while PPO and DDPG baselines without forecasting are outperformed by 25.7–41.3% in episodic return, validating the contribution of spatiotemporal prediction. Sensitivity analysis over the violation penalty (λv ∈ {100–1000}) identifies 700 as a near-optimal tradeoff point, achieving 85.7% fewer violations than the least-penalized setting with only a small throughput loss. Multi-step forecasting analysis shows that single-step (t + 1) prediction is adopted as the primary operating point for its favourable accuracy–latency tradeoff, while the two-step (t + 1 + t + 3) variant yields a marginal return improvement at increased inference cost, and longer horizons (t + 5) degrade performance due to error accumulation, and a fairness-oriented variant increases the fairness index to 0.318–0.401. The proposed framework demonstrates that unified predictive-control learning enables safe and efficient UAV-BS navigation under dynamic uncertainty, achieving a strong safety–throughput balance that reactive methods cannot match in high-risk environments.
Authors
- Talha Mahboob Alam (ORCID: https://orcid.org/0000-0001-7228-0046)
- Ayesha Aslam (ORCID: https://orcid.org/0000-0001-5886-1848)
- Zhu Wei
- Adil Hussain
- Kamran Shaukat
- Tariq
Institutions
- Norwegian University of Science and Technology (NO)
- Chang'an University (CN)
- Torrens University Australia (AU)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1038/s41598-026-67147-z
- Primary Topic
- UAV Applications and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00