A forecast-guided reinforcement learning approach for trajectory planning of unmanned aerial base stations

Abstract The rapid growth of wireless devices and the emergence of dynamic traffic hotspots have increased the need for intelligent trajectory planning in unmanned aerial vehicle base stations (UAV-BSs) operating in complex urban environments, where conventional reactive methods relying only on current system states cannot anticipate near-future changes in demand, obstacle risk, and communication conditions, often resulting in inefficient coverage, higher energy consumption, and unsafe flight behavior. To address this limitation, this paper proposes Forecast-SAC, a forecasting-enhanced deep reinforcement learning framework for proactive 3D UAV-BS trajectory planning that integrates a CNN-LSTM forecasting module with a Soft Actor-Critic (SAC) controller, where the forecasting module learns spatiotemporal demand and obstacle-risk patterns from historical map sequences and the SAC agent uses both predicted and current states to generate continuous control actions under joint service, safety, and energy constraints. The framework is evaluated in a grid-based urban simulation environment with mobile Gaussian hotspot demand, spatial obstacle-risk fields, queue-based service accumulation, and an air-to-ground channel model dependent on elevation angle, while training is performed using sequence-based forecasting and environment interaction under collision-avoidance and energy constraints. Experimental results show stable convergence and demonstrate that Forecast-SAC achieves an average served traffic of 4439.26 ± 532.96 Mbit, energy efficiency of 0.0623 Mbit/J under a revised rotary-wing energy model, average risk of 2.85, a baseline fairness index of 0.229 (a limitation addressed by the proposed fairness-enhancement term, which raises the index to 0.318–0.401; Sect. 4.7), and battery retention of 96.23%, while maintaining low-latency inference (3.07 ± 0.21 ms) compatible with real-time control loops; deployment readiness beyond simulation would additionally require hardware-in-the-loop validation. Comparative results indicate that Forecast-SAC produces smoother and safer trajectories than reactive baselines while maintaining competitive throughput. Ablation studies further show that replacing SAC with heuristic forecast-driven rules increases risk by 57–157%, confirming the necessity of learned continuous control, while PPO and DDPG baselines without forecasting are outperformed by 25.7–41.3% in episodic return, validating the contribution of spatiotemporal prediction. Sensitivity analysis over the violation penalty (λv ∈ {100–1000}) identifies 700 as a near-optimal tradeoff point, achieving 85.7% fewer violations than the least-penalized setting with only a small throughput loss. Multi-step forecasting analysis shows that single-step (t + 1) prediction is adopted as the primary operating point for its favourable accuracy–latency tradeoff, while the two-step (t + 1 + t + 3) variant yields a marginal return improvement at increased inference cost, and longer horizons (t + 5) degrade performance due to error accumulation, and a fairness-oriented variant increases the fairness index to 0.318–0.401. The proposed framework demonstrates that unified predictive-control learning enables safe and efficient UAV-BS navigation under dynamic uncertainty, achieving a strong safety–throughput balance that reactive methods cannot match in high-risk environments.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-28
DOI
https://doi.org/10.1038/s41598-026-67147-z
Primary Topic
UAV Applications and Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A forecast-guided reinforcement learning approach for trajectory planning of unmanned aerial base stations

Talha Mahboob Alam, Ayesha Aslam, Zhu Wei, Adil Hussain et al.
Scientific Reports
UAV Applications and Optimization
article

A forecast-guided reinforcement learning approach for trajectory planning of unmanned aerial base stations

Talha Mahboob Alam, Ayesha Aslam, Zhu Wei, Adil Hussain, Kamran Shaukat, Tariq
article en

Abstract

Abstract The rapid growth of wireless devices and the emergence of dynamic traffic hotspots have increased the need for intelligent trajectory planning in unmanned aerial vehicle base stations (UAV-BSs) operating in complex urban environments, where conventional reactive methods relying only on current system states cannot anticipate near-future changes in demand, obstacle risk, and communication conditions, often resulting in inefficient coverage, higher energy consumption, and unsafe flight behavior. To address this limitation, this paper proposes Forecast-SAC, a forecasting-enhanced deep reinforcement learning framework for proactive 3D UAV-BS trajectory planning that integrates a CNN-LSTM forecasting module with a Soft Actor-Critic (SAC) controller, where the forecasting module learns spatiotemporal demand and obstacle-risk patterns from historical map sequences and the SAC agent uses both predicted and current states to generate continuous control actions under joint service, safety, and energy constraints. The framework is evaluated in a grid-based urban simulation environment with mobile Gaussian hotspot demand, spatial obstacle-risk fields, queue-based service accumulation, and an air-to-ground channel model dependent on elevation angle, while training is performed using sequence-based forecasting and environment interaction under collision-avoidance and energy constraints. Experimental results show stable convergence and demonstrate that Forecast-SAC achieves an average served traffic of 4439.26 ± 532.96 Mbit, energy efficiency of 0.0623 Mbit/J under a revised rotary-wing energy model, average risk of 2.85, a baseline fairness index of 0.229 (a limitation addressed by the proposed fairness-enhancement term, which raises the index to 0.318–0.401; Sect. 4.7), and battery retention of 96.23%, while maintaining low-latency inference (3.07 ± 0.21 ms) compatible with real-time control loops; deployment readiness beyond simulation would additionally require hardware-in-the-loop validation. Comparative results indicate that Forecast-SAC produces smoother and safer trajectories than reactive baselines while maintaining competitive throughput. Ablation studies further show that replacing SAC with heuristic forecast-driven rules increases risk by 57–157%, confirming the necessity of learned continuous control, while PPO and DDPG baselines without forecasting are outperformed by 25.7–41.3% in episodic return, validating the contribution of spatiotemporal prediction. Sensitivity analysis over the violation penalty (λv ∈ {100–1000}) identifies 700 as a near-optimal tradeoff point, achieving 85.7% fewer violations than the least-penalized setting with only a small throughput loss. Multi-step forecasting analysis shows that single-step (t + 1) prediction is adopted as the primary operating point for its favourable accuracy–latency tradeoff, while the two-step (t + 1 + t + 3) variant yields a marginal return improvement at increased inference cost, and longer horizons (t + 5) degrade performance due to error accumulation, and a fairness-oriented variant increases the fairness index to 0.318–0.401. The proposed framework demonstrates that unified predictive-control learning enables safe and efficient UAV-BS navigation under dynamic uncertainty, achieving a strong safety–throughput balance that reactive methods cannot match in high-risk environments.

Scientific ReportsVol. 16(1)
Norwegian University of Science and Technology (NO), Chang'an University (CN), Torrens University Australia (AU)
Affordable and clean energy
Openalex Percentile: Top 8%
UAV Applications and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.