Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration
This paper presents an integrated framework for reusable rocket trajectory optimization that combines deep reinforcement learning with FPGA-based inference acceleration. The proposed Adaptive Hardware-Accelerated Twin Delayed Deep Deterministic Policy Gradient (AHA-TD3) framework couples a hierarchical control policy, an online adaptation mechanism, and an FPGA-oriented inference pipeline. Within the simulation and board-level hardware validation considered in this study, AHA-TD3 improves landing success rate, position accuracy, and fuel consumption relative to the compared PPO, DDPG, SAC, TD3, and SCP-MPC baselines. The FPGA implementation achieves up to 12.7× lower inference latency than the CPU software baseline while maintaining low power consumption. These results indicate the potential of combining adaptive reinforcement learning and hardware-software co-design for real-time reusable launch vehicle guidance, while further high-fidelity hardware-in-the-loop and flight-oriented validation remain necessary before operational use.
Authors
- Limin Mao (ORCID: https://orcid.org/0000-0003-4472-2630)
- Lei-Lei Zhang
- Yang-Ming Guo
- Kuo Guo
- Wen-Gang Yao
Institutions
- Northwestern Polytechnical University (CN)
- House of Representatives (NL)
- Beijing Microelectronics Technology Institute (CN)
- Chinese People's Liberation Army (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-04
- DOI
- https://doi.org/10.1038/s41598-026-65000-x
- Primary Topic
- Adaptive Dynamic Programming Control
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Northwestern University
- Northwestern Polytechnical University