Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration

This paper presents an integrated framework for reusable rocket trajectory optimization that combines deep reinforcement learning with FPGA-based inference acceleration. The proposed Adaptive Hardware-Accelerated Twin Delayed Deep Deterministic Policy Gradient (AHA-TD3) framework couples a hierarchical control policy, an online adaptation mechanism, and an FPGA-oriented inference pipeline. Within the simulation and board-level hardware validation considered in this study, AHA-TD3 improves landing success rate, position accuracy, and fuel consumption relative to the compared PPO, DDPG, SAC, TD3, and SCP-MPC baselines. The FPGA implementation achieves up to 12.7× lower inference latency than the CPU software baseline while maintaining low power consumption. These results indicate the potential of combining adaptive reinforcement learning and hardware-software co-design for real-time reusable launch vehicle guidance, while further high-fidelity hardware-in-the-loop and flight-oriented validation remain necessary before operational use.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-04
DOI
https://doi.org/10.1038/s41598-026-65000-x
Primary Topic
Adaptive Dynamic Programming Control
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration

Limin Mao, Lei-Lei Zhang, Yang-Ming Guo, Kuo Guo et al.
Scientific Reports
Adaptive Dynamic Programming Control
article

Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration

Limin Mao, Lei-Lei Zhang, Yang-Ming Guo, Kuo Guo, Wen-Gang Yao
article en

Abstract

This paper presents an integrated framework for reusable rocket trajectory optimization that combines deep reinforcement learning with FPGA-based inference acceleration. The proposed Adaptive Hardware-Accelerated Twin Delayed Deep Deterministic Policy Gradient (AHA-TD3) framework couples a hierarchical control policy, an online adaptation mechanism, and an FPGA-oriented inference pipeline. Within the simulation and board-level hardware validation considered in this study, AHA-TD3 improves landing success rate, position accuracy, and fuel consumption relative to the compared PPO, DDPG, SAC, TD3, and SCP-MPC baselines. The FPGA implementation achieves up to 12.7× lower inference latency than the CPU software baseline while maintaining low power consumption. These results indicate the potential of combining adaptive reinforcement learning and hardware-software co-design for real-time reusable launch vehicle guidance, while further high-fidelity hardware-in-the-loop and flight-oriented validation remain necessary before operational use.

Scientific Reports
Northwestern Polytechnical University (CN), House of Representatives (NL), Beijing Microelectronics Technology Institute (CN), Chinese People's Liberation Army (CN)
Northwestern University, Northwestern Polytechnical University
Affordable and clean energy
Openalex Percentile: Top 9%
Adaptive Dynamic Programming Control
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adaptive rocket trajectory optimization via reinforcement learning and FPGA acceleration — Limin Mao, Lei-Lei Zhang, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS