PR-Nav: Potential-Shaped Reinforcement Learning with Physics-Modulated Manifolds for Agile Autonomous Navigation
High-speed autonomous navigation of micro aerial vehicles (MAVs) is essential for inspection tasks in restricted spaces (e.g., dense industrial facilities or narrow corridors), where reliable spatial perception and stable flight control are required. However, existing deep reinforcement learning-based navigation methods often suffer from delayed obstacle avoidance responses due to the low fitting efficiency of raw geometric distance representations for dynamic threats, trigger control oscillations that disrupt flight stability due to an over-reliance on external hard-clipping mechanisms, and frequently fall into local optima such as obstacle-edge hovering driven by heuristic penalties, thereby reducing overall navigation success rates and training efficiency. We propose PR-Nav, an agile navigation framework combining potential-field-augmented perceptual representation with a physics-modulated probabilistic action manifold. PR-Nav comprises three components: (1) a potential-field-augmented ray tensor representation that nonlinearly maps raw distance information into physical repulsive gradients to improve the policy network’s fitting efficiency for dynamic threat boundaries; (2) a physics-prior-modulated Beta action manifold that injects local potential-field gradients as residual biases into the probability density generation process, replacing external hard clipping with internal network constraints to reduce the action smoothness index (ASI); and (3) a potential-based reward shaping (PBRS) mechanism that replaces traditional heuristic penalties with global physical potential energy differences to improve the navigation success rate in complex environments and accelerate policy convergence. Experiments on the high-fidelity Isaac Sim simulator and a real-world physical flight platform show that PR-Nav achieves an overall navigation success rate of 96.5%, outperforming the strongest competing method by 10.8 percentage points. In quantitative evaluations of flight smoothness and training efficiency, its action smoothness index (ASI) is reduced to 4.25 m/s3 and it achieves an approximate 54% reduction in convergence steps compared to the baseline, the best results among the compared methods. Ablation studies verify the contribution of each component to its corresponding metrics. These results demonstrate that combining physical potential field modeling with underlying probability distribution reconstruction is an effective route to robust and agile navigation in dynamic restricted spaces.
Authors
- Fengren Jing
- Heng Zhong (ORCID: https://orcid.org/0009-0006-3906-0651)
- 遇常娥
- Tianlong Wang
- Yong Chang
- Ge Qin
Institutions
- Shenyang Institute of Automation (CN)
- Chinese Academy of Sciences (CN)
- State Key Laboratory of Robotics
- China Yangtze Power Co., Ltd. (China) (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/electronics15194454
- Primary Topic
- Robotic Path Planning Algorithms
- Type
- article
- Field-Weighted Citation Impact
- 0.00