CLI-PPO: A Memory-Augmented PPO Framework with Intrinsic Curiosity and Curriculum Learning for Mapless Multi-Obstacle Navigation
Mapless autonomous navigation in unknown multi-obstacle environments remains challenging for deep reinforcement learning due to limited perception, sparse rewards, and distribution shifts between training and deployment scenarios. Existing methods mainly rely on instantaneous observations or individual enhancement strategies, which often struggle to balance exploration robustness and navigation efficiency. This paper proposes a Curriculum Learning and Intrinsic Curiosity-enhanced Temporal Proximal Policy Optimization framework, referred to as CLI-PPO, for end-to-end autonomous navigation. The framework integrates temporal state representation, intrinsic exploration, and progressive task scheduling into PPO, where an LSTM network improves sequential decision-making under partial observability, an Intrinsic Curiosity Module enhances exploration in unfamiliar regions, and Curriculum Learning stabilizes policy optimization by gradually increasing task difficulty. A hybrid reward function combining target guidance, obstacle avoidance, and exploration incentives is designed for policy training. Experiments were conducted using ablation configurations, three independent training seeds, and an evaluation of one SAC policy under the corresponding task protocol, with 5000 randomized episodes considered at each obstacle-density level. CLI-PPO achieved a success rate of 66.24%, a collision rate of 33.58%, and an average return of 3433.14, compared with 46.06%, 53.94%, and 2083.21 for PPO, respectively. In the 10-obstacle setting, the three training seeds produced a CLI-PPO success rate of 62.72±3.59% and an average return of 3337.12±48.93, while PPO+ICM+Curriculum achieved 63.83±6.61% success and a slightly higher Euclidean-reference SPL. As obstacle density increased, PPO+ICM+Curriculum maintained higher success and SPL, whereas CLI-PPO yielded higher returns and more consistent performance across training runs.
Authors
- Shigang Wang (ORCID: https://orcid.org/0000-0002-6868-6542)
- Qifan KANG
- Debing Xie (ORCID: https://orcid.org/0009-0004-5182-4656)
Institutions
- Guangxi University of Science and Technology (CN)
Publication Details
- Journal
- Sensors
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/s26196098
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00