A risk-aware multimodal temporal reinforcement learning approach for collision avoidance of unmanned surface vehicles in multi-dynamic-obstacle scenarios
This study addresses autonomous collision avoidance for unmanned surface vehicles in multi-dynamic-obstacle maritime scenarios, where the number of obstacles varies, risk interactions are complex, and dynamic environmental evolution is difficult to exploit effectively. A Risk-Aware Multimodal Temporal Proximal Policy Optimization algorithm is proposed. First, a collision risk index is defined based on the relative motion between the own ship and the obstacle vessels. Through spatial diffusion and risk fusion, obstacle sets with variable cardinality are transformed into a structured high-dimensional risk representation independent of obstacle population size. Second, a Transformer-based risk-aware multimodal temporal policy network is designed to jointly encode historical risk maps and vehicle states, enabling the extraction of spatiotemporal features and improving the modeling of risk evolution in complex environments. The framework is then integrated with proximal policy optimization and enhanced by curriculum learning to realize progressive training from simple to complex tasks. Simulation results show that the proposed method achieves favorable convergence, adaptability, and robustness, and outperforms baseline methods in success rate, average return, turning responses, and safety.
Authors
- Jiaye Gong (ORCID: https://orcid.org/0000-0003-4422-9432)
- Sijin Yu
- Yunbo Li
Institutions
- Shanghai Ocean University (CN)
- Shanghai Maritime University (CN)
Publication Details
- Journal
- Ocean Engineering
- Published
- 2026-09-13
- DOI
- https://doi.org/10.1016/j.oceaneng.2026.127900
- Primary Topic
- Maritime Navigation and Safety
- Type
- article
- Field-Weighted Citation Impact
- 0.00