Modular dual-critic reinforcement learning control for AUV path tracking and collision avoidance

Autonomous underwater vehicles (AUVs) operating in complex and uncertain ocean environments require reliable path-tracking and obstacle-avoidance capabilities to maintain safety in the presence of currents and obstacles. Traditional end-to-end deep reinforcement learning (DRL) methods often struggle with coupled navigation objectives, which can lead to oscillatory control and unreliable trajectory recovery. This paper proposes a modular dual-critic proximal policy optimization framework with prioritized experience replay (Modular DCPPO-PER) for AUV path tracking and collision avoidance. The framework separates navigation into two specialized policies for path tracking and collision avoidance and coordinates them through a time-to-collision-based soft risk arbitration mechanism. An actor-dual-critic-PER (ADCP) architecture is further introduced to improve learning stability and sample utilization by combining dual buffers with task-specific value estimation. In addition, a direction-sensitive reward function based on the sensor field of view assigns higher weights to high-risk perceptual regions and promotes smoother avoidance maneuvers. Simulation results in complex marine environments show that, compared with the advanced DRL baseline ARAB-PPO, the proposed method improves the average obstacle-avoidance success rate by 49.8% and reduces the path-tracking error by 12.9%. A USV-based physical surrogate experiment is also conducted to examine real-time inference, control continuity, and actuator-level feasibility on real hardware. The results indicate that the proposed framework can provide robust and smooth navigation behavior for AUV-oriented marine robotic applications, while full underwater AUV field validation remains future work.

Authors

Institutions

Publication Details

Journal
Ocean Engineering
Published
2026-10-07
DOI
https://doi.org/10.1016/j.oceaneng.2026.128446
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Modular dual-critic reinforcement learning control for AUV path tracking and collision avoidance

Dongfang Ma, Songjian Lv, Xinren Jiang
Ocean Engineering
Reinforcement Learning in Robotics
article

Modular dual-critic reinforcement learning control for AUV path tracking and collision avoidance

Dongfang Ma, Songjian Lv, Xinren Jiang
article en

Abstract

Autonomous underwater vehicles (AUVs) operating in complex and uncertain ocean environments require reliable path-tracking and obstacle-avoidance capabilities to maintain safety in the presence of currents and obstacles. Traditional end-to-end deep reinforcement learning (DRL) methods often struggle with coupled navigation objectives, which can lead to oscillatory control and unreliable trajectory recovery. This paper proposes a modular dual-critic proximal policy optimization framework with prioritized experience replay (Modular DCPPO-PER) for AUV path tracking and collision avoidance. The framework separates navigation into two specialized policies for path tracking and collision avoidance and coordinates them through a time-to-collision-based soft risk arbitration mechanism. An actor-dual-critic-PER (ADCP) architecture is further introduced to improve learning stability and sample utilization by combining dual buffers with task-specific value estimation. In addition, a direction-sensitive reward function based on the sensor field of view assigns higher weights to high-risk perceptual regions and promotes smoother avoidance maneuvers. Simulation results in complex marine environments show that, compared with the advanced DRL baseline ARAB-PPO, the proposed method improves the average obstacle-avoidance success rate by 49.8% and reduces the path-tracking error by 12.9%. A USV-based physical surrogate experiment is also conducted to examine real-time inference, control continuity, and actuator-level feasibility on real hardware. The results indicate that the proposed framework can provide robust and smooth navigation behavior for AUV-oriented marine robotic applications, while full underwater AUV field validation remains future work.

Ocean EngineeringVol. 368
Zhejiang University (CN)
Openalex Percentile: Top 12%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Modular dual-critic reinforcement learning control for AUV path tracking and collision avoidance — Dongfang Ma, Songjian Lv, et al. · Ocean Engineering (2026) | TGRS Research Map | TGRS