A Deep Reinforcement Learning Framework for Dynamic Routing of Distant-Water Squid-Jigging Vessels: PPO-Based Retrospective Simulation in the Peruvian Jumbo Flying Squid Fishery

Distant-water squid-jigging operations require continuous route decisions that balance expected fishing returns and fuel consumption under spatially heterogeneous fishery resources and oceanographic conditions. This study proposes a Proximal Policy Optimization (PPO)-based dynamic routing framework for the Peruvian jumbo flying squid fishery. A data-informed gridded simulator integrates oceanographic variables, historical fishing-ground information, a wave-induced fuel adjustment, and an MVT-Inspired Local Depletion Mechanism. The routing task is formulated as an augmented-state Markov decision process with a multi-component reward function considering economic return, historical fishing-ground guidance, wave-related operating effects, and operational inefficiency. Vessel-level cross-fitting is used to reduce data reuse between reward-environment construction and historical-prior construction, and the policy is evaluated retrospectively in a held-out 2021 simulation environment. Across three independent training seeds, PPO achieved 257.42 ± 13.92 t of Cumulative Simulated Catch, 256.59 ± 7.83 t of Cumulative Fuel Consumption, and a Simplified Simulated Operating Margin of 129,376.80 ± 24,244.68 USD. Relative to the prespecified representative Historical Trajectory Replay, these results represent 13.77% higher simulated catch, 8.58% lower simulated fuel consumption, and 85.88% higher simulated operating margin. Additional sensitivity analyses showed that the qualitative catch–fuel–margin advantage was retained under moderate perturbations of the held-out catch field and across ±20% squid- and fuel-price variations. The results indicate that PPO provides a more favorable and comparatively stable catch–fuel–margin trade-off within the constructed retrospective simulation framework.

Authors

Institutions

Publication Details

Journal
Journal of Marine Science and Engineering
Published
2026-09-13
DOI
https://doi.org/10.3390/jmse14181701
Primary Topic
Cephalopods and Marine Biology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Deep Reinforcement Learning Framework for Dynamic Routing of Distant-Water Squid-Jigging Vessels: PPO-Based Retrospective Simulation in the Peruvian Jumbo Flying Squid Fishery

Chun‐Hsien Chen, Tianjiao Zhang, Bo Song, Hu Li et al.
Journal of Marine Science and Engineering
Cephalopods and Marine Biology
article

A Deep Reinforcement Learning Framework for Dynamic Routing of Distant-Water Squid-Jigging Vessels: PPO-Based Retrospective Simulation in the Peruvian Jumbo Flying Squid Fishery

Chun‐Hsien Chen, Tianjiao Zhang, Bo Song, Hu Li, Yimeng Zhang
article en

Abstract

Distant-water squid-jigging operations require continuous route decisions that balance expected fishing returns and fuel consumption under spatially heterogeneous fishery resources and oceanographic conditions. This study proposes a Proximal Policy Optimization (PPO)-based dynamic routing framework for the Peruvian jumbo flying squid fishery. A data-informed gridded simulator integrates oceanographic variables, historical fishing-ground information, a wave-induced fuel adjustment, and an MVT-Inspired Local Depletion Mechanism. The routing task is formulated as an augmented-state Markov decision process with a multi-component reward function considering economic return, historical fishing-ground guidance, wave-related operating effects, and operational inefficiency. Vessel-level cross-fitting is used to reduce data reuse between reward-environment construction and historical-prior construction, and the policy is evaluated retrospectively in a held-out 2021 simulation environment. Across three independent training seeds, PPO achieved 257.42 ± 13.92 t of Cumulative Simulated Catch, 256.59 ± 7.83 t of Cumulative Fuel Consumption, and a Simplified Simulated Operating Margin of 129,376.80 ± 24,244.68 USD. Relative to the prespecified representative Historical Trajectory Replay, these results represent 13.77% higher simulated catch, 8.58% lower simulated fuel consumption, and 85.88% higher simulated operating margin. Additional sensitivity analyses showed that the qualitative catch–fuel–margin advantage was retained under moderate perturbations of the held-out catch field and across ±20% squid- and fuel-price variations. The results indicate that PPO provides a more favorable and comparatively stable catch–fuel–margin trade-off within the constructed retrospective simulation framework.

Journal of Marine Science and EngineeringVol. 14(18)
Nanyang Technological University (SG), Shanghai Ocean University (CN), Shanghai Maritime University (CN)
Life below water
Openalex Percentile: Top 7%
Cephalopods and Marine Biology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.