Deep reinforcement learning-driven adaptive search algorithm for emergency rescue routing

Following natural disasters, emergency rescue route planning becomes substantially more challenging due to damaged transportation networks, urgent rescue demands, and the operational complexity of heterogeneous rescue resources. Efficient rescue decision-making under constraints including time windows, regional accessibility, and casualty survival risk is therefore of considerable practical significance. However, existing studies often suffer from inadequate modeling of casualty survival dynamics, excessive reliance on manually designed operator control strategies, and limited search robustness in highly constrained environments. To address these challenges, this paper proposes a reinforcement learning-assisted iterated greedy algorithm for heterogeneous rescue and route planning, termed Dueling Double Deep Q-Network Iterated Greedy (D3QNIG). First, a disaster rescue routing optimization model integrating casualty evacuation and relief material distribution is developed, where an injury-adaptive survival-risk mechanism is incorporated alongside capacity, time-window, and accessibility constraints. Second, a reinforcement learning-assisted iterated greedy framework is designed to adaptively control operator selection and destruction intensity during the destruction, reconstruction, and local search stages. Third, a state representation approach based on set attention pooling is constructed for variable-length and heterogeneous routing solutions, together with a hierarchical Dueling Double DQN decision architecture to enhance policy-learning stability and search effectiveness. Finally, a two-stage reward–acceptance decoupling strategy is introduced to alleviate the credit assignment issue and improve training stability. Experimental results demonstrate that the proposed D3QNIG achieves superior solution quality, faster convergence, and stronger robustness than benchmark methods, indicating its effectiveness for heterogeneous rescue route optimization in complex disaster scenarios.

Authors

Institutions

Publication Details

Journal
Expert Systems with Applications
Published
2026-09-21
DOI
https://doi.org/10.1016/j.eswa.2026.134402
Primary Topic
Evacuation and Crowd Dynamics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Deep reinforcement learning-driven adaptive search algorithm for emergency rescue routing

Biao Zhang, Leilei Meng, Xincheng Tang, Xining Cui et al.
Expert Systems with Applications
Evacuation and Crowd Dynamics
article

Deep reinforcement learning-driven adaptive search algorithm for emergency rescue routing

Biao Zhang, Leilei Meng, Xincheng Tang, Xining Cui, Yanteng Sun
article en

Abstract

Following natural disasters, emergency rescue route planning becomes substantially more challenging due to damaged transportation networks, urgent rescue demands, and the operational complexity of heterogeneous rescue resources. Efficient rescue decision-making under constraints including time windows, regional accessibility, and casualty survival risk is therefore of considerable practical significance. However, existing studies often suffer from inadequate modeling of casualty survival dynamics, excessive reliance on manually designed operator control strategies, and limited search robustness in highly constrained environments. To address these challenges, this paper proposes a reinforcement learning-assisted iterated greedy algorithm for heterogeneous rescue and route planning, termed Dueling Double Deep Q-Network Iterated Greedy (D3QNIG). First, a disaster rescue routing optimization model integrating casualty evacuation and relief material distribution is developed, where an injury-adaptive survival-risk mechanism is incorporated alongside capacity, time-window, and accessibility constraints. Second, a reinforcement learning-assisted iterated greedy framework is designed to adaptively control operator selection and destruction intensity during the destruction, reconstruction, and local search stages. Third, a state representation approach based on set attention pooling is constructed for variable-length and heterogeneous routing solutions, together with a hierarchical Dueling Double DQN decision architecture to enhance policy-learning stability and search effectiveness. Finally, a two-stage reward–acceptance decoupling strategy is introduced to alleviate the credit assignment issue and improve training stability. Experimental results demonstrate that the proposed D3QNIG achieves superior solution quality, faster convergence, and stronger robustness than benchmark methods, indicating its effectiveness for heterogeneous rescue route optimization in complex disaster scenarios.

Expert Systems with ApplicationsVol. 334
Liaocheng University (CN)
Climate action
Openalex Percentile: Top 15%
Evacuation and Crowd Dynamics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.