Hierarchical Multi-Agent Reinforcement Learning for Cooperative Wildfire Suppression with Amphibious Firefighting UAVs

Cooperative task planning for amphibious firefighting UAVs deployed across multiple bases is challenging because of complex resource coupling, evolving fire conditions, and real-time decision requirements. This study proposes an H-MAPPO-LSTM method with a three-level hierarchical multi-agent structure: a Command-Center Agent performs global situation assessment and macroscopic task allocation, Base Agents manage UAV resources and decompose regional tasks, and UAV Agents execute water collection, flight, water-dropping, and return operations. LSTM networks are incorporated into the policy networks and combined with PPO and centralized training with decentralized execution to enable multilevel cooperative learning and online dynamic replanning. During deployment, network parameters remain fixed, while online policy inference, real-time decision-making, and dynamic replanning respond to fire-scene changes. The model jointly considers UAV bases, amphibious firefighting UAVs, natural water sources, dynamic fire hotspots, and their constraints. Compared with Flat MAPPO in the static benchmark, H-MAPPO-LSTM reduces mission completion time by 20.2%, average response time by 14.7%, final burned area by 29.7%, and comprehensive cost by 26.0%. Fire-containment rates reach 88% in highly dynamic scenarios and 81.96% in large-scale scenarios. Ablation results indicate that the hierarchical architecture provides the most pronounced contribution under the tested settings, while LSTM, the level-specific centralized critics, and hierarchical rewards provide additional performance contributions under the tested settings.

Authors

Institutions

Publication Details

Journal
Drones
Published
2026-09-25
DOI
https://doi.org/10.3390/drones10100730
Primary Topic
Fire Detection and Safety Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hierarchical Multi-Agent Reinforcement Learning for Cooperative Wildfire Suppression with Amphibious Firefighting UAVs

Yilin Han, Jiandong Zhang, Qiming Yang, Siyuan Wang et al.
Drones
Fire Detection and Safety Systems
article

Hierarchical Multi-Agent Reinforcement Learning for Cooperative Wildfire Suppression with Amphibious Firefighting UAVs

Yilin Han, Jiandong Zhang, Qiming Yang, Siyuan Wang, Jianfeng Xie, Shuling Dai
article en

Abstract

Cooperative task planning for amphibious firefighting UAVs deployed across multiple bases is challenging because of complex resource coupling, evolving fire conditions, and real-time decision requirements. This study proposes an H-MAPPO-LSTM method with a three-level hierarchical multi-agent structure: a Command-Center Agent performs global situation assessment and macroscopic task allocation, Base Agents manage UAV resources and decompose regional tasks, and UAV Agents execute water collection, flight, water-dropping, and return operations. LSTM networks are incorporated into the policy networks and combined with PPO and centralized training with decentralized execution to enable multilevel cooperative learning and online dynamic replanning. During deployment, network parameters remain fixed, while online policy inference, real-time decision-making, and dynamic replanning respond to fire-scene changes. The model jointly considers UAV bases, amphibious firefighting UAVs, natural water sources, dynamic fire hotspots, and their constraints. Compared with Flat MAPPO in the static benchmark, H-MAPPO-LSTM reduces mission completion time by 20.2%, average response time by 14.7%, final burned area by 29.7%, and comprehensive cost by 26.0%. Fire-containment rates reach 88% in highly dynamic scenarios and 81.96% in large-scale scenarios. Ablation results indicate that the hierarchical architecture provides the most pronounced contribution under the tested settings, while LSTM, the level-specific centralized critics, and hierarchical rewards provide additional performance contributions under the tested settings.

DronesVol. 10(10)
Northwestern Polytechnical University (CN), Beihang University (CN)
Openalex Percentile: Top 12%
Fire Detection and Safety Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Hierarchical Multi-Agent Reinforcement Learning for Cooperative Wildfire Suppression with Amphibious Firefighting UAVs — Yilin Han, Jiandong Zhang, et al. · Drones (2026) | TGRS Research Map | TGRS