Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility

In mass personalized hot rolling, intricate constraints cause load imbalances and low order fulfilment. While order splitting alleviates these bottlenecks, it increases changeover frequency and planning complexity. We propose a bi-level model: the upper level optimizes production cost and time, while the lower minimizes makespan and shutdowns. To address the strong coupling between allocation, splitting and sequencing, a hierarchical cooperative multi-agent reinforcement learning (HCMARL) algorithm is developed using a cross-module closed-loop decoupling mechanism. Decoupled adaptive multi-objective deep deterministic policy gradient (DAM-DDPG) and collaborative constraint optimization Monte Carlo tree search (CCO-MCTS) process continuous allocation and splitting decisions to generate physical constraints. These guide the lower-level flexible coupling production planning optimization (FC-PPO), combined with a corrected generalized advantage estimation (CE-GAE), to synchronize maintenance windows and process adjustments. Experiments show HCMARL closely approaches Gurobi’s exact solutions with significantly higher computational efficiency for industrial-scale problems.

Authors

Institutions

Publication Details

Journal
Engineering Optimization
Published
2026-10-05
DOI
https://doi.org/10.1080/0305215x.2026.2733931
Primary Topic
Scheduling and Optimization Algorithms
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility

Ruilin Pan, Jianhua Cao, Xinyu Jin, Fan Fei
Engineering Optimization
Scheduling and Optimization Algorithms
article

Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility

Ruilin Pan, Jianhua Cao, Xinyu Jin, Fan Fei
article en

Abstract

In mass personalized hot rolling, intricate constraints cause load imbalances and low order fulfilment. While order splitting alleviates these bottlenecks, it increases changeover frequency and planning complexity. We propose a bi-level model: the upper level optimizes production cost and time, while the lower minimizes makespan and shutdowns. To address the strong coupling between allocation, splitting and sequencing, a hierarchical cooperative multi-agent reinforcement learning (HCMARL) algorithm is developed using a cross-module closed-loop decoupling mechanism. Decoupled adaptive multi-objective deep deterministic policy gradient (DAM-DDPG) and collaborative constraint optimization Monte Carlo tree search (CCO-MCTS) process continuous allocation and splitting decisions to generate physical constraints. These guide the lower-level flexible coupling production planning optimization (FC-PPO), combined with a corrected generalized advantage estimation (CE-GAE), to synchronize maintenance windows and process adjustments. Experiments show HCMARL closely approaches Gurobi’s exact solutions with significantly higher computational efficiency for industrial-scale problems.

Engineering Optimization
Anhui University of Technology (CN)
Openalex Percentile: Top 11%
Scheduling and Optimization Algorithms
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility — Ruilin Pan, Jianhua Cao, et al. · Engineering Optimization (2026) | TGRS Research Map | TGRS