Hierarchical collaborative multi-agent reinforcement learning for hot-rolling production planning with order-splitting flexibility
In mass personalized hot rolling, intricate constraints cause load imbalances and low order fulfilment. While order splitting alleviates these bottlenecks, it increases changeover frequency and planning complexity. We propose a bi-level model: the upper level optimizes production cost and time, while the lower minimizes makespan and shutdowns. To address the strong coupling between allocation, splitting and sequencing, a hierarchical cooperative multi-agent reinforcement learning (HCMARL) algorithm is developed using a cross-module closed-loop decoupling mechanism. Decoupled adaptive multi-objective deep deterministic policy gradient (DAM-DDPG) and collaborative constraint optimization Monte Carlo tree search (CCO-MCTS) process continuous allocation and splitting decisions to generate physical constraints. These guide the lower-level flexible coupling production planning optimization (FC-PPO), combined with a corrected generalized advantage estimation (CE-GAE), to synchronize maintenance windows and process adjustments. Experiments show HCMARL closely approaches Gurobi’s exact solutions with significantly higher computational efficiency for industrial-scale problems.
Authors
- Ruilin Pan (ORCID: https://orcid.org/0000-0002-8268-739X)
- Jianhua Cao (ORCID: https://orcid.org/0000-0001-6263-5018)
- Xinyu Jin
- Fan Fei
Institutions
- Anhui University of Technology (CN)
Publication Details
- Journal
- Engineering Optimization
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1080/0305215x.2026.2733931
- Primary Topic
- Scheduling and Optimization Algorithms
- Type
- article
- Field-Weighted Citation Impact
- 0.00