Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions

Supply chain disruptions propagate rapidly through multi-echelon networks, and most learning-based approaches stop at prediction rather than acting on it. This paper advances from disruption forecasting toward autonomous mitigation by framing dynamic inventory rebalancing and last-mile fulfillment as a cooperative multi-agent reinforcement learning problem. Each facility in the network is an independent agent that jointly decides (i) replenishment and lateral transshipment quantities to rebalance inventory across echelons, and (ii) fulfillment assignments that reroute customer orders through available last-mile capacity when a disruption degrades primary routes. The agents are trained under a centralized-training–decentralized-execution paradigm with a graph-neural-network state encoder and a clipped proximal-policy-optimization core and are exposed during training to a stochastic disruption generator so that mitigation policies are learned proactively rather than reactively. Experiments on a three-echelon network under supplier-failure, hub-failure, and transport-disruption scenarios show that the learned policy sustains service levels, shortens recovery time, and reduces total disruption cost relative to base-stock and single-agent baselines. The results demonstrate that prediction must be coupled with autonomous optimization to deliver practical resilience.

Authors

Institutions

Publication Details

Journal
Iconic Research and Engineering Journals
Published
2026-08-26
DOI
https://doi.org/10.64388/irev8i9-1722594
Primary Topic
Supply Chain and Inventory Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions

Sohail Sayed, Nauman Sayed
Iconic Research and Engineering Journals
Supply Chain and Inventory Management
article

Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions

Sohail Sayed, Nauman Sayed
article en

Abstract

Supply chain disruptions propagate rapidly through multi-echelon networks, and most learning-based approaches stop at prediction rather than acting on it. This paper advances from disruption forecasting toward autonomous mitigation by framing dynamic inventory rebalancing and last-mile fulfillment as a cooperative multi-agent reinforcement learning problem. Each facility in the network is an independent agent that jointly decides (i) replenishment and lateral transshipment quantities to rebalance inventory across echelons, and (ii) fulfillment assignments that reroute customer orders through available last-mile capacity when a disruption degrades primary routes. The agents are trained under a centralized-training–decentralized-execution paradigm with a graph-neural-network state encoder and a clipped proximal-policy-optimization core and are exposed during training to a stochastic disruption generator so that mitigation policies are learned proactively rather than reactively. Experiments on a three-echelon network under supplier-failure, hub-failure, and transport-disruption scenarios show that the learned policy sustains service levels, shortens recovery time, and reduces total disruption cost relative to base-stock and single-agent baselines. The results demonstrate that prediction must be coupled with autonomous optimization to deliver practical resilience.

Iconic Research and Engineering JournalsVol. 8(9)
California State University, Fullerton (US)
Openalex Percentile: Top 5%
Supply Chain and Inventory Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.