From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control

Modern supply chains face abrupt supply disruptions, under which the stable lead time assumption behind conventional inventory control no longer holds. Most replenishment models, including existing deep reinforcement learning (DRL) formulations, act as feedback controllers. A disruption is treated as an unobservable event, and correction begins only after deliveries are already delayed. In practice, however, disruptions are often preceded by observable precursor signals. This paper develops a risk sensing DRL framework that incorporates such signals into the state of a proximal policy optimization (PPO) agent, introducing feedforward control into the replenishment decision. The signal quality (detection rate) and the predictive horizon are treated as explicit design parameters. In a two-echelon system with graded disruptions, the risk sensing policy achieves substantial cost savings against both a per-environment calibrated (s, Q) policy and an identically trained reactive DRL policy, even with an imperfect signal. The required horizon is short, as horizons longer than needed bring no further gain. In contrast, the savings depend critically on the signal quality, as a meaningful gain over the calibrated benchmark requires a signal that detects well over half of upcoming disruptions. The advantage is greatest where severe disruptions are infrequent. The learned policy operates a state-dependent reorder point that stays low in calm conditions and rises with the severity and proximity of a predicted threat. A transparent rule driven by the same signal captures about 83% of the gain, indicating that most of the value comes from the information itself rather than from the learning method. These results quantify the value of advance supply information in inventory control and show how it supports the transition from reactive response to proactive mitigation.

Authors

Institutions

Publication Details

Journal
Systems
Published
2026-09-17
DOI
https://doi.org/10.3390/systems14091164
Primary Topic
Supply Chain Resilience and Risk Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control

Yong Won Seo, Hyuksoo Han
Systems
Supply Chain Resilience and Risk Management
article

From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control

Yong Won Seo, Hyuksoo Han
article en

Abstract

Modern supply chains face abrupt supply disruptions, under which the stable lead time assumption behind conventional inventory control no longer holds. Most replenishment models, including existing deep reinforcement learning (DRL) formulations, act as feedback controllers. A disruption is treated as an unobservable event, and correction begins only after deliveries are already delayed. In practice, however, disruptions are often preceded by observable precursor signals. This paper develops a risk sensing DRL framework that incorporates such signals into the state of a proximal policy optimization (PPO) agent, introducing feedforward control into the replenishment decision. The signal quality (detection rate) and the predictive horizon are treated as explicit design parameters. In a two-echelon system with graded disruptions, the risk sensing policy achieves substantial cost savings against both a per-environment calibrated (s, Q) policy and an identically trained reactive DRL policy, even with an imperfect signal. The required horizon is short, as horizons longer than needed bring no further gain. In contrast, the savings depend critically on the signal quality, as a meaningful gain over the calibrated benchmark requires a signal that detects well over half of upcoming disruptions. The advantage is greatest where severe disruptions are infrequent. The learned policy operates a state-dependent reorder point that stays low in calm conditions and rises with the severity and proximity of a predicted threat. A transparent rule driven by the same signal captures about 83% of the gain, indicating that most of the value comes from the information itself rather than from the learning method. These results quantify the value of advance supply information in inventory control and show how it supports the transition from reactive response to proactive mitigation.

SystemsVol. 14(9)
Chung-Ang University (KR)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Supply Chain Resilience and Risk Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control — Yong Won Seo, Hyuksoo Han · Systems (2026) | TGRS Research Map | TGRS