Routing under machine breakdowns: a benchmark of dispatching rules, bandits, and reinforcement learning for multi-server flow shops

Abstract Reinforcement learning is increasingly recommended for routing decisions in flow-shop manufacturing, yet head-to-head comparisons against the dispatching rules it would replace are often weak or absent. This study addresses that gap for machine breakdowns. We benchmark five dispatching rules, two online bandit methods, and Proximal Policy Optimisation (PPO) on a multi-server flow shop simulator with opt-in non-stationarity mechanisms. Five breakdown severities plus a stationary baseline are tested: 4,200 evaluation episodes for rules and bandits across two industrial testbeds and 2,750 for PPO on electronics. On the primary electronics testbed, ShortestQueue and the Hybrid Bandit-Queue (HBQ), a diagnostic queue-aware bandit, lead decisively at mild and moderate severity. Under the heaviest disruption tested they converge to within 3% of the other rules and bandits. On these testbeds availability information confers no measurable benefit over the queue signal; on ten public Carlier-Néron topologies HBQ is the most robust method under heavy breakdowns, consistent with its availability gate. PPO trained under common public defaults performs far worse. Re-training under a corrected protocol (higher discount factor, held-out evaluation-cost checkpointing, throughput-aware reward) narrows the cost-per-unit gap to ShortestQueue on the breakdown grid from 77–137% to 38–72%, without closing it; the stationary gap is larger. All five trained policies remain costlier than ShortestQueue in all 30 cells (one-sided configuration-level signed-rank p = 0.016), and per-seed analysis identifies a bimodal training failure mode that the corrected protocol mitigates without eliminating. We conclude with a method-selection guide bounded to the disruption regimes tested.

Authors

Institutions

Publication Details

Journal
Journal of King Saud University - Engineering Sciences
Published
2026-08-26
DOI
https://doi.org/10.1007/s44444-026-00128-9
Primary Topic
Scheduling and Optimization Algorithms
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Routing under machine breakdowns: a benchmark of dispatching rules, bandits, and reinforcement learning for multi-server flow shops

Khaled R. Alrashdan
Journal of King Saud University - Engineering Sciences
Scheduling and Optimization Algorithms
article

Routing under machine breakdowns: a benchmark of dispatching rules, bandits, and reinforcement learning for multi-server flow shops

Khaled R. Alrashdan
article en

Abstract

Abstract Reinforcement learning is increasingly recommended for routing decisions in flow-shop manufacturing, yet head-to-head comparisons against the dispatching rules it would replace are often weak or absent. This study addresses that gap for machine breakdowns. We benchmark five dispatching rules, two online bandit methods, and Proximal Policy Optimisation (PPO) on a multi-server flow shop simulator with opt-in non-stationarity mechanisms. Five breakdown severities plus a stationary baseline are tested: 4,200 evaluation episodes for rules and bandits across two industrial testbeds and 2,750 for PPO on electronics. On the primary electronics testbed, ShortestQueue and the Hybrid Bandit-Queue (HBQ), a diagnostic queue-aware bandit, lead decisively at mild and moderate severity. Under the heaviest disruption tested they converge to within 3% of the other rules and bandits. On these testbeds availability information confers no measurable benefit over the queue signal; on ten public Carlier-Néron topologies HBQ is the most robust method under heavy breakdowns, consistent with its availability gate. PPO trained under common public defaults performs far worse. Re-training under a corrected protocol (higher discount factor, held-out evaluation-cost checkpointing, throughput-aware reward) narrows the cost-per-unit gap to ShortestQueue on the breakdown grid from 77–137% to 38–72%, without closing it; the stationary gap is larger. All five trained policies remain costlier than ShortestQueue in all 30 cells (one-sided configuration-level signed-rank p = 0.016), and per-seed analysis identifies a bimodal training failure mode that the corrected protocol mitigates without eliminating. We conclude with a method-selection guide bounded to the disruption regimes tested.

Journal of King Saud University - Engineering SciencesVol. 38(7)
Public Authority for Applied Education and Training (KW)
Openalex Percentile: Top 11%
Scheduling and Optimization Algorithms
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.