Explainable reinforcement learning for smart production scheduling in automotive manufacturing

Automotive factories must decide which lot to process next and they must decide fast. Shop floor conditions keep changing so the best choice is rarely obvious. Reinforcement learning can support these decisions but a trained policy acts like a black box so managers cannot see why it picks one lot. This study explains a policy already trained by CMA-ES policy search and finds the factors linked to its choices. Queue-wise SHAP and attention analysis and queue-constrained counterfactuals and decision tree surrogates and heuristic overlap were applied to two benchmark systems MiniPlant and AMT2020. The SHAP results give processing time and machine utilization the largest attributions to the dispatching score. The counterfactuals point to different variables: remaining cycle time and order priority and queue length in MiniPlant and setup time and remaining cycle time in AMT2020. In MiniPlant the decision tree surrogate matched the policy well with F1-scores of 0.8234 to 0.8456 and accuracy of 0.9042 to 0.9674. In AMT2020 it was weaker with F1-scores of 0.5645 to 0.6732 although accuracy stayed high at 0.9345 to 0.9743. The policy behaved like standard rules and matched SRPT in 35.60% of assembly line queues and FIFO in 25.10%. In the paint shop it agreed with the setup rule in 78.5% of queues and a three-way match with EDD reached 40.00%. The two views disagree so score attribution must not be read as causal influence. The work describes one trained policy from several angles and shows that explanations alone cannot prove better performance.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-08
DOI
https://doi.org/10.1038/s41598-026-69106-0
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Explainable reinforcement learning for smart production scheduling in automotive manufacturing

Kazi Md. Tanvir Anzum, Md. Mahfuzur Rahman
Scientific Reports
Explainable Artificial Intelligence (XAI)
article

Explainable reinforcement learning for smart production scheduling in automotive manufacturing

Kazi Md. Tanvir Anzum, Md. Mahfuzur Rahman
article en

Abstract

Automotive factories must decide which lot to process next and they must decide fast. Shop floor conditions keep changing so the best choice is rarely obvious. Reinforcement learning can support these decisions but a trained policy acts like a black box so managers cannot see why it picks one lot. This study explains a policy already trained by CMA-ES policy search and finds the factors linked to its choices. Queue-wise SHAP and attention analysis and queue-constrained counterfactuals and decision tree surrogates and heuristic overlap were applied to two benchmark systems MiniPlant and AMT2020. The SHAP results give processing time and machine utilization the largest attributions to the dispatching score. The counterfactuals point to different variables: remaining cycle time and order priority and queue length in MiniPlant and setup time and remaining cycle time in AMT2020. In MiniPlant the decision tree surrogate matched the policy well with F1-scores of 0.8234 to 0.8456 and accuracy of 0.9042 to 0.9674. In AMT2020 it was weaker with F1-scores of 0.5645 to 0.6732 although accuracy stayed high at 0.9345 to 0.9743. The policy behaved like standard rules and matched SRPT in 35.60% of assembly line queues and FIFO in 25.10%. In the paint shop it agreed with the setup rule in 78.5% of queues and a three-way match with EDD reached 40.00%. The two views disagree so score attribution must not be read as causal influence. The work describes one trained policy from several angles and shows that explanations alone cannot prove better performance.

Scientific Reports
Khulna University of Engineering and Technology (BD)
Openalex Percentile: Top 8%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Explainable reinforcement learning for smart production scheduling in automotive manufacturing — Kazi Md. Tanvir Anzum, Md. Mahfuzur Rahman · Scientific Reports (2026) | TGRS Research Map | TGRS