Explainable reinforcement learning for smart production scheduling in automotive manufacturing
Automotive factories must decide which lot to process next and they must decide fast. Shop floor conditions keep changing so the best choice is rarely obvious. Reinforcement learning can support these decisions but a trained policy acts like a black box so managers cannot see why it picks one lot. This study explains a policy already trained by CMA-ES policy search and finds the factors linked to its choices. Queue-wise SHAP and attention analysis and queue-constrained counterfactuals and decision tree surrogates and heuristic overlap were applied to two benchmark systems MiniPlant and AMT2020. The SHAP results give processing time and machine utilization the largest attributions to the dispatching score. The counterfactuals point to different variables: remaining cycle time and order priority and queue length in MiniPlant and setup time and remaining cycle time in AMT2020. In MiniPlant the decision tree surrogate matched the policy well with F1-scores of 0.8234 to 0.8456 and accuracy of 0.9042 to 0.9674. In AMT2020 it was weaker with F1-scores of 0.5645 to 0.6732 although accuracy stayed high at 0.9345 to 0.9743. The policy behaved like standard rules and matched SRPT in 35.60% of assembly line queues and FIFO in 25.10%. In the paint shop it agreed with the setup rule in 78.5% of queues and a three-way match with EDD reached 40.00%. The two views disagree so score attribution must not be read as causal influence. The work describes one trained policy from several angles and shows that explanations alone cannot prove better performance.
Authors
- Kazi Md. Tanvir Anzum (ORCID: https://orcid.org/0009-0009-7422-350X)
- Md. Mahfuzur Rahman
Institutions
- Khulna University of Engineering and Technology (BD)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-08
- DOI
- https://doi.org/10.1038/s41598-026-69106-0
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00