Machine Learning–based Late Delivery Risk Prediction in Global Supply Chain Operations
This study develops a pre-shipment machine-learning framework for estimating the probability that an order will be delivered late in global supply-chain operations. Using the APL Logistics supply-chain dataset, the study analyzes 180,519 records and applies leakage-aware preprocessing, feature engineering, and machine-learning classification. Personally identifiable information, high-cardinality fields, redundant variables, and post-shipment information were excluded from the modeling boundary. Three classifiers—Logistic Regression, Random Forest, and XGBoost—were evaluated. XGBoost achieved the strongest performance on the independent test set, with a ROC-AUC of 0.7745, precision of 84.12%, recall of 56.35%, and F1 score of 0.6749 at the standard 0.50 decision threshold. The framework additionally translates predicted probabilities into Low (<40%), Medium (40–70%), and High (≥70%) operational risk tiers. The High-Risk tier achieved 89.7% precision on the held-out test set, while predictions at or above 0.80 probability achieved 95.6% precision. SHAP explainability and an interactive Streamlit dashboard are incorporated to support transparent risk interpretation, regional and shipping-mode analysis, and operational prioritization. The work demonstrates how predictive analytics can complement traditional retrospective logistics analysis by identifying potentially high-risk orders before shipment, while recognizing limitations related to historical data, external operational conditions, carrier information, and deployment calibration.
Authors
- Fabian Biju
Institutions
- Siemens (Hungary) (HU)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-11
- DOI
- https://doi.org/10.5281/zenodo.22705410
- Primary Topic
- Supply Chain Resilience and Risk Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00