Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains

Global supply chains face increasing legal, commercial, and stakeholder pressure to verify supplier sustainability rather than rely only on supplier claims. This study develops a safety-gated tabular Q-learning framework for adaptive supplier scoring, limited sourcing-share adjustment, and audit targeting in a stylized Bangladesh ready-made garment sourcing portfolio. The simulation distinguishes direct exporters, subcontractors, and small or informal suppliers; represents safety, labour, and environmental performance separately; prohibits rewards when verified safety or labour falls below 0.60; and imposes a minimum annual audit. Three independently trained agents were evaluated in 75 paired out-of-sample replications against a contextual bandit, a risk-based adaptive heuristic, a static threshold, a combined scheduled-audit-and-threshold policy, scheduled auditing, and random action. Q-learning achieved mean true compliance of 0.862 compared with 0.820 under the static threshold, increased the share of suppliers at or above 0.80 from 58.2 to 81.3%, and reduced mean severe incidents from 3.29 to 1.71. It used 167 audits compared with 362 for the combined scheduled-audit-and-threshold policy, a 53.8% reduction, while mean procurement cost was 1.1% higher than under the static threshold. The contextual bandit produced lower compliance (0.742) and more incidents (2.96), supporting the value of delayed-action learning. The risk-based heuristic produced statistically similar mean compliance but required 55 additional audits and yielded a narrower compliant-supplier share. Paired tests, effect sizes, convergence diagnostics, supplier-stratum results, credit-assignment ablation, time-state testing, and multi-parameter robustness analyses are reported. The findings remain simulation evidence rather than factory-level validation, but they provide a reproducible and technically strengthened test of adaptive supplier governance.

Authors

Institutions

Publication Details

Journal
Discover Sustainability
Published
2026-10-05
DOI
https://doi.org/10.1007/s43621-026-04745-x
Primary Topic
Sustainable Supply Chain Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains

Kazi Md. Tanvir Anzum
Discover Sustainability
Sustainable Supply Chain Management
article

Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains

Kazi Md. Tanvir Anzum
article en

Abstract

Global supply chains face increasing legal, commercial, and stakeholder pressure to verify supplier sustainability rather than rely only on supplier claims. This study develops a safety-gated tabular Q-learning framework for adaptive supplier scoring, limited sourcing-share adjustment, and audit targeting in a stylized Bangladesh ready-made garment sourcing portfolio. The simulation distinguishes direct exporters, subcontractors, and small or informal suppliers; represents safety, labour, and environmental performance separately; prohibits rewards when verified safety or labour falls below 0.60; and imposes a minimum annual audit. Three independently trained agents were evaluated in 75 paired out-of-sample replications against a contextual bandit, a risk-based adaptive heuristic, a static threshold, a combined scheduled-audit-and-threshold policy, scheduled auditing, and random action. Q-learning achieved mean true compliance of 0.862 compared with 0.820 under the static threshold, increased the share of suppliers at or above 0.80 from 58.2 to 81.3%, and reduced mean severe incidents from 3.29 to 1.71. It used 167 audits compared with 362 for the combined scheduled-audit-and-threshold policy, a 53.8% reduction, while mean procurement cost was 1.1% higher than under the static threshold. The contextual bandit produced lower compliance (0.742) and more incidents (2.96), supporting the value of delayed-action learning. The risk-based heuristic produced statistically similar mean compliance but required 55 additional audits and yielded a narrower compliant-supplier share. Paired tests, effect sizes, convergence diagnostics, supplier-stratum results, credit-assignment ablation, time-state testing, and multi-parameter robustness analyses are reported. The findings remain simulation evidence rather than factory-level validation, but they provide a reproducible and technically strengthened test of adaptive supplier governance.

Discover Sustainability
Khulna University of Engineering and Technology (BD)
Openalex Percentile: Top 8%
Sustainable Supply Chain Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains — Kazi Md. Tanvir Anzum · Discover Sustainability (2026) | TGRS Research Map | TGRS