Solver-Independent MAPPO for Disruption-Aware Last-Mile Routing via Candidate Parcel-Locker Consolidation Points
Urban last-mile routes must adapt when demand and network conditions change after dispatch. This study proposes a solver-independent multi-agent proximal policy optimization (MAPPO) framework that combines a shared fleet policy, dynamically ranked candidate parcel-locker consolidation points, disruption-aware observations, fleet-level value learning, and structural feasibility masking. Evaluation on a Los Angeles road-network simulation used chronological data splits, ten random seeds, matched routing baselines, and compound-disruption stress tests. MAPPO served 98.4% of demand versus 96.6% for static OR-Tools and reduced failed packages by 52.2%; under severe compound disruption, service increased from 89.4% to 93.0%. Higher-service adaptive methods reached 99.6–99.7%, with summed online computation times of 0.111–0.503 s per episode compared with 0.029 s for MAPPO. The integrated framework therefore improves disruption recovery over static planning while exposing a clear quality–computation trade-off against adaptive search.
Authors
- Jamal Benhra (ORCID: https://orcid.org/0000-0002-8883-4610)
- Kadim Lahcen Nadime (ORCID: https://orcid.org/0000-0001-9350-7422)
- Mohamed-Ali Ejjanfi (ORCID: https://orcid.org/0009-0001-5492-2815)
Institutions
- University of Hassan II Casablanca (MA)
Publication Details
- Journal
- Computation
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/computation14090223
- Primary Topic
- Transportation Planning and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00