Reinforcement learning for initializing genetic algorithms in vehicle routing
Abstract Vehicle routing problems (VRP) are an extension of the Traveling Salesperson Problem and are a fundamental NP-hard challenge in combinatorial optimization. Solving VRP in real-time at large scale has become critical in numerous applications, from growing markets like last-mile delivery to emerging use-cases like interactive logistics planning. Such applications involve solving similar VRP instances repeatedly, yet current state-of-the-art solvers treat each instance on its own without leveraging previous examples. We introduce an optimization framework where a reinforcement learning agent is trained on prior instances and quickly generates initial solutions, which are then further optimized by a genetic algorithm. This framework, Evolutionary Algorithm with Reinforcement Learning Initialization ( EARLI ), consistently outperforms current state-of-the-art solvers under limited time budgets. For example, EARLI handles vehicle routing with 500 locations within one second, 10x faster than current solvers for the same solution quality, enabling real-time and interactive routing at scale. EARLI can generalize to new data, as demonstrated on real e-commerce delivery data of a previously unseen city.
Authors
- Hugo Linsenmaier
- Rajesh Gandham
- Ido Greenberg
- Shie Mannor (ORCID: https://orcid.org/0000-0003-4439-7647)
- Piotr Sielski (ORCID: https://orcid.org/0000-0002-4594-1854)
- Gal Chechik (ORCID: https://orcid.org/0000-0001-9164-5303)
- Alex Fender (ORCID: https://orcid.org/0000-0002-0903-1407)
- Eli Meirom
Institutions
- Nvidia (United States) (US)
Publication Details
- Journal
- Communications AI & Computing
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1038/s44488-026-00019-7
- Primary Topic
- Vehicle Routing Optimization Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00