Basin-guided Reinforcement Learning for Multi-task Vehicle Routing Problems
Abstract
Multitask neural constructive solvers for vehicle routing are efficient but often lag behind classical solvers in solution quality. Combining neural construction with classical refinement can balance solution quality and efficiency. We propose a basin-guided reinforcement learning framework in which a neural solver generates intermediate solutions that are converted into feasible solutions and refined through bounded local moves. The framework builds on a simple insight: distinct candidates can reach the same refined outcome, so construction-level diversity does not necessarily translate into diversity after post-processing. This many-to-one mapping induces post-processing basins that guide learning. Alongside refined solution costs, training uses two complementary signals: a Refined-Edge Prior (REP) learns customer adjacencies from refined solutions to guide construction, while Refined-Outcome Entropy (ROE) encourages rare refined outcomes at both solution and route levels. During inference, candidates are generated by the neural network and post-processed in parallel. Optional route recombination further exploits diversity across refined candidates, forming the full pipeline. Across 16 VRP variants, the full pipeline reduces macro-average gaps to PyVRP with a 60-second per-instance budget by 69.2–81.7% relative to RouteFinder, PoMtVRS, CCL, MoSES, and RL-RFCS with the same post-processing stages, at mean end-to-end latencies of 0.11 and 0.56 seconds per instance for and , respectively. After fine-tuning on randomly generated 200-customer instances, it achieves an average gap of -0.34% on CVRP200 to the same PyVRP reference at approximately 2.6 seconds per instance with a larger inference budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.