Real-Time Neural Vehicle Routing under Delivery Commitments
Abstract
Vehicle routing in the real world is inherently dynamic, where travel times fluctuate and orders are cancelled or released mid-operation, demanding real-time optimization. Existing Neural Vehicle Routing (NVR) methods can infer high-quality routes within milliseconds, yet they remain static and cannot respect delivery commitments under dynamic events. In this paper, we propose a Real-Time NVR (RT-NVR) approach, which extends NVR from static optimization to online decision-making. During training, a View-Comparative Policy Optimization (VCPO) method is designed to train an encoder-decoder model with state broadcasting, where the performance gap between geometrically equivalent views is turned into a learning signal via a contrast-adaptive baseline. In addition, a Spatiotemporal Compatibility-Guided Sampling (SCGS) strategy is devised to inject spatiotemporal priors into decoding, modulating logits toward promising actions. During real-time deployment, a Delivery Commitment Set-based Masking (DCSM) mechanism derives the vehicle load peak and maximum slack time of the delivery commitment set for a vehicle, establishing a provably safe decision boundary. Extensive experiments demonstrate that VCPO outperforms state-of-the-art training methods, and that the real-time approach reduces travel cost, fleet size, and time delay by 19.16%, 42.14%, and 96.73%, respectively, over heuristic methods. Statistical analyses further reveal that state broadcasting carries the most salient routing knowledge, and that the model parameters evolve along a low-dimensional manifold whose top 10 principal components capture 90.6% of the variance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.