Do We Really Need Multi-Step Refinement for VLA Planners?
Abstract
Pretrained vision-language models bring rich semantic understanding and trajectory reasoning to autonomous driving, spurring a growing family of Vision-Language-Action (VLA) planners. Existing approaches either adopt iterative decoding paradigms inherited from text and image generation, or employ downstream refinement modules widely used in robotic manipulation. However, neither is compatible with the stringent latency of autonomous driving. Since future trajectories are low-dimensional, kinematically constrained, and orders of magnitude simpler than images or text sequences, we challenge the need for such complex pipelines in VLA planning. To this end, we present SnapVLA, a framework that generates the complete trajectory in a single parallel forward pass via masked discrete diffusion. We demonstrate that the perceived demand for multi-step refinement originates from two factors that do not inherently require it: the policy optimization paradigm and the local incoherence among trajectory points. Regarding optimization, single-step diffusion yields an exact sequence likelihood and an unbiased GRPO objective, whereas existing multi-step chains admit no exact likelihood. Regarding coherence, we tokenize trajectories into a structured action space and apply a lightweight closed-form smoother, yielding kinematically consistent trajectories without a learned continuous expert. On NAVSIM v1 and v2, SnapVLA achieves 91.0 PDMS and 90.1 EPDMS, on par with state-of-the-art models. Notably, it takes 180 ms per inference on a single A800 GPU, over 2 faster than the fastest baseline, demonstrating the potential of the single-step paradigm for efficient planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.