acceptodds
Under review as a conference paper at ICLR 2027

Do We Really Need Multi-Step Refinement for VLA Planners?

Abstract

Pretrained vision-language models bring rich semantic understanding and trajectory reasoning to autonomous driving, spurring a growing family of Vision-Language-Action (VLA) planners. Existing approaches either adopt iterative decoding paradigms inherited from text and image generation, or employ downstream refinement modules widely used in robotic manipulation. However, neither is compatible with the stringent latency of autonomous driving. Since future trajectories are low-dimensional, kinematically constrained, and orders of magnitude simpler than images or text sequences, we challenge the need for such complex pipelines in VLA planning. To this end, we present SnapVLA, a framework that generates the complete trajectory in a single parallel forward pass via masked discrete diffusion. We demonstrate that the perceived demand for multi-step refinement originates from two factors that do not inherently require it: the policy optimization paradigm and the local incoherence among trajectory points. Regarding optimization, single-step diffusion yields an exact sequence likelihood and an unbiased GRPO objective, whereas existing multi-step chains admit no exact likelihood. Regarding coherence, we tokenize trajectories into a structured action space and apply a lightweight closed-form smoother, yielding kinematically consistent trajectories without a learned continuous expert. On NAVSIM v1 and v2, SnapVLA achieves 91.0 PDMS and 90.1 EPDMS, on par with state-of-the-art models. Notably, it takes 180 ms per inference on a single A800 GPU, over 2 faster than the fastest baseline, demonstrating the potential of the single-step paradigm for efficient planning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.