acceptodds
Under review as a conference paper at ICLR 2027

WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) policies generate robot controls autoregressively, making closed-loop latency dominated by repeated target model forward passes. Speculative decoding reduces this cost by verifying blocks of draft action tokens in parallel, and recent VLA methods relax token-level acceptance because small differences in action token space often map to similar continuous controls. However, a fixed token-distance tolerance does not account for the scene-dependent consequences of action deviations: deviations that are harmless in free space can cause collisions or grasp failures near contact. We propose WA-SpecDec, a world-aware speculative decoding framework that injects world-model-derived physical scene information during the VLA prefill stage. The resulting world-aware prefill states are shared between draft generation and target verification, while the relaxed acceptance rule remains unchanged. Evaluated across four VLA policies and three relaxed acceptance rules on LIBERO, SIMPLER-Env, and four real-robot tasks, WA-SpecDec improves the trade-off between inference speed and task success. Across five equally weighted policy–benchmark settings, it increases task success by an average of percentage points over the speculative baselines and achieves an average inference speedup of over autoregressive decoding of the base policies. For the two policies evaluated on LIBERO, it reduces the near-contact failure (NCF) metric by on average relative to the speculative baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.