acceptodds
Under review as a conference paper at ICLR 2027

Spec-WAM: Feedback Grounded Speculative Inference for World–Action Models

Abstract

World-action models (WAMs) leverage video-generation priors for robot action generation, but iterative denoising delays their response to new observations. Computing ahead during action execution can reduce this delay, yet predicted video and actions may become inconsistent with the observations that actually arrive. We introduce Spec-WAM, a training-free framework that tightly couples speculation across diffusion denoising and control loops. More specifically, during denoising sampling, extrapolation provides candidate action-denoising states for batched verification, amortizing model evaluation overhead. Within the control loop, the same pretrained WAM autoregressively prepares the next video-action trajectory while the robot executes its current action. After receiving the new observation, Spec-WAM checks whether the drafted video can be reused for action generation. For models that generate video before actions, it can reuse the accepted video together with its paired draft actions. For models that jointly denoise video and actions, a shared model query checks both streams, and action denoising continues from the last state that passes verification instead of restarting from noise. On RoboTwin 2.0 with unseen instructions, Spec-WAM matches the task success of OpenWAM's released inference on paired episodes (92.49% versus 92.65%) while significantly reducing mean feedback-to-action latency by 21.7%. On Next Forcing, it keeps task success and lowers this latency 34-fold, including the effect of fewer video denoising steps and of batched checking, which cuts sequential action-model calls from 51 to 4.02.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.