DiSpecWAM: Distillation and Speculative Inference for World Action Models
Abstract
Current World Action Models (WAMs) face a fundamental tension between world-modeling capability and real-time interaction efficiency. Future video prediction helps model environment dynamics and reason about the consequences of actions, but multi-step video–action generation incurs substantial inference overhead. Existing methods either remove video prediction at the cost of essential world-modeling capabilities, or aggressively compress the generation process, making it difficult to preserve the performance of the full model. To address capability preservation and error recovery in few-step generation, we propose DiSpecWAM. First, F²Distill converts a full-capacity WAM into a complementary hierarchy comprising a low-budget Draft Model and a higher-fidelity Target Model. By coupling teacher-flow preservation with Student-rollout distribution calibration, both retain reliable visual foresight and precise action generation despite extreme denoising reduction. Then, R²Spec organizes them into repairable and reliable speculative inference, where the Target Model selectively verifies, repairs, and falls back on unreliable Draft Model outputs. This design couples high-fidelity compression during training with adaptive recovery during inference. On RoboTwin 2.0, DiSpecWAM achieves 92.65% average success at 205 ms per action chunk, delivering a 16.61× speedup over the full WAM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.