AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
Abstract
World-action models (WAMs) couple predictive visual modeling with action generation, but typically rely on iterative denoising with a fixed inference budget, incurring substantial latency while overlooking variation in generation difficulty across interaction states. We introduce AnyStep World Action Model, a general framework for accurate cross-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk–benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student–teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07, 12.08, and 8.94 percentage points on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.