iWAM: Efficient Action Prediction from World Models without Explicit Future Prediction
Abstract
World-action models (WAMs) equip a policy with a pretrained video generation model, on the premise that predicting how a scene should evolve helps decide how to act. While this simplifies action prediction, imagining future scenes is expensive for real-time control. This motivates efficient proxies for future imagination, but it remains unclear what such a proxy should preserve. Our key insight is that an action policy does not require a high-resolution future; it only needs enough information to specify a plausible future with respect to which an action should be produced. This leads to a simple design principle for *i*WAM: rather than denoising an entire future frame, we condition the action policy on intermediate activations from a single video-model evaluation at a latent noise vector that would otherwise initialize denoising. The key challenge is ensuring these activations correspond to the demonstrated future whose actions supervise the policy. We address this by training the video model to approximately represent an invertible flow map: integrating a demonstrated future in the clean-to-noise direction recovers a latent endpoint corresponding to that future, yielding activations aligned with the demonstrated actions. At inference, we sample this endpoint from the prior, and require only a single video-backbone forward pass for conditioning. Across simulated benchmarks and real robots, *i*WAM recovers the generalization benefits of future-conditioning at the inference cost of methods that discard futures entirely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.