acceptodds
Under review as a conference paper at ICLR 2027

iWAM: Efficient Action Prediction from World Models without Explicit Future Prediction

Abstract

World-action models (WAMs) equip a policy with a pretrained video generation model, on the premise that predicting how a scene should evolve helps decide how to act. While this simplifies action prediction, imagining future scenes is expensive for real-time control. This motivates efficient proxies for future imagination, but it remains unclear what such a proxy should preserve. Our key insight is that an action policy does not require a high-resolution future; it only needs enough information to specify a plausible future with respect to which an action should be produced. This leads to a simple design principle for *i*WAM: rather than denoising an entire future frame, we condition the action policy on intermediate activations from a single video-model evaluation at a latent noise vector that would otherwise initialize denoising. The key challenge is ensuring these activations correspond to the demonstrated future whose actions supervise the policy. We address this by training the video model to approximately represent an invertible flow map: integrating a demonstrated future in the clean-to-noise direction recovers a latent endpoint corresponding to that future, yielding activations aligned with the demonstrated actions. At inference, we sample this endpoint from the prior, and require only a single video-backbone forward pass for conditioning. Across simulated benchmarks and real robots, *i*WAM recovers the generalization benefits of future-conditioning at the inference cost of methods that discard futures entirely.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.