Foresight Without Seeing: Latent Futures for World Action Models
Abstract
World Action Models (WAMs) connect visual prediction with robot control, yet supplying predictive context often comes at the cost of expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. Our key insight is that forecasting and rendering can be separated: a policy can access latent predictive computation without reconstructing future observations. Building on this insight, we introduce ForeWAM, a World Action Model that exposes and shapes latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key–value states throughout action denoising. Because predictive context alone does not ensure relevance to control, we further introduce dynamics registers supervised by latent actions from a frozen teacher during training, encouraging representations of interaction-induced transitions. Concentrating predictive computation in this reusable context supports a lightweight, single-layer action decoder. We evaluate ForeWAM on LIBERO, LIBERO-Plus, RoboCasa, and real-world manipulation tasks. Without additional policy-level embodied pretraining, ForeWAM improves RoboCasa success by 9.7 percentage points over Fast-WAM at the same budget of 50 demonstrations per task, reaching 59.2%. With a single-layer decoder, it achieves 77.6% success on LIBERO-Plus and reduces policy-query latency to 88.7 ms on an NVIDIA A800, delivering a speedup over Fast-WAM. These results show that exposing and shaping latent predictive computation provides useful foresight for robust, efficient control without explicit future-video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.