StreamWAM: Real-Time World-Action Modeling with Closed-Loop Streaming Diffusion
Abstract
World-action models (WAMs) have recently emerged as a promising paradigm for autonomous driving, unifying future video generation and trajectory planning in a single generative framework. However, conventional diffusion-based WAMs require multiple denoising steps to generate each action chunk, leading to high inference latency and limiting the frequency of closed-loop feedback. To address this, we propose StreamWAM, a world-action model capable of real-time driving action planning by closed-loop streaming diffusion. By overlapping the denoising processes of consecutive video-action chunks, the model is able to roll out an action chunk from each inference step, even though the complete diffusion process takes multiple steps. Furthermore, we propose a high-frequency closed-loop planning method to consider the most recent observation for the intermediate action. We employ a latent alignment module to continuously correct the denoised latent states given the new observations. Experiments on the NAVSIM benchmark show that StreamWAM achieves state-of-the-art planning performance with a PDMS of 91.6 and an EPDMS of 94.3. We demonstrate that high-frequency feedback improves planning performance while the model maintains a practical 10Hz planning frequency. We believe StreamWAM represents a concrete step toward deploying world-action models in real autonomous driving systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.