ReWAM: A Unified Framework for Representation-Centric Analysis of World-Action Modeling
Abstract
World-action models (WAMs) learn future scene dynamics alongside actions for autonomous driving, yet how world representations affect planning remains unclear. We propose ReWAM, a unified framework for studying how world representations shape planning. Under a common architecture, ReWAM instantiates inverse dynamics model, world model, and world-action model across nine representations spanning VAE latents, visual foundation features, and driving-specific BEV features, while making even implicit foundation latents visually inspectable as videos. We evaluate these models progressively through action recovery, appearance-agnostic world-action consistency, and planning, revealing aligned representation trends and an advantage for visual foundation representations. Within the same framework, we further compare world-action interaction paradigms and introduce a new method that directs world-prediction supervision to action generation during training while retaining action-only inference. ReWAM achieves strong planning performance, reaching 91.6 PDMS on navtest and 39.3 EPDMS on navhard. Its efficient variant retains 91.6 PDMS on navtest at 59 ms, 3.6× faster than inference with future-world generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.