SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving
Abstract
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. Recent unified world–action models further demonstrate strong planning performance and promising zero-shot transfer capability, highlighting the potential of pretrained video-generation priors for driving. However, generating future videos during inference introduces substantial computational overhead, and many recent driving WMs therefore favor single-front-view inference for efficiency, which limits spatial coverage in safety-critical scenarios. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action–video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves strong planning performance with low inference latency and competitive zero-shot transfer capability. Code: https://anonymous.4open.science/r/SV-WAM-F74E/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.