SwiftWM: An Event-Guided World Model for Fast Motion Planning in Autonomous Driving
Abstract
Timely adaptation to changing traffic is essential for autonomous driving. Dynamic vision sensors (DVS) capture sparse, asynchronous brightness changes with fine temporal resolution. Frame aggregation can obscure event timing and weaken their low-latency advantage. Iterative diffusion planning adds computational cost to trajectory updates. We propose SwiftWM, an event-guided latent world model for fast motion planning in autonomous driving. SwiftWM encodes dynamic spatiotemporal event graphs with a spiking graph neural network. Geometry-guided alignment and gated fusion integrate event features with RGB–LiDAR context and radar observations into a persistent bird's-eye-view latent state. DVS packets incrementally update affected regions while retaining the state elsewhere. A predictor forecasts future latents from the updated state. At each planning update, a diffusion planner conditions on current and predicted future latents to generate a new path with a single solver step, while a separate head predicts target speed. Predictive and planning supervision jointly optimize the world model and planner over causal sequences. We evaluate SwiftWM through closed-loop driving on Bench2Drive and recorded-input diagnostics of future-state prediction and computational cost. SwiftWM achieves a driving score of 89.82 and a success rate of 76.36% on Bench2Drive, while delivering a 76.38% reduction in state-maintenance computation through incremental execution compared with full recomputation in recorded-input evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.