Metis: A Generalizable and Efficient 4D World-Action Model for Autonomous Driving and Urban Navigation
Abstract
World-Action Models (WAMs) jointly model future observations and actions, but existing approaches often require explicit future observation generation during action inference, introducing substantial inference overhead. Moreover, future states are primarily modeled in semantic or 2D visual spaces, leaving their spatial structure insufficiently grounded for reliable planning. We introduce Metis, an end-to-end WAM that jointly learns future observation, action, and spatial geometry prediction during training, while requiring only action prediction at inference time. Metis employs specialized video generation and action experts coupled through asymmetric attention, allowing future video representations to attend to action representations while masking the reverse dependency. This design preserves supervision from future world modeling without requiring stochastic future video generation for action inference. We further introduce training-only spatial registers with complementary depth and occupancy supervision to enhance spatially grounded action representations without additional inference overhead. Extensive experiments demonstrate strong performance on NAVSIM-v2 and CityWalker, together with substantially reduced action-inference latency compared with full world-action generation. Zero-shot real-world deployment further confirms the practical feasibility of our approach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.