InWorld: An Instant-Interactive Multimodal World Model for Autonomous Driving
Abstract
World models have become a scalable alternative to real-world interaction for policy learning and evaluation in autonomous driving. A high-performance driving world model must satisfy three core requirements simultaneously: cross-modal consistency, real-time generation, and pose-controllable interaction. However, existing methods struggle to fulfill all these criteria, frequently requiring trade-offs among them. We present InWorld, an interactive driving world model that synchronously generates cross-modally consistent images and LiDAR point clouds in real time. To faithfully capture intrinsic LiDAR sensor characteristics, we design a dedicated LiDAR VAE optimized for range image representation, which integrates horizontal circular convolution and a differentiable ray-drop module. For cross-modally consistent generation, we present a layout architecture equipped with epipolar masked attention, alongside image-LiDAR consistent Pl\"ucker coordinates encoding. To enable controllable interactive modeling, we propose temporal decay KV caching and embed it into the teacher-student pipeline via distribution matching distillation. Extensive experiments on the nuScenes dataset demonstrate that InWorld achieves state-of-the-art performance in both generation efficiency and cross-modal alignment, while the student enables interactive generation at a 10 Hz streaming inference rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.