acceptodds
Under review as a conference paper at ICLR 2027

InWorld: An Instant-Interactive Multimodal World Model for Autonomous Driving

Abstract

World models have become a scalable alternative to real-world interaction for policy learning and evaluation in autonomous driving. A high-performance driving world model must satisfy three core requirements simultaneously: cross-modal consistency, real-time generation, and pose-controllable interaction. However, existing methods struggle to fulfill all these criteria, frequently requiring trade-offs among them. We present InWorld, an interactive driving world model that synchronously generates cross-modally consistent images and LiDAR point clouds in real time. To faithfully capture intrinsic LiDAR sensor characteristics, we design a dedicated LiDAR VAE optimized for range image representation, which integrates horizontal circular convolution and a differentiable ray-drop module. For cross-modally consistent generation, we present a layout architecture equipped with epipolar masked attention, alongside image-LiDAR consistent Pl\"ucker coordinates encoding. To enable controllable interactive modeling, we propose temporal decay KV caching and embed it into the teacher-student pipeline via distribution matching distillation. Extensive experiments on the nuScenes dataset demonstrate that InWorld achieves state-of-the-art performance in both generation efficiency and cross-modal alignment, while the student enables interactive generation at a 10 Hz streaming inference rate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.