KIO-World: A Kinematics-Informed Sparse Occupancy World Model
Abstract
Occupancy world models reconstruct 3D scene geometry and semantics from past observations and predict their future evolution. Recent sparse models decode multiple future horizons from query anchors, but do not explicitly couple historical evidence alignment and future prediction through a shared motion state. Consequently, independently moving scene content can cause misaligned image feature sampling, while horizon-specific predictions lack a common motion reference. To address these issues, we present KIO-World, a kinematics-informed sparse occupancy world model that jointly models world states across horizons within a unified kinematic framework. Specifically, we equip each query anchor with a horizon-shared kinematic state that compensates for scene dynamics during historical feature sampling and provides world-state references for future query anchor evolution. To preserve spatial structure during scene evolution, we further introduce a hierarchical residual point refinement module, in which query points iteratively inherit parent geometry, and anchor-wise pose updates are decoupled from point-wise shape refinement. These mechanisms introduce complementary structural priors that support temporal reasoning and geometric refinement while retaining the flexibility of direct occupancy decoding. Extensive experiments demonstrate that KIO-World achieves state-of-the-art performance, outperforming existing camera-centric methods by 11.4% and 6.3% in average semantic mIoU and geometric IoU, respectively. Code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.