Supervising 3D Occupancy Where the Sensors Have Not Looked Yet
Abstract
Motion changes not only a driving scene, but also which parts of it can be physically verified. We turn this change in observability into supervision for 3D occupancy and a geometric verifier for driving world models. Later LiDAR sweeps often reveal static space that was hidden at an earlier query frame. Rather than collapsing these measurements into denser temporal evidence, we retain when each voxel becomes observable. This yields a query-time partition: space resolved immediately, space resolved only by later measurements (), and space never resolved in the log. is delayed physical evidence of the hidden present—not an additional model input or a future state to predict. Extending only the loss region of a matched occupancy learner from to raises F1 by on Occ3D-nuScenes; an independent five-seed retraining gives with a scene-clustered 95% confidence interval of . Standard occupancy metrics do not decline, and the gain transfers to Waymo, multiple encoders and domains, and a second lifting architecture. The same delayed evidence also gives generated futures an external geometric reference. We decode each video produced by a driving world model into 3D occupancy and check its implied static geometry at the commanded pose against a sensor-derived map. An occupancy reader trained on roughly doubles the measured separation between rollouts conditioned on logged and counterfactual actions. When choosing among four stochastic rollouts of the same counterfactual action, the resulting verifier yields about 44 more correctly located occupied voxels per frame than random selection, evaluated in a disjoint region by an independently trained reader. This sensor-grounded score is a candidate geometric reward for future world-model post-training; we do not fine-tune the generator or claim a closed-loop benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.