Instance-Centric Unsupervised Monocular 3D Occupancy Prediction
Abstract
Unsupervised monocular 3D occupancy prediction aims to recover dense scene geometry from a single image without 3D dense occupancy annotations, but often produces blurred object boundaries and trailing artifacts under sparse and indirect image-based supervision. We attribute these limitations to the prevailing scene-level representation, which learns unified geometry without explicitly distinguishing individual object geometries, thereby overlooking the 3D shape modeling of individual objects. To address this issue, we propose ICOcc, an instance-centric unsupervised occupancy prediction framework, which reformulates occupancy learning from scene-level reconstruction to instance-centric prediction. Specifically, an instance-centric categorical occupancy formulation directly separates the entangled object geometries, without the need for a unified scene geometry. Based on this formulation, an instance decoupling network is introduced, where learnable queries are employed to capture shape priors and decouple object geometries from scene-level visual features. To supervise the decoupled outputs, an instance-anchored learning strategy further provides stronger constraints through multi-view supervision across dynamically selected views. Extensive experiments on KITTI-360 and CarlaOcc show that ICOcc improves occupancy prediction, achieving a 10.2% relative gain in invisible-empty recall on KITTI-360, while producing clearer object boundaries and fewer trailing artifacts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.