acceptodds
Under review as a conference paper at ICLR 2027

VGGT-Occ: Geometry-Guided Reasoning Across Views and Scales for 3D Semantic Occupancy Prediction

Abstract

3D semantic occupancy prediction requires consistent reasoning across camera views, temporal observations, and spatial scales. Yet reference-point projection alone does not carry camera geometry into subsequent sampling, weighting, and cross-camera aggregation. High-resolution decoding must also balance coarse 3D context with the local detail needed for fine-scale prediction. We present VGGT-Occ, a geometry-guided framework for vision-only 3D semantic occupancy prediction across views, time, and scales. Across views and time, Projection-Aware Deformable Attention (PA-DA) carries projection geometry through three stages: it projects learned 3D offsets into each view, applies a Jacobian-based attention bias, and predicts channel-wise fusion weights conditioned on semantic features, projection geometry, and temporal context to aggregate observations across cameras and time. Across scales, a sequential coarse-to-fine decoder concentrates cross-view attention at coarse resolutions and propagates coarse 3D context through gated fusion to guide high-resolution refinement. VGGT-Occ outperforms the compared state-of-the-art methods in semantic mIoU across all three benchmarks. It achieves 37.24% IoU / 25.36% mIoU on SurroundOcc-nuScenes, 44.47% mIoU on Occ3D-nuScenes, and 45.06% IoU / 18.13% mIoU on SSCBench-KITTI-360. Code will be released publicly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.