RayLift: Beyond Rank-One Lifting for Occlusion-Aware 3D Occupancy Prediction
Abstract
Camera-based 3D semantic occupancy prediction must reconstruct not only visible surfaces, but also geometry and semantics hidden behind them. Consider a pixel ray that first intersects a visible car and then an occluded bus (Fig. 1). Although the image feature at that pixel is dominated by the car, the resulting 3D representation should encode different semantics at the two depths. Standard Lift-Splat-Shoot (LSS) does not provide this distinction during lifting: it copies the same pixel feature to every depth bin and changes only its scalar weight. Consequently, the per-ray lifted representation has rank at most one. When the same depth-independent linear classifier is applied at every depth, all positive-mass depths preserve the same feature-dependent class ordering, so any class-order reversal must arise from subsequent 3D reasoning. We formalize this representational restriction as a rank-collapse obstruction. We introduce **RayLift**, which directly relaxes this restriction and complements it with directed contextual reasoning for occluded regions. Optimal-transport-routed multi-feature lifting assigns each depth bin a depth-dependent mixture of feature heads, enabling different, non-collinear feature vectors along the same ray. Causal ray-radial convolution provides a directed near-to-far pathway for propagating contextual information toward farther, occluded voxels. Visibility-conditioned temporal fusion extends this reasoning across time by conditioning historical voxel features on their predicted visibility. On Occ3D-nuScenes, RayLift achieves and , surpassing ALOcc by points in overall and points in . On Occ3D-Waymo, RayLift achieves , outperforming ALOcc by points under matched camera-only settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.