acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Frustum: Single-View Amodal 3D Scene Reconstruction

Abstract

Single-image 3D methods either reconstruct the visible surface or generate a plausible room, leaving a gap for metric completion tied to the input. We study evidence-grounded amodal occupancy completion: recover the scan-observed geometry of every object supported by a single RGB(-D) view, including occluded and out-of-frustum parts, while predicting no unsupported objects. We present \method, a feed-forward framework that back-projects depth into a sparse, DINO-decorated 3D seed and completes it with a 3D U-Net inside a predicted scene-fitted box. A scalar coverage channel, aggressive seed dropout, and a completion-masked BCE–Dice loss are designed to discourage trivial seed copying and to focus learning on unseen voxels. On 3RScan, \method substantially outperforms reconstructive, generative, and indoor-occupancy baselines on unseen geometry; explicit in-frustum/out-of-frustum evaluation and a volume-matched dilation control confirm that the gain reflects completion rather than surface inflation. The model also transfers zero-shot to ScanNet. Code and trained models will be made public.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.