LaRI3R: Bidirectional Ray Intersections for Occluded 3D Scene Reconstruction from Two Views
Abstract
Recovering complete 3D scene geometry from limited observations remains challenging due to severe occlusions and restricted viewpoints. Although observing a scene from many viewpoints can progressively improve reconstruction completeness, dense multi-view reconstruction often requires costly exploration and computation. Recent single-view feedforward methods address the trade-off between reconstruction quality and efficiency through layered representations, but they still suffer from strong ambiguity in heavily occluded regions. Extending layered reconstruction to two views is also non-trivial because occlusion ordering becomes view-dependent, and visually corresponding evidence can be weak or absent across views. We present LaRI3R, a new feedforward two-view framework for complete 3D scene reconstruction under severe occlusions. Our approach is built on two core designs: 1) We introduce a bidirectional layered representation that predicts geometry in both forward and backward ray orders, which reduces geometric inconsistencies in heavily occluded regions. 2) Instead of relying purely on appearance-based matching, we perform geometry-aware feature aggregation using explicit 3D positional cues. This design enables geometrically related but visually dissimilar regions to communicate across views. Experiments on both object- and scene-level benchmarks show that LaRI3R outperforms representative generative, iterative, and feedforward methods while preserving efficient inference. We further show that the recovered occluded geometry serves as a strong structural cue for downstream 3D vision tasks such as camera pose estimation with notable improvements over existing large geometric models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.