R²: Residual-Guided Attention and Geometric Point Refinement for 3D Reconstruction
Abstract
Recent advances in geometric reconstruction have largely followed from feed-forward approaches that jointly estimate cameras, depth, and 3D point maps from images. Such models output these variables through a loosely parallel architecture despite their strong geometric dependencies. We propose R, an approach based on residual-guided camera attention and camera- and depth-conditioned point refinement. Starting from a pretrained VGGT representation, a frozen teacher first produces provisional camera parameters and geometry. We then construct a multi-view rigid-consistency residual and use it as a signed geometric prior for camera-query attention. The final camera and depth predictions are combined with the initial point map in a camera- and depth-conditioned refinement stage, which we refer to as geometric point refinement. Experiments show improvements in selected camera and depth metrics, while the proposed point refinement reduces median point error by 12.57% in a matched ablation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.