From Visibility to Geometry: Efficient Training for Feed-Forward 3D Reconstruction
Abstract
Feed-forward reconstruction models, such as VGGT- and , have demonstrated impressive reconstruction performance, often leveraging powerful pretrained visual features such as DINOv2 or DINOv3. In this work, we show that restructuring the training process around geometric relationships can substantially improve the efficiency of learning multi-view geometry. Our key insight is that predicting covisibility for multiple pairs of images provides a strong, unified learning signal for geometric relationships across views. Instead of directly predicting camera poses and depth maps, we first pretrain the network to predict covisibility maps derived from the ground-truth geometry, and then fine-tune it to regress camera poses and depth maps. This two-stage training strategy substantially accelerates convergence, yielding significantly better reconstruction performance under a fixed compute budget. Remarkably, our approach achieves competitive reconstruction quality when trained entirely from scratch, without relying on any pretrained visual features. Code and models will be made available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.