On Self-Improving Feed-Forward SfM: A Generalizable Multi-View Pose Estimator
Abstract
Classical Structure-from-Motion (SfM) is an optimization method: it jointly recovers camera poses, verified image correspondences, and depth (i.e., a point cloud) from an image collection without labels. Modern feed-forward SfM models instead regress these quantities in a single pass. Yet they must first be trained on labeled data, part of which is itself produced by classical SfM pipelines. Like other learned models, they fail on inputs beyond the training distribution. We close this gap by making feed-forward SfM produce its own labels: an optional bundle adjustment (BA) stage refines its predictions, and the refined output fine-tunes the network, forming a self-improving loop. We realize this loop for camera intrinsic and extrinsic estimation, formulating SfM as three foundation models: a frozen depth and correspondence estimator, and a tunable feed-forward pose estimator (the SfM pose estimator). We propose a global BA that depends entirely on model outputs: initialized with the feed-forward camera estimates, it optimizes for projective consistency between the predicted depth and correspondence maps, then extracts a subset of accurately registered images whose refined poses serve as pseudo ground truth. The pipeline scales end-to-end: from feed-forward initialization to BA, and from small scenes to 15,000-image unconstrained internet collections. Continual retraining on this self-generated supervision adapts the SfM pose estimator to the new domain. Across small-scene (IMC2021, ETH3D), large-scene (FastMap, T&T), and unconstrained internet (GLDv2, 1DSfM) datasets, our BA-refined poses approach the accuracy of SoTA SfM, making them reliable pseudo labels: fine-tuning on them substantially improves the feed-forward pose estimator. These results point to an emerging self-improving learning paradigm for SfM that generalizes across unseen domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.