Graph3R: Unifying View Graph Initialization with Feed-Forward Reconstruction
Abstract
Feed-forward reconstruction recovers camera poses and dense geometry from image sets whose views are assumed to come from a single scene. In practice, uncurated image collections often contain views from disjoint scenes, and reconstructing these views jointly can corrupt camera poses and geometry. Existing pipelines typically resolve view associations in separate stages before final reconstruction, leaving view graph initialization and reconstruction inference only loosely coupled. We introduce *Graph3R*, which ***unifies view graph initialization with feed-forward reconstruction through a shared geometric representation***. A lightweight verification head predicts pairwise covisibility from the multi-view features of to construct a view graph. This graph classifies the input views into scene groups. For each group, we mask the shared features before decoding camera poses and dense geometry. Reusing the multi-view features allows *Graph3R* to reconstruct disjoint scenes separately with a single backbone pass. Across three benchmarks, *Graph3R* improves average maximum recall at 100% precision by upto 8.99 percentage points over the baseline, while improving reconstruction quality on mixed-scene inputs. Regarding the multi-session mapping task, *Graph3R* improves cross-session localization and achieves competitive pose estimation accuracy, while delivering a runtime speedup of approximately over the baseline when reconstructing the query group. By integrating learned view association with geometric inference, our approach takes a step toward end-to-end reconstruction from uncurated image collections.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.