Latent Bundle Adjustment
Abstract
We present a method for dense bundle adjustment of 3D scenes by optimising the latent representations of feed-forward reconstruction models, such as VGGT-Omega and . These learned models can be more robust than traditional Structure-from-Motion (SfM) pipelines, particularly on challenging scenes. However, combining their robustness with the accuracy of classical bundle adjustment remains challenging and is typically limited to poses and sparse points or patches.We propose a novel parametrisation of the bundle adjustment problem by optimising a set of tokens within a feed-forward reconstruction model. Specifically, we solve for latent updates that minimise reprojection error given a set of sparse correspondences between the input frames. Latents jointly encode the dense scene structure and camera poses, and the change in representation provides implicit regularisation during optimisation, allowing for refinement of dense geometry from sparse matches. We evaluate our approach on 3D reconstruction and camera pose estimation tasks across different datasets, demonstrating state-of-the-art performance. We further analyse parametrisations, both qualitatively and quantitatively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.