FlexGlue: Efficient Sparse-to-Dense Image Matching and Visual Disambiguation
Abstract
Estimating robust correspondences across images is a fundamental building block of geometric computer vision. Existing approaches force a trade-off as they match either sparse points, which is highly efficient, or dense pixel grids, which is more robust but expensive. They also often struggle to reject unrelated image pairs that exhibit visual symmetries. We bridge these gaps by introducing a neural network for asymmetric sparse-to-dense image matching and view-graph filtering. Given two images and query points in the first one, we predict whether the image pair is covisible as well as the subpixel coordinates of the corresponding points in the second image. By searching for correspondences over the entire target image, the sparse-to-dense formulation is not limited by the repeatability of keypoint detectors. Furthermore, aggregated context from images provides global understanding necessary for both robust geometric matching and effective visual disambiguation, while jointly learning these tasks leverages their inherent synergy. Experiments on relative pose estimation, Structure-from-Motion, and visual localization show that our method retains subpixel precision and adds robustness to doppelgangers, while achieving computational efficiency of sparse matchers. Together, our method achieves state-of-the art accuracy on IMC2025, InLoc, Doppelgangers, ViSym, ETH3D SfM with COLMAP, and our proposed indoor doppelgangers benchmark. Weights and code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.