MVLoMa: Multi-View Local Feature Matching with Any Geometry Foundation Model
Abstract
Local feature matching is typically formulated for image pairs, although structure-from-motion and visual localization commonly match one reference image against many retrieved or covisible neighbours. Applying a pairwise matcher independently to these centre–neighbour edges duplicates reference-side computation and prevents the matcher from exploiting complementary multi-view geometric evidence from the remaining observations. To address these limitations, we introduce MVLoMa, a sparse star-graph matcher that extends LoMa from image pairs to multiple views. MVLoMa first uses any frozen multi-view geometry model to jointly describe all images in a star, then injects the resulting cross-view geometric context into local descriptors. It then matches the centre image to all neighbours in a single forward pass through per-view self-attention and bidirectional centre–neighbour cross-attention. For supervision, we derive positive correspondences only from depth-valid, mutually visible dense warps and explicitly penalize high-confidence matches that violate epipolar geometry. On 25-view star graphs from HyperSim, ScanNet++, Mip-NeRF 360, and Tanks and Temples, MVLoMa achieves state-of-the-art average relative-pose AUC@ across the four benchmarks. At 25 views, its matcher is faster than sequential evaluation of the 24 centre–neighbour edges. Furthermore, MVLoMa can be instantiated with any frozen geometry foundation model—including DA3, , VGGT, and VGGT-—and all variants achieve strong performance. Controlled backbone, supervision, and view-count ablations demonstrate the benefits of geometry-aware descriptors, visibility-aware supervision, geometric negative supervision, and joint multi-view inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.