Is It a Good Match? 2D Keypoint Matching with Feedback of a VLM
Abstract
Two-dimensional (2D) keypoint matching is a fundamental problem in computer vision, used in applications such as image registration, 3D reconstruction, and visual localization. Recent Vision Foundation Models, such as DINOv3, have demonstrated strong transferable representations across a range of vision tasks, including keypoint correspondence estimation. However, their performance remains limited by imperfect feature distinctiveness and weak geometric awareness. Their practical usage often relies on task-specific fine-tuning with carefully engineered losses. Vision-Language Models (VLMs) have shown strong capabilities in evaluating and reasoning on multiple vision tasks, including assessing the quality of vision model predictions. However, their high inference cost makes them impractical for direct use in matching pipelines. In this work, we leverage the evaluative strength of VLMs to guide the fine-tuning of vision foundation models for 2D keypoint matching. Specifically, we construct a feedback signal based on a VLM-based evaluation score (VLM-as-Judge) of predicted correspondences. We introduce a feedback-driven training framework using the VLM score to fine-tune a DINOv3 backbone for keypoint matching. By integrating this reward-based feedback loop, our method improves correspondence quality without requiring explicit ground-truth supervision or loss engineering. At test time, we rely only on the fine-tuned model, running 50-160x faster than a single VLM call.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.