acceptodds
Under review as a conference paper at ICLR 2027

MatchVLM: Synergizing Vision-Language Models and Local Feature Matching for Cross-View Point Correspondence

Abstract

Cross-View Point Correspondence (CVPC) aims to identify the same physical point across different viewpoints, providing precise spatial guidance for embodied interaction. Despite recent progress, vision-language models (VLMs) can confuse similar object instances under viewpoint changes, revealing the difficulty of inferring fine-grained correspondences from global visual context alone. To address this limitation, we propose MatchVLM, a collaborative framework that integrates feature matching with a compact VLM, leveraging local correspondence cues to support precise cross-view point prediction. We develop this capability through two training stages. Reference-Oriented Spatial Perception Learning jointly supervises reference reliability assessment and answer prediction through supervised fine-tuning. Distance-Aware Point Refinement employs Group Relative Policy Optimization (GRPO) with a point-level distance reward that directly reflects prediction error, providing spatial feedback beyond token-level supervision. A curriculum within this stage emphasizes cases that remain challenging after supervised fine-tuning. We further introduce CVPC-Bench, comprising 500 manually annotated cross-view point pairs across eight outdoor and two indoor scenes. Experiments demonstrate that MatchVLM-3B achieves the highest overall Score on CrossPoint-Bench among the evaluated models and outperforms both CroPond-3B and CroPond-7B on CVPC-Bench under zero-shot evaluation, supporting its effectiveness in cross-view point correspondence and generalization to diverse scenes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.