One-Shot Relational Adaptation of Vision Foundation Models for 3D Correspondence
Abstract
Vision foundation models (VFMs) have demonstrated generalization across diverse 2D vision tasks, yet adapting them for 3D correspondence tasks, such as multi-view matching and pose estimation, remains challenging due to the lack of large-scale 3D correspondence annotations. Existing adaptation strategies typically rely on extensive training data or task-specific supervision, limiting their applicability in data-scarce scenarios. In this work, we propose a one-shot adaptation framework that enhances the 3D correspondence capability of VFMs using only a single overlapping image pair with cross-view correspondences. Our key insight is that a correspondence-annotated image pair provides not only point-level supervision but also rich relational structures among matched points. To exploit this property, we introduce a two-stage adaptation framework. First, an attention-gated convolutional adapter is designed to improve local cross-view feature consistency while introducing minimal modification to the pretrained VFM. Second, a relational adaptation module explicitly models structural relationships among corresponding points and enforces cross-view relational consistency, enabling the adapted representation to preserve relational structures across viewpoints. Experiments on multiple 3D correspondence benchmarks show that our approach generally improves pretrained VFMs and outperforms recent VFM adaptation methods in many settings, while requiring only a single image pair for adaptation. These results highlight the effectiveness of exploiting intrinsic correspondence and relational information for efficient 3D-aware VFM adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.