acceptodds
Under review as a conference paper at ICLR 2027

P2G: From Point-wise Matching to Relational Graphs for Training-Free Cross-Modal Image Registration

Abstract

Cross-modal image registration is challenging because modality discrepancies and large geometric transformations weaken the reliability of point-wise matching evidence. Rather than further strengthening individual point descriptors, we observe that intra-image relations among corresponding points remain consistent across modalities, providing complementary evidence of compatibility. Based on this observation, we propose P2G (Points-to-Graph), a task-training-free framework that progressively transforms corresponding point matches into a weighted correspondence relational graph. P2G first estimates a coarse affine transformation and preserves multiple plausible matches under geometric guidance, and then resolves their ambiguity using complementary structural and frozen foundation-model features to reach one-to-one correspondences as weighted graph nodes. Graph edges are constructed from multi-layer features of a frozen foundation model and encode how consistently intra-image query feature relations are preserved across modalities. Node reliability and edge compatibility jointly guide precise correspondence-triplet sampling for robust affine estimation. Experiments across multiple cross-modal datasets and geometric difficulty settings demonstrate that P2G achieves strong registration accuracy and robustness to shifts in scene distribution and geometric transformation than dataset-specific alternatives, offering a new graph perspective for cross-modal feature representation. Our source codes will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.