RVHomo: Unsupervised Cross-Modal Homography Estimation with Geometry-Aware VFM Features
Abstract
Unsupervised cross-modal homography estimation is challenging because modality-dependent appearances obscure shared geometric structures, while common surrogate objectives do not explicitly enforce geometry-preserving behavior. We propose the Recurrent VFM-guided Homography Estimation Network (RVHomo), which combines broad structural cues from a frozen vision foundation model with the spatial precision of local CNN features. Rather than directly injecting coarse VFM tokens, RVHomo transforms hierarchical DINOv2 features into scale-aligned geometric representations and fuses them with CNN features for coarse-to-fine recurrent estimation. A dual-scale reconstruction objective further encourages the adapted representations to retain spatial structure. We additionally enforce geometric consistency at the modality-transfer and homography-prediction levels, imposing warp equivariance and composition-consistent predictions without cross-modal transformation labels. To reduce model complexity, we derive RVHomo-D through two-stage progressive distillation, eliminating modality transfer and online VFM extraction at inference. Experiments on GoogleMap, OPT-SAR, FSD, and VIS-IR demonstrate that RVHomo achieves MACE values of 0.65, 1.95, 4.32, and 4.55, reducing error by 17.4–46.7% relative to the strongest competing estimator on each dataset. RVHomo-D remains more accurate than all evaluated competitors while reducing RVHomo's total parameter count by 98.5% and computation by 70.4%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.