acceptodds
Under review as a conference paper at ICLR 2027

OSCAR: A Foundation Model for Optical-SAR Cross-modal Image Registration

Abstract

Optical–SAR cross-modal image registration is a fundamental yet challenging problem in remote sensing. Existing methods often perform well on individual datasets but generalize poorly across sensors, scenes, and geometric transformations. In this work, we show that part of the apparent success on current registration benchmarks can arise from an overlooked geometric shortcut rather than genuine cross-modal correspondence learning. We identify an overlooked shortcut in common affine-warping protocols: zero padding creates black-border patterns that are deterministically related to the transformation, allowing models to infer geometry without learning genuine cross-modal correspondences. To address this issue, we introduce a shortcut-free data construction protocol that removes transformation-dependent padding artifacts while preserving exact dense supervision. Leveraging on this rigorous setting, we propose OSCAR, a generalizable optical-SAR registration framework designed to address the two challenges exposed after shortcut removal: severe cross-modal appearance discrepancy and large global displacement. OSCAR combines foundation-model semantic features with trainable fine-grained representations, performs long-range bidirectional cross-modal interaction, explicitly estimates the dominant affine motion, and recurrently refines the remaining displacement. Jointly trained on seven heterogeneous datasets, our model consistently outperforms all jointly/separately trained baselines and improves zero-shot performance on two unseen datasets, highlighting the importance of eliminating geometric shortcuts and cross-domain trainging of our framework. All codes and datasets will be open-sourced.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.