GeoSync: Unified Representation Learning for End-to-End Cross-View Geolocalization
Abstract
We propose GeoSync, a unified framework for cross-view image geolocalization that jointly learns satellite image retrieval and fine-grained localization from a shared cross-view representation. Existing approaches typically separate these tasks, using global image-level matching to retrieve a satellite reference and a dedicated localizer to estimate the position within it. This decoupling limits interaction between retrieval and localization while adding computation during inference and increasing system cost. GeoSync learns a shared representation that captures global scene understanding for retrieval and local feature correlations for precise localization. We further introduce consistency supervision across positive and semi-positive satellite references and a sub-meter refinement module to improve accuracy and robustness. Localization-aware re-ranking then uses localization evidence to refine retrieved candidates, allowing localization to reinforce retrieval. Experiments on VIGOR and DReSS demonstrate strong performance in both retrieval and fine-grained localization, yielding accurate and fast end-to-end geolocalization from a single shared representation and maintaining robustness as the target moves toward the boundary of the satellite reference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.