Rethinking Cross-View Geo-Localization: A Probabilistic Motivation
Abstract
Cross-view geo-localization (CVGL) matches a street-view query, typically a panorama, to a geo-tagged satellite reference image of the same location. Existing methods primarily address the severe viewpoint gap between the two modalities, but pay much less attention to appearance variation caused by weather and daylight. We model a CVGL image as , where is an image formation function and denotes location-preserving content, viewpoint, appearance, and residual nuisance. This formulation clarifies the learning objective: preserve information about while reducing sensitivity to the remaining factors. This probabilistic view motivates ProbInfoNCE, which replaces the shared temperature in InfoNCE with a learned pair-dependent scale to account for sample-dependent ambiguity during training. We also introduce CAV-Geo, a large-scale global dataset with appearance variation, with 256K street-view locations and 440K satellite-view images. CAV-Geo further includes a cross-continental split designed to evaluate out-of-distribution generalization. It is the largest CVGL benchmark with explicit weather and daylight variation, and its cross-continental protocol provides a substantially more challenging evaluation setting than existing standard splits. With the same DINOv3 initialization, ProbInfoNCE improves R@1 over InfoNCE by on VIGOR cross-area and on CVACTCVUSA. Appearance augmentation with dual-matrix training and query-to-query alignment further improves R@1 on original VIGOR cross-area queries by relative to ProbInfoNCE trained without appearance augmentation. All percentage gains in this paper are computed as ratios to their stated baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.