acceptodds
Under review as a conference paper at ICLR 2027

CVGL-CueBench: A Benchmark for Visual-Cue Dependence and Representation Transfer in Cross-View Geo-Localization

Abstract

Cross-view geo-localization (CVGL) models achieve high retrieval recall on benchmark RGB ground queries. However, standard RGB images intrinsically couple surface appearance with spatial geometric layouts. Evaluating RGB inputs alone obscures which visual cues actually drive alignment, as well as how sensitively models rely on each cue. To systematically diagnose visual cue reliance and model sensitivity, we introduce CVGL-CueBench, a multi-representation diagnostic benchmark comprising two sub-benchmarks, CVUSA-Cue and CVACT-Cue, derived from CVUSA and CVACT. While retaining ground–aerial spatial correspondences and the invariant RGB aerial gallery, CVGL-CueBench provides five parallel query representations: the original RGB image alongside four derived structural representations (class-agnostic region segmentation, estimated relative depth, edge maps, and sketches). We establish a systematic diagnostic protocol evaluating five representative CVGL methods across representation-matched training () to measure localization capacity preserved within each visual cue, and cross-representation transfer () to assess sensitivity to representation shifts, complemented by test-time spatial perturbations on an RGB baseline to examine reliance on high-frequency textures, spatial layout, and regional context. Evaluations demonstrate that derived structural representations retain sufficient geometric cues to support high retrieval recall when paired with capable backbones: on north-aligned CVUSA-Cue, Sample4Geo achieves 90.13%–98.68% Recall@1 across all query representations. However, cross-representation transfer () reveals pronounced sensitivity, varying from 68.12% on sketches to a near-total collapse of 0.06% on relative depth queries. Furthermore, under a field of view on CVACT-Cue, models trained on sketch queries surpass their RGB counterparts across all five evaluated methods by 0.96–15.48 percentage points, though this ranking advantage varies across methods on CVUSA-Cue. These findings demonstrate that high RGB recall masks acute sensitivity to representation shifts, while confirming that structural representations retain sufficient geometric cues to support cross-view alignment, particularly under restricted viewing constraints. CVGL-CueBench thus provides a controlled diagnostic foundation to decouple visual cues and evaluate structural robustness in cross-view geo-localization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.