Beyond Invariance: Learning the Geographic Core for Generalizable Cross-View Geo-Localization
Abstract
Cross-view geo-localization (CVGL) identifies locations by matching images across different viewpoints. Existing methods prioritize view-invariant representations, but invariant semantics may exhibit high visual similarity across different locations, limiting geographic discrimination under domain shifts. Although scene semantics can narrow plausible matches, they cannot fully resolve ambiguous places, motivating complementary evidence from fine-grained regional structure. Inspired by sufficient invariant learning, we propose GeoCore to learn a geographic core of complementary semantic and structural evidence. The semantic branch can capture view-invariant scene semantics for coarse discrimination. To resolve the spatial ambiguity, the structural branch learns fine-grained regional structure that remains discriminative across environmental variations. Extensive experiments on UAV and ground-view benchmarks demonstrate that our model consistently outperforms existing approaches, achieving superior performance in cross-area generalization and robustness evaluations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.