acceptodds
Under review as a conference paper at ICLR 2027

How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization

Abstract

Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image of its surroundings. The predominant approach to this task maps features from ground-view images and satellite references into a shared bird's-eye-view (BEV) space and then establishes spatial correspondences between them. Although effective, this approach still faces ambiguity in feature placement and descriptor matching. When mapping ground-view features into BEV space, insufficient depth constraints allow the same feature to be assigned to different distances along the viewing direction, creating geometric ambiguity in BEV feature placement. Meanwhile, similar appearances at different locations create descriptor matching ambiguity, and existing descriptor learning lacks explicit semantic supervision to distinguish these locations. To reduce these ambiguities, we propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. First, we geometrically constrain ground-view BEV feature placement through radial depth supervision for distance assignment and vertical height supervision for height aggregation. Then, shared explicit semantic supervision further promotes consistent semantic predictions across views and helps distinguish different locations with similar semantics. These constraints jointly improve feature placement and descriptor discriminability, thus improving the satellite-ground localization performance of state-of-the-art models by a large margin. Qualitative and quantitative results demonstrate the effectiveness of GeoSem-BEV in enhancing these models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error relative to the corresponding state-of-the-art method by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors were reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, GeoSem-BEV reduces the same-area mean orientation error by 26.8% relative to the corresponding baseline under orientation noise.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.