acceptodds
Under review as a conference paper at ICLR 2027

AnchorFuse: Robust Anchor-Centered Temporal Fusion for Multimodal Aerial-Ground Place Recognition

Abstract

Multimodal aerial-ground place recognition localizes a ground platform by matching camera–LiDAR observations to georeferenced aerial imagery. However, descriptor-level temporal aggregation does not explicitly distinguish evidence relevant to the queried location from observations of neighboring locations, and retrieval can degrade when context is unavailable. We propose AnchorFuse, an anchor-centered framework that explicitly separates the retrieval target from its temporal context. An anchor observation defines both the positive aerial references and the coordinate frame for fusion, keeping the retrieval target fixed as auxiliary observations change. Building on this formulation, Anchor-Aligned Cross-Frame Fusion (AACF) transforms local camera–LiDAR features into the anchor frame and selects nearby evidence according to spatial proximity, temporal distance, and learned reliability. A gated residual integrates the selected evidence into the query descriptor. To accommodate missing context, Robust Anchor-Memory Consistency (RAMC) jointly supervises clean and corrupted queries against the same retrieval target and aligns corrupted descriptors with a stop-gradient clean target. Context corruption preserves the multimodal anchor, and RAMC adds no inference-time branch. AnchorFuse achieves R@1 of 44.28% on KITTI360-AG and 78.35% on nuScenes-AG. On KITTI360-AG, RAMC improves R@1 over AACF by 17.03 percentage points when two non-anchor frames are removed, compared with 1.49 points on clean queries.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.