acceptodds
Under review as a conference paper at ICLR 2027

SceneMoE-VPR: Scene-Adaptive Mixture-of-Experts for Visual Place Recognition across Heterogeneous Scenes

Abstract

Visual Place Recognition (VPR) aims to retrieve reference images depicting the same location as a query image. Most existing methods are trained for specific scenes and remain limited in their ability to generalize across heterogeneous scenes. Training VPR models on diverse datasets is a natural route toward robust recognition across heterogeneous environments. However, we observe that extending joint training with indoor data improves indoor recognition but degrades several outdoor benchmarks under fully shared adaptation, revealing a cross-environment trade-off. To address this issue, we propose SceneMoE-VPR, a heterogeneous Mixture-of-Experts framework for input-dependent adaptation. It keeps the DINOv2 backbone frozen and adapts its features through an always-active shared expert and sparsely routed heterogeneous experts designed to model viewpoint-induced geometric variation, local structure, directional layout, and multi-scale context. An image-level router dynamically selects experts for each input without scene labels, enabling input-dependent adaptation across heterogeneous environments. Under the same four-dataset training setting, SceneMoE-VPR consistently outperforms the fully shared baseline on all four benchmarks, improving average R@1 from 88.9% to 89.5%. It further achieves the highest average R@1 among the compared single-image methods with a compact 2,048-dimensional descriptor. These results demonstrate the effectiveness of conditional heterogeneous adaptation for multi-environment VPR across diverse environments. Codes will be publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.