acceptodds
Under review as a conference paper at ICLR 2027

SMLA-Nav: End-to-End Learning of Spatial Memory, Localization, and Action for Embodied Navigation

Abstract

Service robots in long-term operational environments are often expected to navigate from arbitrary starting positions to destinations specified by text, images, or coordinates. However, jointly learning scene-specific spatial memory and navigation behavior from sparse goal-directed demonstrations remains challenging. We propose SMLA-Nav, an end-to-end spatial memory navigation framework that complements sparse expert trajectories with readily obtainable image–pose observations to acquire scene knowledge and adapt pretrained navigation skills. A consistent-prompt multi-stage training strategy bridges navigation pretraining and scene-specific post-training through a common spatial representation. This connects scene-specific localization and goal grounding with transferable point-goal navigation, composing these capabilities into a hierarchical spatial reasoning chain within a shared multimodal large language model (MLLM). At inference, the policy uses implicit spatial memory encoded in its parameters and deterministic geometric computation to navigate without querying external maps or localization databases. We establish SMLA-Bench, comprising SMLA-HM3D and SMLA-Gauss, on which SMLA-Nav achieves success rates of 86.9% and 85.7%, respectively. Comparisons using the same sparse scene trajectories further show improved text-goal navigation over direct MLLM post-training, while deployment on a quadruped robot demonstrates real-world feasibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.