Learning When to Recall and Where to Search for Lifelong Multi-Modal Navigation
Abstract
Lifelong multi-modal navigation requires reusing past observations under changing goals. Yet storing observations selectively or in compressed form can omit details that later requests need, and retrieved evidence must still be interpreted to decide whether to revisit a known location or explore further. We introduce RecallNav, which separates goal-agnostic observation storage from goal-conditioned memory access and learns when to recall and where to search. It retains original observations across goals and uses a frozen multimodal encoder to retrieve a small set of candidate frames. A multimodal large language model (MLLM) reads retrieved, current, and frontier views within a bounded policy context and learns from expert decisions to ground targets or select destinations for further search. Explicit recall supervision ties return decisions to supporting retrieved views. Even with a 3B MLLM backbone, RecallNav outperforms all compared baselines in navigation success on every GOAT-Bench validation split. It also surpasses the reported GPT-4o-based snapshot baselines in both success rate and path efficiency on Val-Unseen. On LMEE-Bench, the same 3B policy exceeds the 7B memory-based systems in both navigation success and path efficiency. The same RGB archives support later question answering through a separate pretrained MLLM without QA-specific fine-tuning. Ablations show that goal-relevant memory access and explicit recall supervision each contribute to navigation performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.