Situating Memory for Embodied Navigation
Abstract
Memory-based navigation requires an agent to relate stored environmental knowledge to its current observation. When navigation fails, however, task success alone cannot distinguish missing knowledge from errors in using it. We introduce SCOPE, a controlled evaluation framework for studying this distinction across 128 navigation tasks in 64 indoor scenes. We vary frozen scene memory and online episode memory while holding the navigation policy and execution pipeline fixed, and then replace native scene memory with audited ground-truth scene knowledge. The strongest ground-truth condition succeeds on 93 of 128 tasks. Among its 35 residual failures, trace-based audits attribute 24 to current localization, goal localization, or egocentric correspondence. One-shot corrections at diagnosed checkpoints recover 15 episodes, compared with five under matched sham interventions, providing behavioral evidence that the identified errors are consequential. Motivated by these findings, we introduce Experience Skeleton Memory (ESM), which uses simulator-provided poses to organize executed trajectories and express historical observations relative to the current view. With the same ground-truth scene memory, ESM achieves 85.94% success on the full suite, compared with 72.66% for the original structured episode memory. Together, these results identify situated correspondence between memory and current observations as a key navigation bottleneck and show that explicitly maintaining this correspondence can substantially improve closed-loop navigation under accurate pose information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.