ReMaP: Retrieval-Augmented Multimodal Planning with Dense Robot Memory
Abstract
General-purpose robotic systems must not only perform low-level skills, but also accomplish long-horizon goals. In partially observable settings, like navigation, this requires maintaining a memory or “mental map” that supports efficient, flexible retrieval of task-relevant prior experience. To this end, we present \MethodName, or Retrieval-Augmented Multimodal Planning, a navigation memory and planning framework inspired by retrieval-augmented generation (RAG). \MethodName uses a learned embedding model that captures both semantic and geometric information to retrieve context from a dense memory of prior experience. This context is used to ground a plan of sub-goals, generated by a VLM, for a low-level navigation policy to execute. To assess \MethodName's capabilities for long-horizon navigation, we evaluate \MethodName as part of a full robotic system on real-world robot tasks, beating other navigation baselines by over 20% in success rate and 27 in speed, and on GOAT-bench, where it outperforms prior RGB-based methods by 3% in success rate and 7.8% in SPL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.