EMAG: Event Memory Retrieval and Spatial Affordance Grounding for Non-Markovian Robot Manipulation
Abstract
Many real-world manipulation tasks are non-Markovian: the target of the current action may be determined by an event observed much earlier and therefore cannot be inferred from the current observation alone. Existing memory-augmented policies compress past observations into compact visual representations or textual summaries that track task progress, but do not jointly retain explicit event semantics and the visual evidence needed to re-ground a history-dependent target in the current scene. To address this issue, we introduce EMAG, which uses memory for semantic retrieval to identify which past events are relevant and for visual grounding to determine where the corresponding target is. Specifically, EMAG stores memory as captioned keyframes, pairing RGB snapshots of salient events with concise textual descriptions. Each caption serves as a semantic handle for memory retrieval, while the corresponding keyframe preserves the visual evidence needed for grounding. Given the task goal, progress summary, and recent observations, EMAG queries the stored captions to retrieve relevant memory entries, re-grounds the corresponding target in the current observation, and predicts an affordance map that specifies the target region to act on. The predicted affordance map modulates the visual features of a low-level vision-language-action (VLA) policy to guide action generation. Across all 16 RoboMME tasks, EMAG achieves an average success rate of 73.4%, exceeding the strongest official RoboMME baseline by 28.9 percentage points. On three real-robot tasks requiring object counting, object permanence, and sequence imitation, EMAG achieves an average success rate of 77.2%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.