acceptodds
Under review as a conference paper at ICLR 2027

MomentMem: Memorizing Key Moments for Robot Manipulation

Abstract

Many manipulation tasks require recalling visual evidence that has left the robot's view. Yet most vision-language-action (VLA) policies see only the current observation or a short history. Such tasks call for visual moment memory, which captures the scene at the moments that matter. To evaluate this capability, we introduce VMBench, whose tasks require recalling patterns, locations, or both. We then propose MomentMem, which instantiates this memory with keyframes. A VLM selects keyframes, and a VLA executor predicts actions from them and the current observation. Both are trained with segment-level annotations and representative-frame sampling. MomentMem achieves state-of-the-art results on VMBench and RMBench, and improves success on a real robot. Its memory also transfers zero-shot to unseen visual appearances and adapts to a held-out task with few annotations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.