acceptodds
Under review as a conference paper at ICLR 2027

Reasoning to Remember: Structured Keyframe Memory for Long-Horizon Robotic Manipulation

Abstract

Long-horizon robotic manipulation requires historical information to guide decisions beyond what current observations reveal. Keyframe-based memory provides visual context for a high-level vision-language model (VLM), which predicts subgoals for a low-level vision-language-action (VLA) policy. However, whether selecting better keyframes or retaining more frames effectively improves planning remains unclear. To investigate this, we replace predicted keyframes with ground-truth keyframes and separately double the keyframe output frequency. Ground-truth keyframes do not improve average success, while increased frequency yields only modest gains. Further analysis identifies two limitations that visual selection alone does not directly address. First, task-relevant conclusions remain implicit in visual content, requiring subsequent planning to reconstruct them even when the original supporting context has changed or become unavailable. Second, keyframe selection can discard intervening changes and interactions, leaving distinct histories indistinguishable from the retained snapshots. To address these limitations, we propose R2Mem (Reasoning to Remember), built around Structured Keyframe Memory (SKM). SKM pairs keyframes with reasoning text organized around task-relevant key events. These events specify which object changes and completed operations to record beyond what individual snapshots convey. The records form observation memory and action memory, explicitly preserving object histories and execution progress, respectively. The high-level VLM jointly constructs SKM and predicts subgoals using these explicit records. Experiments on RoboMME, RoboMemArena, and real-world manipulation tasks show that R2Mem improves average task success over memory-augmented baselines, highlighting the value of explicitly recording task-relevant events for subsequent subgoal planning. The code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.