Beyond Context: Executable Memory for Embodied Agents
Abstract
Long-horizon robotic tasks require memory because the same current observation can call for different actions depending on earlier interactions. Existing memory-augmented systems improve how history is stored and retrieved, but often pass it to a learned model whose interpretation determines the next decision—an interface we call Context-Memory. Yet recalling the relevant past does not by itself determine its role in execution: identifying a previously used target and deciding whether to repeat an interaction require different operations on history. Our insight is that these operations can be made explicit, allowing remembered information to update task state rather than merely provide context for inferring it. We introduce Execute-Memory, which couples persistent entity and event memory with compositional task programs that resolve historical references, maintain progress, and determine the next execution requirement. By keeping uncertain world beliefs separate from accepted events, it links progress updates to observed interactions without assuming perfect perception. The resulting requirement is grounded in the current scene and passed as a local subgoal to a frozen vision-language-action policy. Experiments on RoboMME and challenging stress subsets derived from it show substantial gains over memory-augmented and policy baselines, supporting a shift from remembered context to executable memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.