Mem4Act: Learning Effective Latent Memory for Long-Horizon Robotic Manipulation
Abstract
Long-horizon robotic manipulation fundamentally depends on memory: agents must retain information from past experiences and interactions and take correct actions. Existing approaches typically preserve historical observations, compressed visual tokens, or language summaries, but these representations either expose the action policy to increasingly long histories or risk discarding task-relevant temporal, semantic, and perceptual information. We introduce Mem4Act, a framework for learning effective latent memory as a fixed-size representation explicitly optimized for long-horizon robotic control. Mem4Act addresses three coupled challenges: how to cover long histories, what information to retain, and how to ensure the policy actually uses memory. It organizes history through multi-granularity temporal aggregation, preserves complementary semantic and perceptual information through question answering and key-frame DINO feature reconstruction, and adopts memory-first two-stage training to mitigate reactive shortcuts. The resulting representation captures both high-level task state and fine-grained perceptual evidence while maintaining a bounded interface to the action policy. Across RoboMME, RMBench, and four real-world manipulation tasks, Mem4Act achieves state-of-the-art performance in memory-dependent control, reaching success rates of 49.88%, 61.0%, and 81.50%, respectively. Ablations further demonstrate the importance of temporal structure, complementary memory supervision, and memory-first two-stage learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.