LATCH: Action-Guided Perceptual Memory for Vision-Language-Action Models
Abstract
Non-Markovian manipulation requires a vision-language-action (VLA) policy to know which past observations still matter for the action being produced. Most VLAs act from the current observation or a short visual window; existing mem- ory systems either add symbolic task state through a separate high-level policy trained on subtask or keyframe annotation, or retain perceptual history with no task-general way to prioritize useful past observations. We introduce LATCH, which learns this priority inside the policy: lightweight read and write gates con- trol where the action queries attend within a uniformly sampled visual history and how strongly each frame contributes, optimized jointly with the VLA ac- tion objective. The action objective is the only supervision signal: no annotation pipeline, no retrieval policy, and a single differentiable policy at inference. We evaluate LATCH on the π0.5 and GR00T N1.6 backbones across RoboMME’s 16 memory-centric tasks, the LIBERO suites, dual-arm RMBench, and real-world memory-centric tasks on an SO-101 arm. On RoboMME, LATCH lifts π0.5 from 17.9% to 47.2% average success, higher than all other evaluated memory methods for VLAs. These results indicate that a bounded perceptual memory can be orga- nized by the action objective alone and can effectively provide historical context without keyframe or subtask annotation and without an external memory policy. Project page: https://latch-robomem.github.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.