acceptodds
Under review as a conference paper at ICLR 2027

LATCH: Action-Guided Perceptual Memory for Vision-Language-Action Models

Abstract

Non-Markovian manipulation requires a vision-language-action (VLA) policy to know which past observations still matter for the action being produced. Most VLAs act from the current observation or a short visual window; existing mem- ory systems either add symbolic task state through a separate high-level policy trained on subtask or keyframe annotation, or retain perceptual history with no task-general way to prioritize useful past observations. We introduce LATCH, which learns this priority inside the policy: lightweight read and write gates con- trol where the action queries attend within a uniformly sampled visual history and how strongly each frame contributes, optimized jointly with the VLA ac- tion objective. The action objective is the only supervision signal: no annotation pipeline, no retrieval policy, and a single differentiable policy at inference. We evaluate LATCH on the π0.5 and GR00T N1.6 backbones across RoboMME’s 16 memory-centric tasks, the LIBERO suites, dual-arm RMBench, and real-world memory-centric tasks on an SO-101 arm. On RoboMME, LATCH lifts π0.5 from 17.9% to 47.2% average success, higher than all other evaluated memory methods for VLAs. These results indicate that a bounded perceptual memory can be orga- nized by the action objective alone and can effectively provide historical context without keyframe or subtask annotation and without an external memory policy. Project page: https://latch-robomem.github.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.