LTM: Latent Transition Based In-Context Memory for Vision-Language-Action Models
Abstract
Robotic manipulation requires remembering interactions that are no longer observable in the current scene. Existing memory mechanisms face a trade-off between retaining detailed histories and maintaining a compact policy context. We introduce Latent Transition Memory (LTM), which represents interaction history as an ordered sequence of compact latent transitions between consecutive observations. A frozen latent action encoder extracts each transition, and a lightweight adapter maps it to a single memory token, preserving a record of observed changes at one token per transition. We instantiate LTM on as LTM-, using causal attention and persistent KV caching for incremental inference, with sparse visual anchors to ground video demonstrations. Across two pretrained latent encoders, LTM- achieves up to 58.54% success on RoboMME and 80.5% on LIBERO-Mem, exceeding the reported memory-free baseline by 40.61 percentage points and the strongest compared LIBERO-Mem baseline by 38.0 points, respectively. On a real robot, it improves average exact-count accuracy from 11.25% to 82.50% in a pick-and-place \(x\) times task. Ablations support the contribution of transition content beyond temporal cues, while a controlled swap diagnostic shows that relevant motion information can remain recoverable even when downstream state tracking fails. These findings establish latent transitions as an effective, compact memory representation and highlight temporal composition as a remaining challenge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.