acceptodds
Under review as a conference paper at ICLR 2027

LTM: Latent Transition Based In-Context Memory for Vision-Language-Action Models

Abstract

Robotic manipulation requires remembering interactions that are no longer observable in the current scene. Existing memory mechanisms face a trade-off between retaining detailed histories and maintaining a compact policy context. We introduce Latent Transition Memory (LTM), which represents interaction history as an ordered sequence of compact latent transitions between consecutive observations. A frozen latent action encoder extracts each transition, and a lightweight adapter maps it to a single memory token, preserving a record of observed changes at one token per transition. We instantiate LTM on as LTM-, using causal attention and persistent KV caching for incremental inference, with sparse visual anchors to ground video demonstrations. Across two pretrained latent encoders, LTM- achieves up to 58.54% success on RoboMME and 80.5% on LIBERO-Mem, exceeding the reported memory-free baseline by 40.61 percentage points and the strongest compared LIBERO-Mem baseline by 38.0 points, respectively. On a real robot, it improves average exact-count accuracy from 11.25% to 82.50% in a pick-and-place \(x\) times task. Ablations support the contribution of transition content beyond temporal cues, while a controlled swap diagnostic shows that relevant motion information can remain recoverable even when downstream state tracking fails. These findings establish latent transitions as an effective, compact memory representation and highlight temporal composition as a remaining challenge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.