acceptodds
Under review as a conference paper at ICLR 2027

Memento: Distilling Memory Capability from VLMs to VLAs

Abstract

Sequential multimodal decision making often requires retaining information that is no longer available in the current observation. This poses a fundamental challenge for Vision-Language-Action (VLA) policies, whose predictions are typically conditioned on limited context despite many tasks depending on information accumulated over interaction histories. Addressing such history dependence requires retaining relevant context over time, yet naively extending the context is computationally prohibitive, and hierarchical architectures that pair a high-level Vision-Language-Model (VLM) with a low-level VLA incur the cost of an additional large model throughout inference. In this work, we introduce Memento, a framework that distills multi-layer representations of interaction histories from a pretrained VLM into a lightweight online memory module based on Test-Time Training (TTT). Memento learns to encode task-relevant historical information into fast weights through online updates. The resulting memory readout is projected into memory tokens, providing the downstream VLA policy with compact access to accumulated context at bounded inference cost. In our experiments, we instantiate Memento by distilling multi-layer representations from Qwen3.5-4B and integrating the resulting memory module with . Memento improves the average success rate from 19.33% to 41.04% on RoboMME and from 10.56% to 51.67% on RMBench, while maintaining comparable performance on LIBERO, where long-term interaction history is less critical. These results demonstrate that memory capability can be distilled from a pretrained VLM, endowing VLA policies with the ability to predict actions conditioned on interaction histories.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.