Beyond the Current Observation: History-Informed Adaptation for Vision-Language-Action Models
Abstract
In robotic manipulation, the task-relevant scene state is often shaped by prior interactions. A robot must infer prior events, retain relevant evidence, and use it to guide subsequent actions. Most existing Vision-Language-Action (VLA) models, however, implicitly assume that the current observation alone provides sufficient information for action prediction. We introduce HI-Adapter, a lightweight and transferable history-informed adaptation framework that equips pretrained VLA policies to understand and utilize recent interaction history without robot-video pretraining. HI-Adapter comprises three components: a spatiotemporal vision adapter that aggregates multi-view observations across time to contextualize visual evidence; a reconstruction-guided memory adapter that distills task-relevant historical information into compact memory tokens; and an action expert adapter that selectively incorporates relevant memory representations into the action-generation process. Together, they improve temporal grounding and memory-guided manipulation in a 1B-parameter VLA policy while training only around 10% of its parameters. Our method achieves high success rates on RMBench and RoboMME, outperforming the baselines. It further reaches high performance on real-world tasks requiring recent historical context. Its consistent improvements when integrated into demonstrate transferability across VLA architectures. These results establish history-informed adaptation as an effective and parameter-efficient approach for overcoming the reliance of existing VLA policies on the current observation alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.