acceptodds
Under review as a conference paper at ICLR 2027

FAME-VLA: Full-history Adapter for Memory -Enhanced vision-language-action Models

Abstract

Pretrained Vision-Language-Action (VLA) models have shown strong promise for robot manipulation. However, most existing VLAs condition primarily on the current observation or a short window of recent observations, limiting their ability to resolve history-dependent ambiguities in long-horizon tasks. To address this, we introduce , a lightweight full-history adapter that equips pretrained VLAs with dual-stream memory for history-aware action generation. Our adapter combines two complementary memory branches: full-history temporal memory, which compresses the complete observation history into a compact task context, and persistent initial-scene memory, which preserves access to spatial details that may become occluded or altered during execution. These history-aware representations are injected exclusively into the action expert, making our adapter a lightweight, plug-and-play module for pretrained VLAs. To efficiently scale to full observation histories, we precompute history features once per trajectory during training. Experiments on simulation benchmarks and real-world manipulation tasks demonstrate consistent improvements across pretrained VLA backbones, achieving average success rates of 56.11% with StarVLA and 64.89% with on RMBench. Further analyses demonstrate the complementary roles of the two memory branches and the effectiveness of full-history conditioning, while history sharing provides up to a training speedup over independently encoding each observation prefix.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.