acceptodds
Under review as a conference paper at ICLR 2027

Hybrid Memory-Augmented Vision-Language-Action Models for Long-Horizon Manipulation

Abstract

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for closed-loop robot control, yet their limited memory remains a critical bottleneck for long-horizon embodied manipulation. Existing memory-augmented VLAs typically rely on short temporal windows, recurrent memory, or explicit visual memories. These approaches either emphasize recent context, compress away long-range evidence, or retain redundant historical observations, making it difficult to jointly maintain temporal continuity and preserve task-relevant evidence over long horizons. To address these limitations, we propose HERM-VLA, a Hybrid Episodic–Recurrent Memory Vision-Language-Action Model that combines recurrent state propagation with explicit episodic evidence retention. The Recurrent Working Memory maintains a compact latent state to track task evolution, while the Episodic Evidence Memory preserves informative historical context as compact snapshots. An action-aware evidence consolidation mechanism selects action-relevant visual features together with cognitive and spatial context, reducing redundant historical information. The resulting episodic memory is summarized and integrated with the recurrent state and current context for action prediction, enabling effective long-horizon control. Experiments on long-horizon manipulation tasks demonstrate the effectiveness of HERM-VLA and the complementary roles of its two memory components.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.