MemAlign: Memory–Action Alignment for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models have achieved remarkable progress in robotic manipulation, particularly when augmented with memory to support history-dependent decision-making under partial observability. However, memory-augmented policies may fail to translate historical evidence into appropriate actions, a gap we refer to as memory-action misalignment. In this work, we introduce MemAlign, a post-training framework that aligns action generation with historical evidence retained in memory, enabling policies to resolve ambiguous action choices using historical context. Specifically, we first introduce counterfactual supervision to train the policy to select appropriate actions for different histories under the same current observation and instruction. Then, we propose a history-aware low-rank modulation module to dynamically modulate the action network based on historical evidence, allowing different histories to induce different effective action mappings. Finally, we complement this modulation with an optional visual rebinding module that locates remembered targets in the current scene. Their estimated positions then guide the refinement of the generated actions. Extensive experiments on two long-horizon manipulation benchmarks (MIKASA-Robo and RoboMME) demonstrate that MemAlign significantly improves the success rate of completing history-specified targets, e.g., raising MemoryVLA’s result from to on MIKASA-Robo tasks. Notably, MemAlign exhibits strong generalization capabilities, consistently improving performance when applied to different memory-augmented VLA models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.