SCAR: Supervised, Compressed, and Aligned Representations for Memory-Augmented Vision-Language-Action Models
Abstract
In long-horizon robotic manipulation, vision-language-action (VLA) policies must retain task-relevant information that may not be observable from the current input. Recent works on perceptual memory approach this problem by conditioning VLAs on past observations, but learning effective memory representations remains challenging. Action prediction provides only sparse supervision for information that must be retained over long horizons, raw observation histories contain temporal redundancy, and newly introduced memory representations are often underutilized by the pretrained policies. To address these challenges, we introduce SCAR (Supervised, Compressed, and Aligned memory Representations), which combines textual supervision of memory-relevant events, learned temporal compression, and staged memory alignment. We evaluate SCAR on 16 memory-dependent manipulation tasks from the RoboMME benchmark, spanning temporal, spatial, object, and procedural memory. SCAR improves average success rate from to over the closest architectural perceptual-memory baseline, with controlled ablations demonstrating the contribution of all three components. Across four real-world manipulation tasks, SCAR further improves average success from to over the same baseline. These results demonstrate that explicitly supervising, compressing, and aligning perceptual memory improves memory-dependent manipulation in both simulation and real-world settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.