Controlled Predictive Memory for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models need compact memories to distinguish execution histories that require different decisions despite similar current observations. Predicting future observations can guide this compression, but their dependence on subsequent behavior complicates comparisons between histories. We introduce Controlled Predictive Memory (CPM), which learns memory by predicting expected joint outcome features under a fixed set of short closed-loop interaction programs. The fixed programs provide a common behavioral reference for comparing histories, while joint features expose selected cross-step dependencies that separate endpoint means can miss. Response prediction and action imitation jointly train a fixed-size recurrent memory from additional program interactions and expert demonstrations; the programs and response decoder are used only during training. A prediction-risk decomposition characterizes the response information lost through history compression. CPM improves average success on each of five simulation benchmarks over two history-based baselines with the same VLA backbone and matched action-update budgets. On five single-arm and two bimanual real-world tasks, CPM achieves average success rates of 73.3% and 63.3%, respectively, exceeding the stronger history-based baseline on each platform by 6.7 and 8.3 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.