Learning Action–Consequence Representations for Vision-Language-Action Models
Abstract
Existing future-oriented vision-language-action (VLA) methods typically predict actions and future observations in parallel from the same context, without explicitly modeling how actions lead to subsequent changes in the environment. Consequently, a model may learn to anticipate what will happen next largely from scene context, without learning what will happen as a consequence of a particular action. To address this limitation, we propose a framework for learning action–consequence representations (ACE) from future supervision. ACE models future states as consequences of actions and contrasts demonstrated actions with alternative actions under the same context and observed consequence, enabling future representations to distinguish different actions and their corresponding environmental changes. ACE further emphasizes action-relevant changes by weighting interaction regions and uses task-terminal states to constrain the task relevance of consequence representations. We evaluate different forms of future supervision on LIBERO and RoboCasa under a unified training budget. Under a unified training and evaluation protocol, ACE improves the average success rate on LIBERO by 6.4 percentage points over a matched VLA baseline. On RoboCasa, ACE outperforms VLA and FOCA baselines by 5.0 and 8.0 percentage points, respectively. These results suggest that modeling how actions lead to subsequent environmental changes enables future supervision to translate more effectively into control performance than independent future prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.