acceptodds
Under review as a conference paper at ICLR 2027

Learning Action–Consequence Representations for Vision-Language-Action Models

Abstract

Existing future-oriented vision-language-action (VLA) methods typically predict actions and future observations in parallel from the same context, without explicitly modeling how actions lead to subsequent changes in the environment. Consequently, a model may learn to anticipate what will happen next largely from scene context, without learning what will happen as a consequence of a particular action. To address this limitation, we propose a framework for learning action–consequence representations (ACE) from future supervision. ACE models future states as consequences of actions and contrasts demonstrated actions with alternative actions under the same context and observed consequence, enabling future representations to distinguish different actions and their corresponding environmental changes. ACE further emphasizes action-relevant changes by weighting interaction regions and uses task-terminal states to constrain the task relevance of consequence representations. We evaluate different forms of future supervision on LIBERO and RoboCasa under a unified training budget. Under a unified training and evaluation protocol, ACE improves the average success rate on LIBERO by 6.4 percentage points over a matched VLA baseline. On RoboCasa, ACE outperforms VLA and FOCA baselines by 5.0 and 8.0 percentage points, respectively. These results suggest that modeling how actions lead to subsequent environmental changes enables future supervision to translate more effectively into control performance than independent future prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.