acceptodds
Under review as a conference paper at ICLR 2027

Learning Action-Effect Latents for Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) policies are typically trained with embodiment-specific motor commands, creating a mismatch between rich multimodal representations and the low-level signals used for action supervision. We introduce AEL (Action-Effect Latents), a two-stage framework that augments action representation learning with the observed visual consequences of executed actions. In Stage I, AEL jointly encodes the current visual state and an executed action chunk into a latent representation that preserves the motor trajectory through action reconstruction while reconstructing the corresponding future visual representation through a residual effect formulation. In Stage II, this latent becomes the prediction target of a VLA policy conditioned on the current visual observations, language instruction, and robot state, and is decoded into continuous actions. Future observations are used only as supervision during Stage I and are not required at deployment. Separating representation learning from policy learning also allows suboptimal interactions to contribute action-effect supervision without being treated as imitation targets. Experiments on LIBERO, CALVIN, and real-world single-arm and bimanual manipulation show consistent improvements over matched direct-action baselines, supporting visual effects as a complementary supervision signal for learning robot action representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.