Learning Action-Effect Latents for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) policies are typically trained with embodiment-specific motor commands, creating a mismatch between rich multimodal representations and the low-level signals used for action supervision. We introduce AEL (Action-Effect Latents), a two-stage framework that augments action representation learning with the observed visual consequences of executed actions. In Stage I, AEL jointly encodes the current visual state and an executed action chunk into a latent representation that preserves the motor trajectory through action reconstruction while reconstructing the corresponding future visual representation through a residual effect formulation. In Stage II, this latent becomes the prediction target of a VLA policy conditioned on the current visual observations, language instruction, and robot state, and is decoded into continuous actions. Future observations are used only as supervision during Stage I and are not required at deployment. Separating representation learning from policy learning also allows suboptimal interactions to contribute action-effect supervision without being treated as imitation targets. Experiments on LIBERO, CALVIN, and real-world single-arm and bimanual manipulation show consistent improvements over matched direct-action baselines, supporting visual effects as a complementary supervision signal for learning robot action representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.