Revealing State Aliasing in Visual Representations of Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models are a promising framework toward general-purpose robot manipulation that leverages the rich prior knowledge of pretrained vision-language models for diverse tasks. However, they often struggle with tasks that require high-precision control. Our analysis across various visual encoders in VLA shows that a key source of the limitation is state aliasing, where the visual encoder fails to distinguish fine-grained states that require different actions even though the states are visually similar. Motivated by this observation, we propose a simple yet effective auxiliary inverse dynamics objective for VLA training. By predicting intervening actions between two frames in a demonstration trajectory, our auxiliary task encourages the visual encoder to preserve action-relevant state distinctions. Our approach requires no additional manual annotations and the inverse dynamics prediction is used only during training, introducing no additional computational overhead at inference. Our experiments on LIBERO, CALVIN ABCD, and SimplerEnv demonstrate consistent improvements in manipulation performance over standard VLA training across diverse VLA architectures. Representation-level analyses further show reduced state aliasing and more accurate action prediction in the learned visual representations, consistent with the trend observed across visual encoders.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.