Vision-Language Models Provide Task-Centric Promptable Representations for Latent Action Learning
Abstract
Latent Action Models (LAMs) have rapidly emerged as a promising approach for addressing the scarcity of high-quality action-labeled data. However, they fail when observations contain action-correlated distractors, often encoding noise instead of meaningful latent actions. Humans, on the other hand, can effortlessly distinguish task-relevant motions from irrelevant details in any video given only a brief task description. In this work, we propose to utilize the common-sense reasoning abilities of Vision-Language Models (VLMs) to provide promptable representations, effectively separating controllable changes from the noise. We use these representations as targets during LAM training and benchmark a wide variety of popular VLMs, revealing substantial variation in the quality of promptable representations as well as their robustness to different prompts and hyperparameters. Interestingly, we find that newer VLMs do not necessarily outperform older ones; however, all evaluated VLMs consistently outperform CLIP and DINOv2, highlighting the vital importance of language conditioning. Finally, we show that simply asking VLMs to ignore distractors can substantially improve latent action quality and downstream performance on Distracting MetaWorld.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.