Transferable Behavior Embeddings from Unlabeled Videos for Policy Learning
Abstract
A robot policy chooses better actions when it knows how the scene is changing, where that change leads, and how far the current step has progressed. We learn this information from unlabeled manipulation videos as behavior embeddings that describe subprocesses: coherent video segments discovered from the embeddings at several granularities. Each embedding combines a motion channel describing ongoing motion with a scene configuration channel describing the configuration a subprocess reaches and its relation to the current scene. Three tasks give these channels their meaning: next-frame prediction, endpoint reconstruction, and status scoring. Learning requires no action, language, subtask, or progress annotations. The frozen encoder then computes embeddings for policy demonstrations. Predicting these embeddings from the policy’s action features regularizes those features with no additional inference cost. Experiments show consistent improvements in task success across vision–language–action and world-action models. On a real robot, behavior supervision also improves task success and generalization to unseen instruction combinations and appearance changes. Probes also show that motion, reached configuration, and execution status are readable from the embeddings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.