WALA: Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Abstract
Human videos provide rich information about object manipulation at a scale difficult to reproduce with robots, but often lack action annotations that can directly supervise robot policies. Realizing this potential requires distinguishing interaction-relevant changes from static appearance and connecting the learned interaction knowledge to executable robot actions. We present , a framework named for . Its semantic-geometric latent action model (LAM) learns latent actions by encoding and predicting semantic and geometric changes, emphasizing interactions over static appearance while retaining spatial detail. During robot policy training, the LAM encoder and decoder provide guidance through distillation and visual prediction, respectively. Together with robot action supervision, these signals shape policy-generated latent actions to support world prediction and capture the information needed to directly generate executable robot actions. Because LAM supervision requires no action labels, WALA enables co-training on action-labeled robot demonstrations and task-relevant action-free videos. In simulation experiments, our method achieves % average success on RoboTwin and % on RoboCasa-GR1-Tabletop, exceeding the previous state of the art on RoboCasa by percentage points. WALA also shows strong real-robot performance, further enhanced by task-relevant action-free human videos that improve data efficiency and support target-task zero-shot and few-shot transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.