acceptodds
Under review as a conference paper at ICLR 2027

WALA: Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

Abstract

Human videos provide rich information about object manipulation at a scale difficult to reproduce with robots, but often lack action annotations that can directly supervise robot policies. Realizing this potential requires distinguishing interaction-relevant changes from static appearance and connecting the learned interaction knowledge to executable robot actions. We present , a framework named for . Its semantic-geometric latent action model (LAM) learns latent actions by encoding and predicting semantic and geometric changes, emphasizing interactions over static appearance while retaining spatial detail. During robot policy training, the LAM encoder and decoder provide guidance through distillation and visual prediction, respectively. Together with robot action supervision, these signals shape policy-generated latent actions to support world prediction and capture the information needed to directly generate executable robot actions. Because LAM supervision requires no action labels, WALA enables co-training on action-labeled robot demonstrations and task-relevant action-free videos. In simulation experiments, our method achieves % average success on RoboTwin and % on RoboCasa-GR1-Tabletop, exceeding the previous state of the art on RoboCasa by percentage points. WALA also shows strong real-robot performance, further enhanced by task-relevant action-free human videos that improve data efficiency and support target-task zero-shot and few-shot transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.