acceptodds
Under review as a conference paper at ICLR 2027

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Abstract

World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their future prediction targets typically come from recorded robot trajectories, limiting supervision when target-task demonstrations are scarce. Therefore, we ask whether future supervision for WAMs must originate from target-task expert trajectories. In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. Given an observation and language instruction, Vid2WAM distills supervision through two complementary channels: task-conditioned future rollouts directly supervise the student's future prediction branch, while an inverse dynamics model recovers embodiment-specific pseudo-actions for action learning. To integrate synthetic and real supervision, we introduce source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from noisy pseudo-actions. During inference, both the video teacher and inverse dynamics model are discarded, leaving only the WAM student for efficient deployment. Simulation and real-world experiments show gains under limited expert demonstrations and in target-conditioned adaptation to novel tasks without target-task action trajectories, while preserving low-latency inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.