acceptodds
Under review as a conference paper at ICLR 2027

Self-Improving Video-to-Action Models

Abstract

Video generative models are an emerging tool for embodied agents: they can serve as scalable visual planners, generating videos that depict how to complete previously unseen tasks. However, executing these plans typically requires a video-to-action or "inverse dynamics" model trained on large, human-collected video-action datasets, severely limiting scalability. We introduce Mirror, a self-supervised approach that learns video-to-action mappings through environment interactions, without prior action data or hand-crafted rewards. Mirror learns by attempting to follow a visual plan by executing its best guess of the underlying actions in the environment. Mirror then trains on the resulting video-action pairs, incrementally self-improving and following the visual planner more and more closely. Across Atari, OGBench, and DeepMind Control, Mirror achieves overall performance comparable to leading learning-from-observation methods with learned video plans, and near-expert performance with oracle plans. On Minecraft BASALT and text-conditioned tasks, it outperforms methods with access to expert actions or hand-engineered task knowledge. Using generated video plans, it surpasses VPT-BC in survival mode, even though VPT-BC was trained on thousands of hours of action-labeled human gameplay. With oracle video plans, it learns to acquire diamonds without task rewards, using an order of magnitude fewer environment interactions than DreamerV3. Our results suggest that advances in video models can translate into increasingly capable agents without a corresponding need for hand-crafted rewards or large collections of expert action data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.