Learning to Act and Edit in Video Worlds
Abstract
Interactive video world models have made substantial progress in generating explorable environments with navigation control. However, extending such control to diverse character actions and environmental events is hindered by fragmented supervision across heterogeneous video sources. To bridge this gap, we develop a unified data pipeline that transforms real-world videos, gameplay video, edited videos and synthetic videos into a temporally aligned dataset for camera control, character actions, and environmental events. The resulting dataset contains 180,162 clips, covering 91 action types across 14 visual domains. Using this dataset, we train an autoregressive world model initialized from a pretrained bidirectional video generator . Experiments show that the resulting model (i) achieves strong instruction following across everyday interactions, gameplay actions, visual effects, and scene-level edits while preserving spatial navigation fidelity, (ii) exhibits zero-shot compositional generalization across unseen instruction combinations alongside precise camera-actor disentanglement, and (iii) reveals the distinct roles of different data sources, yielding a reproducible, resource-efficient recipe for training semantically controllable video world models. We will release the code and dataset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.