World Model Self-Distillation: Turning Video Generators into Task Solvers
Abstract
Pretrained video generators are promising visual world models for planning, as they can depict how actions change a scene, but solving tasks from high-level goals remains difficult without detailed execution descriptions. Common approaches outsource execution reasoning to language or vision-language models (VLMs), or require costly paired task-execution videos. We introduce World Model Self- Distillation (WMSD), combining self-distillation and reinforcement learning with- out curated task-video supervision. From an unlabeled scene image, a VLM generates a task and a step-by-step execution description. This description guides a frozen pretrained video generator, the Demonstrator; self-distillation transfers its behavior into an Executor conditioned only on the image and a short task prompt. VLM feedback further improves the Executor, exploiting the asymmetry between generating a solution and judging a sampled execution. Inference requires only the Executor. On WorldTasks-Bench, eight-step LTX-2 reaches 60.5% task completion, surpassing the Demonstrator’s 49.5% and Base’s 28.5% under VLM evaluation. Transfer to robotic video generation suggests generalization without robotics-specific fine-tuning. Independent and human evaluations support the task-solving gains. Verified maze solving shows how lightweight Demonstrator adaptation can enable transfer to new domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.