acceptodds
Under review as a conference paper at ICLR 2027

Distilling World Understanding from Video Generators for Robot Control

Abstract

Video generation models have shown promising results in robot control through future scene generation supported by their implicit understanding of the physical world. However, generating realistic textures does not guarantee an accurate representation of physical properties such as depth and motion, which are important for robot control. To leverage **K**nowledge for **R**obot control from the implicit **U**nderstanding in video **G**enerators, we propose **KRUG**, a two-stage framework based on generation as understanding. To activate this implicit world understanding, we apply multimodal instruction tuning to a pretrained video generator using a small amount of multimodal data, yielding a single World Understanding Expert that generates depth, surface normals, segmentation masks, or optical flow conditioned on modality-specific instructions. The challenge is to transfer this understanding to a student policy that uses RGB as its only visual input. We introduce World-to-Action Distillation (W2AD), which uses the expert as a frozen teacher and aligns features retrieved from the teacher and student using the same action queries to guide action prediction. For model training, we construct a multimodal embodied AI dataset by annotating RGB videos from nine sources, with 600,000 training clips and over 1,200 hours counted across available modalities. **KRUG** achieves strong performance on RoboTwin 2.0 and advances the **state of the art** on DexJoCo with an average success rate of 60.4%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.