acceptodds
Under review as a conference paper at ICLR 2027

X-World: Real-Time Cross-Embodiment Policies with Scalable Video World Priors

Abstract

Language-instructed manipulation requires robust closed-loop policies, yet action-labeled robot trajectories are costly and limited in coverage. Human and robot videos provide rich spatiotemporal interaction priors that action supervision alone cannot easily capture. Transferring video diffusion priors to policy learning must accommodate variable camera layouts, heterogeneous state and action interfaces, and real-time control. We propose X-World, which adapts a pretrained video diffusion backbone to variable-view robot observations, distills its hidden states into a compact token sequence with a Q-Former, and trains a flow-matching action expert that attends to this sequence to predict action chunks. Dataset-specific modules align heterogeneous state and action interfaces. We first adapt the video diffusion backbone on large-scale robot and human videos, freeze it while training the Q-Former and action expert with an action-only objective, and then post-train on the target platform. At inference, X-World maintains a 20Hz model response rate and uses asynchronous real-time chunking to decouple model updates from action execution, avoiding waiting latency. It shows strong performance on in-distribution LIBERO, zero-shot LIBERO-Plus distribution shifts, and real-world manipulation tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.