MiX-World: A Unified World Simulator for Manipulators and Locomoting Humanoids
Abstract
Action-conditioned world models forecast how visual observations evolve in response to robot actions, yet existing approaches are largely confined to a single robot embodiment. Although robots differ substantially in morphology and control, they interact with the same physical world, suggesting that experience collected across embodiments may provide complementary supervision. However, bringing these experiences together in a single model is challenging. Low-level actions differ across embodiments in both control semantics and action-space parameterization, while trajectories span different physical timescales. In this paper, we present MiX-World, a unified action-conditioned world simulator for heterogeneous robot embodiments. To our knowledge, it is the first such model to include full-body locomoting humanoids in joint training with mobile-base and fixed-base bimanual manipulators. MiX-World uses Semantically Structured Conditioning (SSC) to represent heterogeneous robot controls, while a shared generator predicts synchronized future observations from multiple views. To capture both transient interactions and long-horizon evolution, MiX-World introduces Event-Anchored Scheduling (EAS) with Time-Anchored RoPE (TA-RoPE). MiX-World achieves better aggregate visual prediction performance than the evaluated baselines at both short and long horizons. With one demonstration per task and 100 adaptation updates, pretraining also improves full-rollout prediction on three unseen embodiments. Controlled temporal ablations further demonstrate that variable-time training improves both short-timescale event prediction and long-horizon forecasting. MiX-World thus provides a scalable foundation for policy evaluation, interactive data collection, and policy post-training across diverse robot embodiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.