Scaling Interactive Omnimodal World Models from Heterogeneous Supervision
Abstract
Interactive omnimodal world models should respond to continuous control while generating the synchronized video, environmental sound, speech, and music that make a world feel alive. Existing approaches typically use source-specific action spaces or focus on visual prediction, making it difficult to combine heterogeneous interaction and audiovisual supervision. We present EchoWM, an interactive omnimodal world model that canonicalizes interaction as a dataset-calibrated relative 6-DoF camera trajectory. The same camera-intent interface supports first-person observer motion and third-person camera-subject-world evolution, with viewpoint semantics learned from data. To reconcile complementary supervision, we progressively train an audiovisual prior, trajectory conditioning, and their joint realization on AV-rich, control-clean, and balanced data mixtures. We then convert the model for causal few-step deployment through audiovisual teacher forcing, self-generated distribution matching, and long-horizon rollouts, exposing it to the errors that accumulate during interactive generation. On WBench Navigation and SANA-WM-Bench, \model achieves leading interactive and visual-quality performance, while its four-step causal variant largely preserves these capabilities over long-horizon rollout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.