acceptodds
Under review as a conference paper at ICLR 2027

Scaling Interactive Omnimodal World Models from Heterogeneous Supervision

Abstract

Interactive omnimodal world models should respond to continuous control while generating the synchronized video, environmental sound, speech, and music that make a world feel alive. Existing approaches typically use source-specific action spaces or focus on visual prediction, making it difficult to combine heterogeneous interaction and audiovisual supervision. We present EchoWM, an interactive omnimodal world model that canonicalizes interaction as a dataset-calibrated relative 6-DoF camera trajectory. The same camera-intent interface supports first-person observer motion and third-person camera-subject-world evolution, with viewpoint semantics learned from data. To reconcile complementary supervision, we progressively train an audiovisual prior, trajectory conditioning, and their joint realization on AV-rich, control-clean, and balanced data mixtures. We then convert the model for causal few-step deployment through audiovisual teacher forcing, self-generated distribution matching, and long-horizon rollouts, exposing it to the errors that accumulate during interactive generation. On WBench Navigation and SANA-WM-Bench, \model achieves leading interactive and visual-quality performance, while its four-step causal variant largely preserves these capabilities over long-horizon rollout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.