OpenWorldAct: Open-Vocabulary Event Control for Real-Time Interactive World Models
Abstract
We introduce **OpenWorldAct**, a real-time interactive world model that executes temporally grounded open-vocabulary events while achieving accurate navigation control from a third-person view. Most existing world models support only exploratory action spaces and do not allow users to control open-vocabulary events. The central obstacle is data: gameplay videos provide frame-synchronous controls but cover only a narrow range of semantics, while open or generated videos contain diverse events but can provide only inaccurate control signals obtained through estimation. Data that satisfy both requirements are difficult to obtain. To overcome this bottleneck, we build **ActMosaic**, a complementary data engine that divides its data into two streams. The gameplay stream retains frame-synchronous keyboard and mouse inputs to represent local navigation intent, while the generated event stream is organized around controllable characters, environments, and character–world interactions. At inference time, OpenWorldAct uses a unified time-aligned prompt interface to specify navigation controls and open-vocabulary events during the same rollout. We further develop a dual-resolution streaming few-step autoregressive model, where a low-resolution causal rollout is refined online by a high-resolution branch, enabling long-horizon real-time 720p generation. As an application showcase, we build a Gameplay Harness that maps a user's world setting to state-conditioned skills and tool-assisted observations. Comprehensive experiments show that OpenWorldAct substantially outperforms the evaluated systems in event control while achieving comparable wandering control and visual quality. **We will open source the model weights, code, data and benchmark**.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.