World Intelligence Agent: Unifying World Modeling and World Action for Digital Agents
Abstract
Digital agents built on large language models have made rapid progress, with world models emerging as their fundamental predictive engines for understanding complex environments. However, current world models are largely fragmented: they lack a unified data pipeline, are often tied to specific domains, and are treated as passive, isolated simulators. We argue that advancing digital agents requires a true **foundation world model**: a universal cognitive engine that not only masters cross-domain simulation but also yields predictive representations robust enough to extend naturally into downstream applications. To realize this vision of **World Intelligence**, we first construct **WIA-Corpus** (5M continual pre-training examples, 9k SFT trajectories, and 134.6k RL prompts) and **WIA-Bench**, the first paired in-domain (1.8k) and out-of-domain (1.7k) evaluation benchmark spanning seven digital domains. Building on this cohesive data foundation, we develop **WIA-WM**, a foundation world model family (4B / 9B / 27B / 35B-A3B), for digital agents, that masters systematic state-transition modeling. WIA-WM achieves the highest average world-model score among all evaluated models on WIA-Bench, surpassing GPT-5.6-sol and Claude-Opus-4.8 by **** and **** on the seven-domain OOD suite, and empowers downstream agents via simulation SFT, simulation RL, and inference lookahead. Finally, as a natural extension of this predictive engine, we introduce **WIA-WAM**, which couples prediction and execution via Dual-Perspective Inversion (DPI) and Counterfactual Simulation-Guided Policy Optimization (CSPO). This integration allows the agent to implicitly internalize environment dynamics, moving beyond passive simulation to proactive, robust decision-making. Our dataset, code, and models will be released to facilitate future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.