World Boundary Bench: Probing the Failure Mode of Language World Models
Abstract
Language World Models (LWMs) are evolving from next state decision reference to interactive world simulators. Their reliability depends on whether they could maintain decisive world knowledge, preserve the consistent states in long horizon, and maintain their simulator role when environment data contains instruction like text. Therefore, to further investigate current LWMs' failure modes, we introduce World Boundary Bench (WBB), a benchmark spanning Terminal, MCP, and Web environments. WBB contains three tracks: 1) Causal Forking Track tests whether an LWM notices a small but decisive difference, such as whether incompatible torch and torchvision versions cause pip installation to fail. 2) World Commitment Track simulates scenarios with limited context and observes what kind of next state the LWM will output, and whether its further outputs remain consistent with that state. 3) Role Boundary Track tests whether embedded instructions within terminal or tool commands detach the LWM from its world model role. Across 13 proprietary and open source models, Kimi K3 leads Causal Forking and Role Boundary and ties with Qwen 3.8 Max 0902 on World Commitment, while models exhibit distinct failure modes across tracks. In our sim to real study, rejection finetuning on LWM self-reported success trajectories improves real environment success from 2.81% to 4.41% after simulated practice, illustrating potential agentic training utility beyond simulation failures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.