acceptodds
Under review as a conference paper at ICLR 2027

World Boundary Bench: Probing the Failure Mode of Language World Models

Abstract

Language World Models (LWMs) are evolving from next state decision reference to interactive world simulators. Their reliability depends on whether they could maintain decisive world knowledge, preserve the consistent states in long horizon, and maintain their simulator role when environment data contains instruction like text. Therefore, to further investigate current LWMs' failure modes, we introduce World Boundary Bench (WBB), a benchmark spanning Terminal, MCP, and Web environments. WBB contains three tracks: 1) Causal Forking Track tests whether an LWM notices a small but decisive difference, such as whether incompatible torch and torchvision versions cause pip installation to fail. 2) World Commitment Track simulates scenarios with limited context and observes what kind of next state the LWM will output, and whether its further outputs remain consistent with that state. 3) Role Boundary Track tests whether embedded instructions within terminal or tool commands detach the LWM from its world model role. Across 13 proprietary and open source models, Kimi K3 leads Causal Forking and Role Boundary and ties with Qwen 3.8 Max 0902 on World Commitment, while models exhibit distinct failure modes across tracks. In our sim to real study, rejection finetuning on LWM self-reported success trajectories improves real environment success from 2.81% to 4.41% after simulated practice, illustrating potential agentic training utility beyond simulation failures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.