The Simulator Says Yes: Language World Models Narrate the Executions Their Environment Refused
Abstract
Language world models—models fine-tuned on agent trajectories to predict the next tool observation—are now used as environments for training agents with reinforcement learning. We show that they systematically answer the wrong question. An environment must say what an action does and whether it is allowed; trajectories supervise the first on every step and give few examples of the second (1.5% of rows in the corpus we study). On 399 held-out actions the real executor refused, world models fine-tuned from Qwen3 (1.7B–8B) report a normal outcome instead of the refusal on 57 / 44 / 32% and the released 35B Qwen-AgentWorld on 25%—a failure we call whitewash—and the pretrained checkpoints, prompted, whitewash as much or more: fine-tuning removes only part of a deficit that predates it. Even on the rows an LLM reader (checked by one author on 100) judges decidable from the shown context, the 1.7B model still whitewashes 0.54 and the released model 0.14: much of the deficit is not missing information but a failure to use it. More refusal data helps, but not for the right reason: raising the 1.7B's refusal share from 1.5% to 21.3% lowers its whitewash from 0.57 to 0.22 (0.54 to 0.16 on the decidable rows, level with the released 35B) at a false-refusal cost, yet leaves a linear probe of the state unchanged and barely reaches planted state-dependent violations; in a τ-bench corpus where every refusal follows a missing read, the simulator learns that cue and no share moves it, which a two-prefix test exposes. Trained inside a simulator that never refuses, a GRPO policy stops refusing within twenty steps, because the task reward makes that optimal; the executor's refusals, passed through per call or used to threshold a learned head, restore it. What intermediate whitewash costs a policy is not established here. We release the refusal set, the planted battery and the instrument. Code and data: https://anonymous.4open.science/r/whitewash-audit-6635
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.