WorldSimBench: Benchmarking Large Language Models as Environment Simulators
Abstract
Training agents at scale requires large volumes of interaction trajectories, which are costly to collect from real environments. LLM-based generative simulators offer a scalable alternative, but existing public evaluations are often limited to narrow interaction settings and short trajectories, or rely on open-ended LLM-based scoring. We introduce WorldSimBench, a benchmark for evaluating LLMs as environment simulators across diverse tasks and long-horizon interactions. It comprises 3,580 tasks from 322 environments across ToolUse, Terminal, Game, Web, and Search. To construct long-horizon trajectories with fixed reference outcomes, we use story-driven plans to organize interactions and executable environments to produce the resulting states and observations. ToolUse, Terminal, Game, and Search use deterministic code-based or rule-based verifiers, whereas Web uses a fixed LLM verifier to check predefined requirements for components and values. Human assessment of representative tasks across all five environments yields an average rating of 9.15 out of 10. Across 32 model, the best achieves an unweighted macro-average success rate of 81.39%. Reasoning configurations achieve higher macro-average success rates in all 13 paired comparisons under the reported mode-specific inference settings. We find that world model performance is positively correlated with agent performance. Context truncation experiments on two models further show that early history is important for ToolUse, Terminal, and Search, whereas retaining only recent interactions improves accuracy over full histories in Game. These findings suggest that reliable simulation depends on preserving task-relevant state and limiting interference from irrelevant history. Code and sample data are available at https://anonymous.4open.science/r/worldsimbench/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.