EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
Abstract
LLM-based environment simulators offer a flexible substrate for agent training, but their usefulness depends on whether generated feedback and state updates faithfully reflect agent actions. In practice, LLM-simulated environments suffer from hallucinations, logical inconsistencies, and silent state drift—failures that corrupt agent reward signals and compound the construction costs that the paradigm was designed to eliminate. To address this gap, we propose EnvSimBench with four contributions: (1) We provide the first formal definition and operationalization of Environment Simulation Ability (EnvSim Ability) as a quantifiable research objective. (2) We construct EnvSimBench, a rigorous benchmark covering 400 samples across 167 diverse environments, equipped with verifiable labels and fine-grained difficulty stratification along three axes. (3) Evaluations of seven frontier language models reveal a pronounced decline in Config Match on state-changing operations relative to state-preserving operations, which we term the state-change cliff. This finding identifies a limitation in predicting the state updates required by the evaluated environments. (4) We design a constraint-driven simulation pipeline that substantially reduces hallucination, boosts environment synthesis yield by 6.8%, and cuts costs by over 90%. Overall, EnvSimBench serves as both a diagnostic framework and a practical optimization path for reliable LLM-based environment simulation, establishing a solid foundation for scalable agent training. Code and data are available at \url{https://anonymous.4open.science/r/EnvSimBench-
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.