CodeWorldBench: How Well Can Coding Agents Build Executable Agent Worlds?
Abstract
Modern coding agents increasingly demonstrate the capability to scale and generate complex interactive agent environments. However, existing benchmarks evaluate the ability to generate repositories from scratch or repair localized bugs, leaving its capacity to construct fully functional, stateful environments largely unmeasured. Therefore, we introduce CodeWorldBench, a benchmark designed to assess the quality of code-driven world construction. We construct 118 environments spanning common workflows and professional software domains, requiring agents to generate or repair repositories that realize a stateful world, including its data, tool interfaces, and transition logic. Rather than relying on code similarity, we test behavioral equivalence: submitted and reference worlds start from the same seed, receive hidden tool calls, and are compared on their outputs and resulting states. Our evaluation reveals vulnerabilities in current systems. While models preserve already-correct behavior, they struggle to construct environments from scratch and fail at repairing broken behavior. Agents fail on state transitions and validation order rather than syntax. Because agent-built environments increasingly serve as the grounding for training subsequent models, an environment that runs but transitions incorrectly supplies a damaging false signal. Uncovering these flaws provides insights and directions for robust Recursive Self-Improvement and the self-evolution of future agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.