GameEngineBench: Full-Session Verification of Coding Agents in a Networked Game Engine
Abstract
Coding agents face new challenges in game development, where they must produce code that is not only syntactically correct but also consistent with engine-managed lifecycle, state, and networking behavior. Existing game-development benchmarks, however, do not combine assertions over engine-based runtime states with cross-client consistency, failing to evaluate how well generated code adheres to the desired behavioral specifications. To address these limitations, we present GameEngineBench, a benchmark of 110 C++ implementation tasks derived from 9 existing Unreal Engine 5 projects. Each task requires an agent to implement missing behavior within an established gameplay architecture and a restricted editable scope. We also introduce Full-Session Verification, a testing framework that injects hidden tests after the agent finishes and evaluates the resulting implementation inside networked Play-in-Editor (PIE) sessions. The tests evaluate Unreal's game loop, actor lifecycle, subsystems, and replication stack while directly inspecting runtime state. Across 17 benchmark tasks, tests additionally inspect a distinct client world to determine whether the authoritative server state is correctly replicated. To our knowledge, this is the first coding benchmark to evaluate cross-client replication. A detailed analysis of 768 mutated implementations finds that 74% of distributed-property bugs are detected only by end-to-end tests involving full PIE sessions. Across model–harness configurations, the strongest achieves 55.5% pass@1, while 31 tasks remain unsolved by every evaluated configuration. These results demonstrate that compilation and shallow verification overestimate agents' ability to implement behavior that remains correct during full engine execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.