acceptodds
Under review as a conference paper at ICLR 2027

VGBench: Can Agent Schedule a Real-Engine Game Development?

Abstract

Game generation is a challenging long-horizon task for modern coding agents. Existing game-generation benchmarks mostly measure the final artifact, leaving the development process unmeasured. In practice, the final artifact is only part of the result: an agent must also coordinate goals at different levels and produce useful intermediate builds. A sound process builds a playable core early, adds dependent features on top of it, and preserves earlier behavior as the game becomes more complex. We introduce VGBench to evaluate this ability through two complementary settings. In the single-round setting, an agent receives a task with many requirements and a fixed development budget, while the benchmark embeds several intermediate checks; the score depends on the quality of the staged deliverables produced at these checkpoints. In the multi-round setting, the benchmark progressively adds requests and measures whether the agent can complete each new stage while maintaining the stability of existing functionality. Both modes use multiple rubrics for evaluation. An instrumentation agent prepares game states that are difficult to reach, and a player agent records the evidence used for scoring. VGBench contains 107 natural-language queries and nine stage increments, covering 17 game types and multiple harnesses. Evaluations of frontier coding agents show that end-to-end game generation remains challenging: the strongest multi-round configuration reaches 84.82 overall, while two thirds of the configurations remain below 72. Further analysis shows that agents with similar final game scores can differ sharply in their ability to deliver useful intermediate results, and this difference is reflected in the clarity and extensibility of the resulting games. VGBench provides a way to study long-horizon agents at the process level beyond final products.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.