GAMEASG-BENCH: Benchmarking Autonomous Software Generation for Game Development
Abstract
Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements. We introduce GameASG-Bench, a benchmark that makes behavioral testability part of the generation task for game development. Our design declares an evaluation interface specification before generation, fixing legal starting scenarios, player-level actions, stable snapshots, rejection behavior, and invariants while leaving private implementations open. Concretely, we include: (i) static L1 checks that assess source-level compliance; and (ii) browser-executed L2 checks that combine semantic observations with real input and runtime evidence. We implement this protocol as 47 browser-native game-generation tasks spanning 12 primary genres and both 2D and 3D interaction, each with executable checks and an independently verified reference implementation. Experiments show that our predefined interface enables more accurate and efficient behavioral evaluation. We evaluate nine agent stacks on this benchmark and find that the highest observed mean L2 check pass rate is 93.2%, yet the highest observed strict task success rate, requiring all L1 and applicable L2 prerequisite and core requirement checks, is only 55.3% (26/47 tasks). Further comparisons examine the effects of tool access and harness choice. Together, these evaluations support the predefined-interface design and reveal task-level compliance gaps that high average check pass rates can obscure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.