GameCode-Bench: Benchmarking Repository-Scale Game Development with Coding Agents
Abstract
As large language models rapidly evolve, the capabilities of autonomous coding agents are advancing to new limits. Recent software engineering benchmarks, however, primarily evaluate conventional software projects, while game benchmarks focus on generating miniature, tutorial, or demo-level games. Agents' capacity to handle mature, highly interactive game systems, therefore, remains largely underexplored. To address this gap, we introduce GameCode-Bench, a benchmark for repository-scale gameplay engineering comprising 86 tasks across eight independent, open-source browser game repositories. Our benchmark captures complementary developer- and community-facing development workflows, and agents operate directly within mature codebases containing a median of 162,055 non-test source lines, with the largest exceeding one million. Alongside frontend, backend, and API integration maintenance tasks, GameCode-Bench centers on core game logic and challenges models to localize relevant implementation paths, infer intended behavior, and modify existing systems while preserving surrounding functionality. Across eight evaluated models, we find that failures concentrate in state integration, conditional logic, and incomplete behavioral branches; trajectory analysis further shows that additional exploration or testing is beneficial only when accurately targeted. Evaluated against hidden, deterministic, human-reviewed scoreable variants, GameCode-Bench provides a rigorous testbed for measuring and diagnosing coding-agent capabilities in complex game development.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.