Agent's GameDev Exam
Abstract
Evaluating AI agents for game development requires testing whether they can translate requirements into working gameplay within complex engine environments. We introduce Agent's GameDev Exam, a production-oriented benchmark spanning Unreal Engine 5, Unity, Godot, and RPG Maker, with an estimated inventory of 448 candidate tasks from 55 source projects and packages. Two complementary tracks assess component-level feature development within existing projects and End2end-level game creation from supplied assets, functional specifications, and reference screenshots. Our construction pipeline combines expert-defined requirements, LLM-assisted task preparation, and structural and in-engine validation; component-level tasks remove complete feature implementations while preserving the surrounding project. Evaluation combines deterministic checks with ground-truth-referenced video assessment to test requested functionality and, where applicable, preservation of existing behaviors. Across nine model–harness configurations, the current evaluation yields a best component-level full-pass rate of 14.23%, while no configuration achieves a full pass on the End2end-level track. Analysis of UE5 component-level submissions identifies incomplete execution paths and disconnected state-driven effects as the dominant issue categories, accounting for 48.3% and 18.3% of categorized issues, respectively. These findings expose a gap between locally plausible implementations and complete feature delivery, positioning Agent's GameDev Exam as a testbed for improving system-level reasoning, integration, and runtime verification in game-development agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.