acceptodds
Under review as a conference paper at ICLR 2027

Agent's GameDev Exam

Abstract

Evaluating AI agents for game development requires testing whether they can translate requirements into working gameplay within complex engine environments. We introduce Agent's GameDev Exam, a production-oriented benchmark spanning Unreal Engine 5, Unity, Godot, and RPG Maker, with an estimated inventory of 448 candidate tasks from 55 source projects and packages. Two complementary tracks assess component-level feature development within existing projects and End2end-level game creation from supplied assets, functional specifications, and reference screenshots. Our construction pipeline combines expert-defined requirements, LLM-assisted task preparation, and structural and in-engine validation; component-level tasks remove complete feature implementations while preserving the surrounding project. Evaluation combines deterministic checks with ground-truth-referenced video assessment to test requested functionality and, where applicable, preservation of existing behaviors. Across nine model–harness configurations, the current evaluation yields a best component-level full-pass rate of 14.23%, while no configuration achieves a full pass on the End2end-level track. Analysis of UE5 component-level submissions identifies incomplete execution paths and disconnected state-driven effects as the dominant issue categories, accounting for 48.3% and 18.3% of categorized issues, respectively. These findings expose a gap between locally plausible implementations and complete feature delivery, positioning Agent's GameDev Exam as a testbed for improving system-level reasoning, integration, and runtime verification in game-development agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.