Coding4Games: Can Modern Coding Agents Play Modern Video Games?
Abstract
Games, especially modern AAA games, challenge language models to perceive complex scenes, control actions, and make decisions over long horizons. Existing game benchmarks primarily test language models as computer-use agents that query a language model throughout play, incurring repeated inference cost and, in real-time games, acting while the environment continues to evolve. We consider the abilities of the large language model via the aspect of modern coding agents: they can utilize tools and skills, develop programs that subsequently play without online model inference. We thus introduce *Coding4Games*, a benchmark for evaluating whether coding agents can develop standalone programs that play modern games through visual observations and standard game inputs. It has three defining features: *1) Diverse Modern Games*, spanning visually and behaviorally varied tasks; compared with classic games, these games offer richer visuals, more complex states, and broader action spaces, placing greater demands on perception, control, and long-horizon decision-making. *2) Development–Test Separation*, in which a coding agent uses a sandbox and generic GUI interaction to iteratively develop a program, after which the selected program is frozen and independently tested without language-model inference; and *3) Outcome and Process Evaluation*, which measures not only game performance and reliability but also how agents use feedback, revise programs, and resolve problems during development. The resulting programs outperform computer-use agents that query a model at every step in some games, yet remain far below human performance. We additionally examine the effects of development budgets. Coding4Games evaluates whether coding agents can turn their understanding of complex interactive environments into programs that work reliably on their own. Project page: [Coding4Games](https://anonymous.4open.science/w/Coding4Games-Page-Review-DE42/)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.