Compiled Agency: Coding Agents as Game AI Researchers — from a Roguelike to StarCraft II and Civilization
Abstract
Coding agents are increasingly capable of sustained engineering and empirical research. Greater autonomy makes their research decisions themselves a target for evaluation: what to investigate, which experiments to run, and when to stop. Games provide a controlled, interactive setting with objectively measurable outcomes. We introduce Gauntlet, a develop–freeze–evaluate protocol for studying these decisions as agents build standalone game-playing programs. Starting from a game description, a raw observation/action interface, and an empty policy file, an off-the-shelf agent develops a controller from bare interaction in one autonomous session. We call the capability under study compiled agency: the shipped program plays with zero model calls. Every intermediate version is frozen for later scoring on held-out instances, connecting the agent's research process to independently measured performance. The capability is real and advancing: environment access adds 10 to 78 percentage points of held-out success over construction-only controls, and progress across model generations comes in steps, with tiers that defeat one generation entirely falling to the next. The protocol yields StarCraft II controllers that defeat every fair built-in AI, and in the newest generation the strongest cheating tier as well, and Civilization controllers that win complete games by conquest. Replaying more than 5,000 frozen versions exposes the research behind the programs: gains that plateau early; rigorous local investigation beside sparse validation of what actually ships; and stops that follow a race between the agent's own evidence and a model-specific transcript budget it was never given. When validation panels are refreshed mid-session, so that only the evidence changes, shipped success rises 12.4 points in nine of nine completed pairs of twelve initiated. The experiments an agent designs are part of the capability it delivers; Gauntlet makes them measurable and improvable—a step toward agents whose research practice, not just whose code, can be engineered. We release the benchmark, the corpus, and the replay tooling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.