GamePilot: A Game Agent That Learns How to Observe and Act from Its Own Play
Abstract
Human players progress through unfamiliar games by connecting individual actions and using earlier experience to decide what to do next. Existing game agents often simplify this process through access to game-state APIs, pauses during inference, or repeated task attempts. We study game playing in the wild, where agents interact with a continuously running game through its graphical user interface and learn from their own play to sustain progress. We propose GamePilot, a game agent that learns how to observe and act from its own play. It places a fixed-weight vision-language model inside the GamePilot Harness, editable code that controls what the model observes and how its decisions reach the game. While play continues, the GamePilot Evolver localizes failures in execution records, admits a revised harness only after it passes offline replay tests, and selects which harness version plays next by its estimated advantage in the state it inherits. GamePilot Bench evaluates agents in the wild on four games, pairing one-hour Milestone campaigns that measure human-time-weighted progress with Atomic and Basic tasks that test general interaction and game knowledge. With Gemini 3.6 Flash, GamePilot reaches 1.85–2.00× the Milestone progress of the strongest baseline in every game. Revising how the agent observes and acts, together with game knowledge, adds 10.6–11.4 percentage points over revising game knowledge alone. Across three backbones, the final GamePilot harnesses receive the highest mean expert ratings in 10 of 12 Atomic and all 12 Basic game–backbone settings. Supplementary videos: https://vocal-sawine-d5fe2d.netlify.app/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.