acceptodds
Under review as a conference paper at ICLR 2027

GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games

Abstract

Powered by expert guidance, embodied agents can operate in interactive environments; however, it is unclear whether they can learn autonomously and continually from their own experience. To evaluate such self-improvement methods, we introduce a two-legged benchmark: , the first testbed for agentic self-improvement in complex video games. evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. At test time, agents must complete short-horizon execution tasks that evaluate their ability to navigate, interact, and engage with game-specific mechanics (e.g., combat) in potentially unseen games. Out-of-the-box frontier models complete fewer than 50% of the 500 tasks due to failures in multimodal grounding, establishing that self-improvement methods have considerable room to push performance. We demonstrate that contemporary approaches to self-improvement are lacking, with standard implementations of world modelling and autonomous skill discovery failing, and a novel strategy that uses curiosity-based exploration to write guides achieving only partial success. tests end-to-end game completion in two fan-made Pokémon games. We show that while frontier models have been pre-exposed to official releases such as Pokémon Red, they lack essential information on the games in our testbed. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough. We show that a sophisticated agentic pipeline with multimodal memory and hierarchical subgoals fails to reach even the first major milestone in both games, establishing GameBoyWorlds as an ambitious target for self-improving agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.