Learning in a World That Remembers: When Past Performance Shapes Future Reward
Abstract
An agent can improve its general competence by directing itself towards what it can learn but has not yet learned. Yet an agent may continue to receive reward for repeating previously learned behaviour, without any incentive to broaden its competence. In this paper, we introduce the ArcadeWorld problem, where an agent chooses both its actions within individual environments and which environments in a suite to engage with. The agent’s goal is to develop general competence by improving across the entire suite within a single lifetime. We propose leaderboard rewards to formalize this goal: the agent’s past performance is remembered, so that improvement is rewarded while previously attained performance receives diminishing rewards. To understand the challenges of learning with leaderboard rewards, we conduct experiments on a simple abstraction of ArcadeWorld where continual improvement is possible. We find that deep reinforcement learning algorithms can fail to match the performance of simple memory-bounded agents, with exploration being a key challenge. To investigate these challenges at a larger scale, we introduce AtariWorld, where an agent can play and switch among Atari games. Standard deep reinforcement learning algorithms similarly struggle in AtariWorld and are even outperformed by exploration methods that receive no external reward based on score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.