Counting Capabilities, Not States: Capability Representations for Exploration in Open-Ended Worlds
Abstract
Intrinsic exploration bonuses reward novelty, and novelty is nearly always measured over states. In open-ended environments with deep achievement hierarchies, we find that the choice of novelty space is the dominant effect. We hold the standard count bonus fixed and vary only the space it counts. Every state-derived space (pixels, hashed features, prediction error) stays near the no-bonus floor, because appearance novelty is abundant yet orthogonal to progress: an agent can sightsee forever without gaining a capability. Counting the agent's capability configuration instead, the binary vector of which items it holds, read from the environment's inventory for counting only, turns the same bonus into a climbing signal: it is worth +15.7 points on a DreamerV3-class agent that already observes its achievement state, and +0.5 on the same agent stripped of that input and its crafting-action prior. With that scaffold, the agent sets a new state of the art among agents trained from scratch on Crafter (67.7 vs. the prior best 58.1 at 10M steps) and raises collection of the deepest milestone, the diamond, from 0.5% to 22% of episodes. A ladder of single-flag ablations attributes every point of the improvement and uncovers a second phenomenon: without extended replay retention the agent forgets its own breakthroughs: a rare deep-tier trajectory leaves the default buffer within steps, and holding it four times longer halves the relapses and adds 6.9 points. The same bonus lifts success on sparse-reward Meta-World pick-place from 0.22 to 0.83, and it is inert on dense-reward Meta-World and on Craftax, where the reward is already dense enough to learn from. Curiosity, in progression-structured worlds, should be about what you can do, not what you see.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.