Reward-Free Exploration through Calibrated Ordering
Abstract
Reinforcement learning agents are often guided towards exploratory behavior with the help of intrinsic rewards. Maximizing the entropy of the state visitation distribution is a principled choice for such a reward. The objective pays the agent , where is the visitation density. However, density estimation in high-dimensional spaces is unreliable. Furthermore, the shape of the latent space from which distances are measured is crucial, as it determines the semantic meaning of proximity, yet there is little methodology for selecting or validating such spaces, and we are left relying on assumptions about their semantic structure. We introduce CALOR (CALibrated ORdering), which extracts from the latent space only a pairwise ordering among states, thereby relying less on the exact geometry of that space. A predictor network is trained to reproduce that ordering with a pairwise ranking loss, and at the same time to place its own outputs on a calibrated scale, which restores the reward magnitude that an ordering alone leaves undetermined. The construction accepts any source of pairwise rarity comparisons, which we demonstrate by adding a second ordering derived from how long the agent has stayed alive. On reward-free Craftax-Classic at one billion interactions, CALOR reaches a Crafter score of against for the strongest baseline, and unlocks every one of the achievements in every seed at least once. On three continuous-control exploration tasks it matches or exceeds every baseline, and on the humanoid maze it visits five times as many distinct states as the strongest of them. Ablations show that the ordinal signal and its calibration each contribute, and that neither on its own accounts for the result.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.