acceptodds
Under review as a conference paper at ICLR 2027

Exploration by Maximizing the Entropy of Random-Network Residuals

Abstract

State-entropy maximization is a standard objective for reward-free exploration: it rewards states that are rare among those the current policy visits, so collecting diverse data is the objective itself. But the objective equates explored with visited. A rarely visited state stays rewarding after the agent's models have fit it, and a frequently visited one loses its reward while they still fail there. We propose to maximize the entropy of prediction residuals instead of states. We use the residual vector of random network distillation (RND), the difference between the outputs of a fixed random network and of a predictor trained to match it on the agent's experience, and apply to it the nearest-neighbour entropy estimator of state-entropy methods. The residuals of states the predictor has fit are small and lie close together, so these states earn little reward. The new objective still rewards diversity, now measured in what the predictor has yet to fit. We show that the resulting reward keeps RND's preference for large residuals and adds a preference for residuals that are rare among those the current policy produces. We implement the method in a world-model agent and evaluate it in the reward-free setting on a range of continuous-control tasks. It is competitive with the state of the art and nearly always significantly improves over state-entropy maximization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.