acceptodds
Under review as a conference paper at ICLR 2027

EGGS: Entropy-Guided Joint Goal Selection for Long-Horizon Unsupervised Goal-Conditioned Reinforcement Learning

Abstract

We consider unsupervised goal-conditioned reinforcement learning, where the agent aims to learn a policy that can reach diverse goals without task-specific rewards. To this end, existing methods select goals using visitation novelty or predicted exploration value to guide the agent toward new states, but when several goals are used in the same collection batch, they may be close to one another, leading to overlapping exploration across rollouts. We introduce EGGS (Entropy-Guided Joint Goal Selection), which jointly selects goals to reduce redundancy across exploration rollouts. EGGS combines a visitation novelty score motivated by visitation entropy with a mutual-information term that measures distinguishability among the selected goals in a learned temporal representation. Each selected goal guides a Go-Explore rollout, which first moves toward the goal and then continues exploration from the resulting state. Compared with selecting goals using visitation novelty alone, EGGS selects more distinguishable goals and collects trajectories with less overlap while retaining high novelty scores. In controlled 2-D environments, EGGS achieves the highest mean coverage AUC in all six tasks and the highest mean final normalized visitation entropy in five of six, demonstrating rapid coverage and broadly distributed visitation. Across AntMaze and locomotion benchmarks, EGGS attains the highest mean coverage AUC in four of six environments, and its pretrained behaviors support zero-shot goal reaching and sequential multi-goal control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.