Scaling Contrastive Reinforcement Learning to Open-Ended Environments
Abstract
Developing generally capable agents requires scalable methods for learning a wide range of goal-directed behaviours. Contrastive reinforcement learning (CRL) offers a promising approach by reframing goal-conditioned value learning as self-supervised classification of an agent's own future states. We extend CRL to open-ended environments, where agents pursue semantic goals that represent desired outcomes rather than particular target states. We formalize these goals as predicates over the state space, which makes CRL's training labels computable in closed form, and derive a Rao-Blackwellized estimator of the resulting discounted goal-occupancy measure. We further find that this estimator is biased when an agent's actions can terminate the episode, as it implicitly conditions on the episode continuing and fails to penalize actions that cause termination. We address this bias with an absorbing class and termination-discounted targets that account for the probability mass beyond the episode end. The resulting critic serves as an effective teacher for a policy-gradient student. On Craftax, our method achieves comparable success to an agent trained with a per-goal Q-learning teacher across thousands of semantic goals while using up to less training compute. These results demonstrate that CRL can scale to large semantic goal spaces while accounting for irreversible failures in open-ended environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.