CURE: Contrastive Pretraining for Escaping Suboptimal Entrapment in Skill Discovery
Abstract
Unsupervised skill discovery enables agents to learn diverse and transferable behaviors without external rewards. Distance-maximizing skill discovery aligns representation spaces with state distances, which enables the agent to achieve strong exploration. However, its effectiveness in environments with irreversible failure modes (e.g., a humanoid robot falls and is unable to recover) is limited by the suboptimal entrapment issue: The coupled optimization of representation and policy causes the encoder to overfit to early-failure transitions, which assigns these transitions high intrinsic rewards, and thereby traps the policy in absorbing states. In this paper, we validate the suboptimal entrapment issue through theoretical analysis and empirical experiments. To address this issue, we propose CURE, a simple yet effective method that injects a robust geometric prior into the representation space. Before the joint optimization begins, CURE pretrains the state encoder on a minimal set of unlabeled trajectories using the InfoNCE contrastive objective. We prove that this pretraining is asymptotically equivalent to performing principal component analysis (PCA) on the transition dynamics, naturally favors forward-progressing transitions over failure ones, and successfully resolves the entrapment issue. Extensive experiments demonstrate that CURE significantly mitigates suboptimal entrapment, achieves broader state coverage, faster convergence, and robust behavioral priors for downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.