acceptodds
Under review as a conference paper at ICLR 2027

CURE: Contrastive Pretraining for Escaping Suboptimal Entrapment in Skill Discovery

Abstract

Unsupervised skill discovery enables agents to learn diverse and transferable behaviors without external rewards. Distance-maximizing skill discovery aligns representation spaces with state distances, which enables the agent to achieve strong exploration. However, its effectiveness in environments with irreversible failure modes (e.g., a humanoid robot falls and is unable to recover) is limited by the suboptimal entrapment issue: The coupled optimization of representation and policy causes the encoder to overfit to early-failure transitions, which assigns these transitions high intrinsic rewards, and thereby traps the policy in absorbing states. In this paper, we validate the suboptimal entrapment issue through theoretical analysis and empirical experiments. To address this issue, we propose CURE, a simple yet effective method that injects a robust geometric prior into the representation space. Before the joint optimization begins, CURE pretrains the state encoder on a minimal set of unlabeled trajectories using the InfoNCE contrastive objective. We prove that this pretraining is asymptotically equivalent to performing principal component analysis (PCA) on the transition dynamics, naturally favors forward-progressing transitions over failure ones, and successfully resolves the entrapment issue. Extensive experiments demonstrate that CURE significantly mitigates suboptimal entrapment, achieves broader state coverage, faster convergence, and robust behavioral priors for downstream tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.