acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Unsupervised Skill Discovery: A Trajectory-Level Perspective

Abstract

Unsupervised skill discovery enables reinforcement learning agents to acquire reusable behaviors without relying on laborious, task-specific reward engineering. However, existing MI-based approaches typically discriminate skills at the level of individual states. Consequently, this state-centric formulation often forces skills to partition the environment and specialize in disjoint regions, preventing agents from learning temporally extended behaviors that generalize across varying initial conditions and making downstream skill chaining highly challenging. In this work, we propose TIME (Trajectory-level Information Maximization and Exploration), a novel framework that discovers reusable skills by modeling the temporal evolution of agent-generated trajectories rather than isolated state occupancy. TIME introduces a conditional MI objective that allows all skills to access the full state space while encouraging them to induce distinguishable long-horizon behaviors. Furthermore, to overcome the limited state coverage of pure MI methods, TIME incorporates a simple yet effective Lipschitz-constrained training mechanism that prevents the exploitation of trivial differences. Evaluated across a diverse set of challenging continuous-control environments, TIME learns coherent and highly reusable behaviors while outperforming prior unsupervised skill discovery methods, yielding substantial performance gains on downstream tasks through hierarchical skill chaining. Code is available at https://anonymous.4open.science/r/TIME-32C3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.