acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Policy-Reward Mismatch in Skill Discovery

Abstract

Unsupervised skill discovery (SD) aims to enable agents to explore in environments without task rewards, where a key factor determining their utility on downstream tasks is the joint state space coverage of all learned skills. However, we find that even state-of-the-art methods cannot sufficiently cover the state space. Gaps between skills formed in the early learning stage are never filled, which significantly reduces the utility of skills on downstream tasks. In this paper, we provide an in-depth theoretical analysis of this dead-zone phenomenon. We identify that a core issue is policy-reward mismatch: the intrinsic reward mismatches the policy’s exploration needs. When the policy initially explores a new region, the skill discriminator is dominated by early data and fails to update the reward for the new region in time, which suppresses further exploration. To address this issue, we propose RASD, a simple yet efficient method that provides more on-policy data for the policy and skill discriminator to ensure quick reward adaptation. Extensive experiments show that our method effectively improves the state-space coverage of skills and achieves better performance on downstream tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.