MIM-DICE: Unsupervised Skill Discovery from Arbitrary Off-policy Experience
Abstract
Unsupervised reinforcement learning (URL) aims to acquire reusable behaviors without external supervision for rapid adaptation to downstream tasks. Within URL, skill learning achieves this by discovering skills whose induced state distributions are mutually distinct yet collectively provide broad state coverage. A common way to induce these properties is to maximize the mutual information (MI) between skills and policy-induced states. However, existing methods entail a state distribution mismatch between the off-policy replay data and the current policy. This leads to two critical limitations in practice. First, as we empirically show, their biased estimation of MI impedes effective MI maximization. Second, they strictly rely on replay data collected by skill-conditioned policies, precluding skill learning from arbitrary offline datasets. We introduce Mutual Information Maximization Distribution Correction Estimation (MIM-DICE), a DICE-based method that corrects this distribution mismatch, enabling off-policy unbiased and effective MI maximization. Furthermore, by eliminating the reliance on skill-conditioned data, MIM-DICE enables skill learning directly from arbitrary offline datasets. In online off-policy experiments on tabular environments and Pendulum, MIM-DICE better maximizes MI under the current policy than prior methods. Across six MAZE domains, it learns diverse skills directly from arbitrary offline datasets, achieving higher MI and inducing more clearly separated state-visitation regions. Finally, MIM-DICE outperforms a broad range of state-of-the-art baselines on the URLB benchmark suite.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.