acceptodds
Under review as a conference paper at ICLR 2027

LEARNING WHICH CHANGES MATTER: PREFERENCE-CONDITIONED SKILL DISCOVERY

Abstract

Unsupervised Skill Discovery (USD) aims to autonomously learn a diverse set of reusable skills without relying on extrinsic rewards. However, existing methods mainly encourage skills to be diverse or distinguishable, which can result in behaviors that are easy to differentiate but lack behavioral significance. We hypothesize that leveraging the extensive knowledge encoded in pretrained large language models (LLMs) can help identify behaviorally consequential skills based on their executions. In this sense, we introduce Preference-Conditioned Skill Discovery(PCSD), a skill discovery framework that incorporates task-agnostic preference guidance to discover skills with greater behavioral significance. PCSD decomposes skill executions into factor-specific behavioral changes and learns preferences separately for each factor to avoid bias arising from their different physical roles. PCSD defines a preference-conditioned mutual-information objective to prioritize behaviorally consequential skill executions, and provides a variational interpretation of this objective under a induced distribution. Across multiple environments, we show that PCSD assigns higher preference to coherent and behaviorally consequential skill executions. PCSD consequently learns skill repertoires that are more useful for downstream tasks and achieves improved performance over existing skill discovery baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.