acceptodds
Under review as a conference paper at ICLR 2027

Controllable Skill Discovery with a Skill-Linear Critic

Abstract

Unsupervised skill discovery (USD) learns diverse behaviors without task rewards, but those behaviors are not necessarily controllable. Existing methods can obtain diversity through small state changes or varied locomotion directions, yet varying their skills can move state properties together or switch them among a few extremes. We ask whether the skill vector can provide a controllable interface for bounded state descriptors such as joint angles. We use controllable to mean that one descriptor can be varied gradually (graded control), other descriptors can be held approximately fixed (isolated control), and multiple descriptor targets can be achieved together (composable control). We propose the Skill-Linear Critic (SLC), which retains a mutual-information-based intrinsic reward but restricts skill dependence to a linear combination of skill-independent basis critics. This turns the skill vector from an arbitrary conditioning input into a coefficient vector for a skill-independent family of critics. We add a pairwise orthogonality penalty between the basis embeddings to mitigate observed critic-training instability. In static-pose evaluations of four simulated robots, SLC yields graded, isolated, and composable joint-angle behavior on HalfCheetah, Ant, and Humanoid. It is also the only evaluated learned pipeline with a reliable affine fit in a majority of evaluations on every platform, although this all-platform conclusion is threshold-sensitive for Franka.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.