Returning Skill Discovery
Abstract
Unsupervised skill discovery (USD) learns a repertoire of distinct behaviors without external reward, and across every major approach the learned skills are one-way: a skill leaves its starting state and stops wherever it ends, because nothing in the objective rewards coming back. We first show this is structural and then introduce Returning Skill Discovery (RSD), a method-agnostic add-on that makes skills return without changing the base objective. RSD adds three pieces: a phase input, without which a skill that retraces its path is provably inexpressible; a reward weight that flips sign at mid-episode, under which the optimal skill travels out for half the episode and home for the other half; and a dimensionless return constraint that catches everything the weighted reward cannot see. The weight is the same for every method, and a checkable property of the reward decides only what it multiplies. Paired with five USD methods (DIAYN, DADS, LSD, CSD, METRA) on locomotion and manipulation, with state and pixel inputs, RSD produces skills that return to their starting position and posture, chain zero-shot, and outperform one-way libraries as long-horizon options. Videos of returning skills are at https://sites.google.com/view/rsd-usd.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.