Guiding Skill Discovery with Foundation Model
Abstract
Unsupervised skill discovery removes the need for hand-designed rewards, but maximizing behavioral diversity alone readily produces undesirable or dangerous skills: a cheetah robot learns to roll in every direction when we would prefer it to run without flipping or entering hazardous areas. We propose Foundation model Guided (FoG) skill discovery, which incorporates such preferences without reward functions or demonstrations. A preference is stated once, in language; a foundation model turns it into a binary label on states zero-shot, by writing the labeler as a short program from a description of the state vector or by scoring raw pixels against the sentence. Starting from a trajectory-level Wasserstein-dependency formulation, we show how this label enters the skill-discovery reward as a multiplicative gate through the critic of the Kantorovich-Rubinstein dual, and prove that the resulting objective lower-bounds the trajectory-level dependency. Across state- and pixel-based benchmarks, FoG eliminates flipping and rolling, avoids hazardous areas, discovers behaviors that are hard to specify explicitly, and yields skill libraries that adapt quickly and safely to downstream goals. Interactive visualizations are available from https://sites.google.com/view/submission-fog.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.