acceptodds
Under review as a conference paper at ICLR 2027

Put Your Skills To Work: Exploration via Random Walks in Skill Space

Abstract

Self-supervised pre-training in reinforcement learning aims to discover a diverse set of behaviours without extrinsic reward, which can then accelerate a wide range of downstream tasks. Recent skill discovery methods such as METRA and CSF learn skills as directions in a latent space and produce impressive behaviours in locomotion domains. However, the exploration of existing skill learning methods degrades in complex environments, and existing remedies add machinery on top of skill learning, such as an explicit skill tree or a learned guide policy. Leveraging the latent space that METRA/CSF already learn, we instead change only how skills are executed during training. Rather than committing to one skill per episode, we chain skills of a fixed horizon, resampling the latent direction at each boundary, so that a rollout performs a random walk whose steps are skill executions rather than primitive actions. Applying this to CSF gives Chained-CSF (CCSF), which leaves the base algorithm, its losses and its networks untouched and adds only the skill horizon as a hyperparameter. On a bottleneck maze, skill chaining increases state coverage from to . On Craftax-Classic, a Minecraft-like domain with a deep crafting tree, CCSF unlocks of the achievements per episode without task rewards, the highest we are aware of for an unsupervised RL method. Removing the survival pressures that end episodes early lets it climb further, obtaining an iron pickaxe (depth 6; of episodes) and diamonds (depth 7; ), while a baseline without skill switching does not reliably collect stone (depth 3). Without any reward or additional learning, the skills learned during pre-training can be used to solve downstream tasks: goals specified as a change in the raw observation are converted into latent directions and commanded from the frozen policy, reaching every stage up to iron in over of episodes and doing so several times faster than random skill selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.