EXPAND: Extensible Policy Capacity for Unsupervised Skill Discovery
Abstract
Unsupervised reinforcement learning pretrains agents without task rewards for downstream learning and adaptation. Active exploration broadens state coverage but often lacks a reusable skill interface; skill discovery learns controllable behaviors whose coverage can plateau. Fixed policy capacity may limit exploration even as intrinsic objectives and representations improve. We introduce **EX**tensible **P**olicy c**A**pacity for u**N**supervised skill **D**iscovery (EXPAND), a framework that adds skill-conditioned cascade stages during pretraining. Stages train jointly under the original objective; gates conditioned on state and skill recursively combine action distribution parameters. EXPAND increases METRA's final historical coverage by **71.1%** on average across seven environments and generally outperforms parameter-matched wider, deeper, and parallel gated policies. EXPAND also improves coverage with LSD and DIAYN, with gains varying by environment and metric. EXPAND also improves METRA's overall zero-shot goal reaching and hierarchical control with frozen skills. Applying the same architecture to LSD improves coverage in three environments. The skills learned with EXPAND applied to METRA also improve overall performance in zero-shot goal reaching and hierarchical control with frozen low-level skills. These results show that organizing policy capacity as a cascade improves exploration coverage and supports downstream skill reuse in the evaluated settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.