Discovery of Sticky and Responsible Options for Frozen-Option Transfer
Abstract
A fundamental problem in hierarchical reinforcement learning is finding a small number of reusable and easily composable behaviors that transfer with little or no fine-tuning to solve novel tasks. Option and skill discovery frameworks are two prominent approaches that address this problem. However, option discovery typically focuses on learning specialized behaviors for a training task, leading to a low diversity of behaviors with limited transferability, and unsupervised skill discovery learns diverse behaviors that can be misaligned with downstream tasks. To address these issues, we introduce Task-Agnostic Sticky and Responsible Options (TASRO), an option-learning framework that combines (1) Markovian option switching at every time step, (2) an objective promoting persistent, diverse, and specialized options, and (3) information asymmetry, where the high-level controller observes both task and task-agnostic information while the low-level options observe only task-agnostic information. We evaluate under a strict frozen-option transfer protocol: pretrain options on a goal-reaching training task, freeze them, and relearn only the controller on downstream tasks. Across our navigation, locomotion, and manipulation benchmarks, TASRO improves frozen-option transfer over the tested skill-discovery and task-exposed option baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.