acceptodds
Under review as a conference paper at ICLR 2027

Skill-Flow: Learning Skill-based Hierarchical Policies via Flow Models for Offline Goal-Conditioned Reinforcement Learning

Abstract

We address the problem of learning policies for long-horizon goal-conditioned tasks with sparse rewards from offline datasets. Such problems are difficult for two primary reasons: sparse rewards in a long-horizon setting hinder value function learning, and limited data induces out-of-distribution (OOD) errors during policy learning. We tackle these challenges using a skill-based hierarchical framework with flow policies at two levels. The high-level policy selects temporally extended skills to guide long-horizon planning, while the low-level policy executes actions conditioned on the selected skill. Temporally extended skills alleviate sparse rewards by enabling the learning of goal-conditioned value functions, while hierarchical decomposition reduces error accumulation over long horizons. To mitigate OOD errors, we learn behavior-cloning flow policies over skills and actions, then steer their bounded input noise using learned value functions towards high value regions. This construction keeps the resulting policies within the support of the learned behavior-cloning policies. We evaluate our approach on the OGBench benchmark and demonstrate superior performance over state-of-the-art hierarchical and value learning based baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.