acceptodds
Under review as a conference paper at ICLR 2027

Doubly Dual-Weighted Flow Matching for Hierarchical Offline Goal-Conditioned RL

Abstract

Offline goal-conditioned reinforcement learning (GCRL) learns goal-reaching policies from static datasets without further environment interaction. We propose *Doubly Dual-Weighted Flow Matching* (DFM), a hierarchical goal-conditioned flow policy that mitigates distant goals by decomposing long-horizon goal reaching into a high-level module that proposes nearer subgoals and a low-level flow module that generates actions conditioned on those subgoals. To mitigate sparse rewards, we introduce a doubly dual-weight mechanism that derives two optimal correction ratios from a single reward-tilted target distribution: a state–action weight ratio, obtained from a primal–dual Bellman-flow program for low-level policy extraction; and a state weight ratio, recovered through an auxiliary off-policy evaluation objective, that reweights high-level subgoal selection. On OGBench, a benchmark covering navigation, manipulation, and combinatorial-puzzle tasks, DFM achieves higher average success rates than the evaluated GCRL baselines on long-horizon goal-reaching tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.