acceptodds
Under review as a conference paper at ICLR 2027

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

Abstract

Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction—treating environment steps as a single transition—restores this signal at long range, but no single fixed suits all state-goal distances: large preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on . GITA trains one policy by aggregating advantage-weighted supervision across multiple values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single . On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed- method, OTA, by 7 percentage points (14% relative).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.