acceptodds
Under review as a conference paper at ICLR 2027

Reachable Information Gain: Task-Relevant Goal Curricula for Sparse-Reward Long-Horizon Reinforcement Learning

Abstract

Sparse-reward long-horizon reinforcement learning requires agents to discover intermediate behaviors that make distant rewards reachable. Existing methods often select subgoals by novelty, controllability, or reachability, but such signals may not reveal progress toward final success, especially in tasks with distractors. To address this limitation, we introduce Reachable Information Gain (RIG), an evaluation criterion for candidate progress signals in sparse-reward long-horizon reinforcement learning. RIG is motivated by an ideal latent progress criterion, but uses observable multi-step progress over task predicates or goal potentials because latent task progress is not directly identifiable from observations alone. We instantiate RIG as an online teacher for goal-conditioned off-policy learning. The teacher learns context-dependent utilities from delayed multi-step evidence and supplies curriculum rewards and replay priorities without requiring successful trajectories. Experiments across sparse-reward manipulation, navigation, and discrete-control benchmarks show that RIG improves both sample efficiency and success rate over strong baselines, demonstrating its effectiveness across diverse long-horizon tasks. The code is available at https://anonymous.4open.science/r/RIG.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.