acceptodds
Under review as a conference paper at ICLR 2027

Intermediate-Goal-Informed Scorer for Long-Horizon Goal-Conditioned RL

Abstract

Long-horizon goal-conditioned reinforcement learning faces a central challenge in reliably distinguishing which primitive actions are useful for distant goals. As the goal becomes more distant, action-dependent differences can become weak relative to estimation noise, making primitive-action discrimination increasingly difficult. Prior work has addressed this challenge through hierarchical methods using intermediate goals and flat methods learning long-horizon values or goal-reaching relations. However, hierarchical methods introduce additional complexity, while accurate long-horizon estimation in flat methods does not necessarily ensure effective primitive-action discrimination. To address this challenge, we propose the Intermediate-Goal-Informed Scorer (IGIS), which learns a more discriminative representation for primitive-action scoring. IGIS constructs a goal-conditioned action scorer from state-action and state-goal representations and strengthens the distant-goal representation by distilling shorter-horizon information from intermediate-goal representations along offline trajectories. This enables more effective discrimination among candidate actions while requiring neither subgoal generation nor hierarchical execution at test time. Across challenging long-horizon locomotion and manipulation benchmarks, IGIS shows consistent gains over strong baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.