acceptodds
Under review as a conference paper at ICLR 2027

Local Evidence, Joint Allocation: Planning Guidance for Goal-Conditioned Reinforcement Learning

Abstract

Planning can accelerate sparse-reward goal-conditioned reinforcement learning by using subgoal-conditioned behavior to regularize policy updates. Yet assigning guidance weights independently across samples couples their relative priorities with the overall balance between planning supervision and return maximization. We propose Support-Proportional Guidance (SPG), which treats guidance strength and sample allocation as a joint decision within each minibatch. A critic compares planning-induced actions with the current policy's sampled action to obtain a support count for each state-goal pair. The prevalence of sufficient support sets the mean guidance weight, while capped proportional allocation distributes the corresponding total according to support strength. The actor retains the complete planning-induced action mixture as its reference, so critic evidence changes guidance strength rather than selecting individual target actions. We establish that weak local support limits an individual guidance weight and that one changed candidate comparison has bounded influence on the entire allocation. Across six navigation and manipulation tasks, SPG improves sample efficiency over the evaluated baselines. Further results show that both excessively weak and excessively strong guidance reduce the success rate and that SPG outperforms general sample allocation methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.