Adaptive Value Anchoring for Long-Horizon Stitching in Offline Reinforcement Learning
Abstract
An important capability of offline goal-conditioned reinforcement learning is stitching, which can improve policies by connecting suboptimal trajectory segments and enable agents to solve long-horizon tasks by combining shorter trajectories. Stitching commonly relies on temporal difference (TD) learning, whose repeated bootstrapping propagates reward signals between trajectories but can also lead to bias accumulation in value estimation. Although longer backups can mitigate bias accumulation by reducing bootstrapping depth, we show that they can hinder stitching. Specifically, longer backups can skip junction states where reward signals propagate between trajectories. Therefore, we propose Adaptive Value Anchoring (AVA), which adapts backup lengths to bootstrap from junction states. AVA can thereby mitigate bias accumulation while preserving reward propagation between trajectories. We evaluate AVA on long-horizon OGBench tasks that require stitching. AVA achieves the highest average success rate among the compared methods. Moreover, it achieves the highest mean success rate in every subsampled setting, where reduced trajectory overlap makes stitching harder.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.