Cross-Trajectory Value Composition for Offline Goal-Conditioned Reinforcement Learning
Abstract
Offline goal-conditioned reinforcement learning (GCRL) learns goal-reaching policies from fixed datasets. In long-horizon tasks, the data may contain useful parts of a solution across different trajectories without containing a complete path from the start to the goal. Standard temporal-difference learning propagates value information through local one-step backups, so supervision from a distant goal must pass through many updates, making long-horizon values slow and difficult to learn. Existing structured methods address this limitation but often require specialized value models or compose only segments from the same trajectory. We propose Cross-Trajectory Value Composition (CVC), which samples intermediate waypoints across trajectories and uses a discounted triangle relation to combine two shorter goal-reaching problems. This divide-and-conquer approach propagates value information over long distances without imposing a specialized structure on the value function. We incorporate CVC into a hierarchical offline GCRL framework and evaluate it on 24 long-horizon OGBench datasets. CVC achieves the best or tied-best performance on 20 tasks, learns substantially faster than prior methods, and performs especially well on stitching and noisy datasets. These results show that cross-trajectory composition can efficiently learn long-horizon behaviors from fragmented offline experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.