When is Composition Trustworthy? Grounded Value Composition for Goal-Conditioned RL
Abstract
Long-horizon goal-conditioned reinforcement learning increasingly relies on composition: estimating how to reach a distant goal by chaining reachability through intermediate points such as subgoals, waypoints, or prerequisite goals. Yet no existing method gives a general criterion for when a composed estimate can be trusted. We show that composing through an intermediate always yields a lower bound on optimal reachability, and that the error of any composed estimate splits into exactly three terms: the error in reaching the intermediate, the error in continuing from it, and an arrival-mismatch term, the divergence between the distribution at which the continuation is evaluated and the states the agent actually reaches after arriving at the intermediate. This decomposition reveals a structural obstacle: a goal predicate specifies what has been achieved, but not the state distribution from which continuation begins. Transitive methods ask which intermediates are safe to compose; their answer is states on the same trajectory. We ask a different question: under what distribution must the continuation value be correct? The answer is the arrival distribution induced by actually achieving the intermediate. This yields a single design rule: compose only through intermediates whose arrival distributions are supported by data. We instantiate this rule as GRAFT (Grounded, Arrival-Faithful Transitivity). For online GCRL, GRAFT reads an all-goals critic at the states where the agent actually arrives after commanding each goal, and commands an intermediate only when the resulting grounded lower bound certifies that the direct estimate is underestimated. For offline GCRL, GRAFT extends transitive value learning across trajectories through crossings, data-supported states where trajectories meet, so that every factor in a cross-trajectory target stays in support. On Craftax and OGBench, GRAFT substantially improves over prior methods, with gains concentrated in the prerequisite-heavy goal families where composition is most needed, and ablating arrival-grounding removes the gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.