acceptodds
Under review as a conference paper at ICLR 2027

ReStitch: Offline Trajectory Stitching via Retrieval-Based Policy Improvement

Abstract

Return-Conditioned Supervised Learning (RCSL), notably Decision Transformers (DT), has emerged as a stable paradigm for offline reinforcement learning. However, these methods can struggle to stitch suboptimal trajectories, limiting their performance in sparse-reward environments. While recent methods incorporate value-based guidance into DTs, they typically represent policy improvement through conditioning signals or optimization objectives rather than explicit trajectory-level continuations. In this work, we propose ReStitch, which formulates retrieval as an instance-based mechanism for offline policy improvement. ReStitch retrieves trajectory segments that are relevant, feasible, and advantageous, and injects them as explicit state-action guidance through cross-attention. Experiments on D4RL show substantial gains on sparse-reward AntMaze tasks, while improvements are more limited on dense-reward MuJoCo control. Mechanism analyses show that these gains require an informative value representation and the joint use of relevance, feasibility, and advantage in retrieval. These results highlight both the effectiveness and the empirical capability boundary of retrieval-based trajectory stitching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.