acceptodds
Under review as a conference paper at ICLR 2027

Mitigating the Need for Bellman Completeness: Multi-Fragment Trajectory Stitching for Decision Transformer

Abstract

The Decision Transformer (DT) has achieved promising results in offline reinforcement learning by reframing it as sequence modeling. However, DT fundamentally struggles with trajectory stitching—the ability to combine trajectory fragments into globally superior trajectories. This limitation stems from two intertwined factors: (i) DT is a supervised learning model whose policy space is confined to the support of the training data distribution, and (ii) offline datasets, especially suboptimal ones, often lack coverage of the compositional trajectories that stitch together high-return fragments. In contrast, dynamic programming (DP)-based approaches, such as Q-learning, naturally support trajectory stitching through value iteration. However, they rely on Bellman completeness and accurate function approximation, which can be challenging to satisfy in practice. To bridge this gap, we propose Multi-Fragment Trajectory Stitching (MFTS), a diffusion-based data augmentation framework that constructs synthetic trajectories by stitching high-return fragments from suboptimal data. MFTS stitches high-return fragments together by generating forward and backward state trajectories with a pre-trained conditional diffusion model, then fusing them via a time-aware asymmetric blending strategy. The resulting bridge is completed with actions and rewards through auxiliary models. MFTS thus grants DT the benefits of dynamic programming while mitigating the need for Bellman completeness. Extensive experiments on the D4RL benchmark demonstrate that MFTS achieves superior or competitive performance in most settings, with the largest gains on challenging medium-replay datasets where stitching matters most.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.