Trajectory-Aware Value Guidance with a Co-Trained Critic for Offline RL
Abstract
Offline reinforcement learning (RL) faces the significant challenge of deriving optimal policies from static, often suboptimal datasets. Sequence modeling methods like Decision Transformer (DT) have shown promise, but their performance is often constrained by the quality of the best trajectories in the dataset. In contrast, value-based methods have the potential to stitch together suboptimal trajectories to achieve performance beyond that of any single trajectory. Recent work has explored integrating DT with value-based approaches to enhance trajectory stitching capabilities while avoiding distribution shift in offline RL. However, existing methods typically rely on static pre-trained Q-functions, which can become misaligned with the evolving policy during training, limiting their effectiveness. We introduce Adaptive Q-Guided Decision Transformer (AQDT), an approach that dynamically co-trains the Q-function alongside DT training, creating a mutually adaptive relationship between policy and critic. AQDT adaptively balances DT learning with value-based guidance based on trajectory quality, while ensuring that the critic remains aligned with the evolving policy. On standard D4RL benchmarks, AQDT outperforms prior DT and hybrid methods, achieving a 4.6% gain in average normalized return on locomotion tasks and demonstrating robust generalization across datasets of varying quality. The code implementation is available at :Anonymous.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.