State-Value Transformers with In-Sample Bellman Backups
Abstract
Recent advances in Transformer-based critics have shown the benefits of scaling value functions in reinforcement learning (RL). However, these methods rely on specialized architectures or attention entropy control to scale effectively. In this paper, we identify the value learning objective as a key factor in the effectiveness of Transformer-based critics and demonstrate that in-sample learning improves performance without attention entropy control. Building on these findings, we propose State-Value Transformer Implicit Q-Learning (VTIQL), a method that combines a Transformer-based *V*-function with an MLP-based *Q*-function under in-sample learning. Extensive numerical experiments on OGBench demonstrate that VTIQL consistently outperforms prior offline RL methods, achieving the best performance on **7** out of 9 environments. Notably, VTIQL exhibits high performance on robotic manipulation tasks and maze navigation tasks with high-dimensional states, whereas prior methods struggle to achieve high performance in both task categories. VTIQL also performs strongly with a compact Transformer-based -function and remains robust across a wide range of model sizes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.