acceptodds
Under review as a conference paper at ICLR 2027

From Policy Evaluation to Off-Policy Control: Provable Emergence of In-Context Q-Learning

Abstract

Transformers can implement reinforcement-learning algorithms in context, yet how off-policy control emerges within Transformers remains poorly understood. We study Q-learning, a foundational model-free control algorithm that combines off-policy bootstrapping with greedy policy improvement. We prove that a two-layer attention-only Transformer can implement a finite-temperature form of Q-learning, where the temperature controls the approximation to greedy action selection, recovering window Q-learning as . We train the structured Transformer on randomly generated MDP tasks using output-only supervision. Experiments show that the theoretically predicted parameter structure emerges during training, including the internal greedy action-selection mechanism, and that the trained Transformer achieves performance comparable to Q-learning on unseen tasks. Our theoretical analysis further shows that the equivalent realizations form a solution manifold and that, under suitable task-diversity conditions, the training iterates converge locally to this manifold at an exponential rate. For the random-task model analyzed here, the local worst-case number of gradient steps needed to reduce the distance to the manifold by a factor is , quantifying the increasing optimization cost as the learned update approaches window Q-learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.