Constructive Softmax Training for Certified In-Context Policy Improvement
Abstract
We establish a constructive connection between finite noisy training and policy improvement with frozen Transformer parameters. For finite discounted Markov decision processes sharing a state–action space and discount factor, we train a structured Transformer with standard softmax attention to reproduce an explicit temporal-difference teacher update. A two-stage projected curriculum of designed probes learns three update coefficients and the free attention scores from arbitrary initialization within prescribed parameter boxes. The resulting uniform operator-error bound controls repeated evaluation as target policies and value estimates change using the same deployment record. An outer greedy policy-improvement loop repeatedly calls the frozen evaluator, reusing a single trajectory collected by a fixed behavior policy. Under the stated coverage, mixing, and training conditions, the complete procedure returns an -optimal policy with a high-probability certificate bounding its suboptimality. At fixed confidence and structural and conditioning constants, sufficient budgets are labeled operator examples and unique deployment transitions, where bounds label noise. Training iterations, evaluation depth, and improvement rounds each grow logarithmically in . A two-task lower bound matches the deployment accuracy exponent when pretraining provides no information about the new task. Controlled experiments recover the teacher update from multiple initializations, show that training accuracy matters for repeated evaluation, and demonstrate policy improvement on unseen tasks. The RMS policy-uniform teacher Bellman residual decays approximately as .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.