acceptodds
Under review as a conference paper at ICLR 2027

Improving Bayesian In-Context Learning in Markov Transformers

Abstract

A transformer can enter a generalizing in-context learning regime while remaining far from the corresponding optimal predictor. What prevents it from closing this residual gap? We study this question in Markov-chain prediction, where the Bayes posterior is available exactly and the learned attention circuit can be inspected. We introduce Symmetry-Aware In-context Pre-training (SAIP) for two-block transformers. SAIP guides the formation of previous-token attention early in training and lightly regularizes predictions toward the closed-form posterior throughout training. Inference uses the two-block model without evaluating auxiliary losses. Across 18 task-pool/seed settings, SAIP reduces mean KL divergence to the bigram posterior by 89.3%, with at least a 69.8% reduction in every setting. It also strengthens previous-token attention and moves the empirical generalization boundary to smaller task pools. Ablations attribute most of the accuracy gain to posterior regularization, while attention measurements reveal a sharper previous-token circuit. These results show how training can improve the accuracy of Bayesian in-context prediction while changing the attention circuit that implements it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.