A Bayesian Perspective on Task Retrieval and Task Learning in In-Context Learning on Markovian Data
Abstract
In-context learning enables pretrained transformers to adapt from a prompt without parameter updates, but it remains unclear when this behavior arises from retrieving memorized pretraining tasks (i.e., task retrieval) versus learning a rule from in-context data (i.e., task learning). We study this question through a finite-mixture Markov data model, where each latent Markov chain represents a pretraining task. Our analysis identifies the ratio of prompt length to the number of pretraining tasks as a key factor. We first establish quantitative guarantees showing that the finite-prior Bayesian predictor associated with the pretraining tasks favors retrieval at large ratios and learning from the prompt at sufficiently small ratios. We further show that, at small ratios, well-trained transformers make predictions close to those of the finite-prior Bayes predictor. At larger ratios, our experiments show increasing deviations from Bayes, with transformers exhibiting a stronger relative preference for learning. Our findings clarify how prompt length and pretraining diversity jointly shape in-context learning and when Bayesian prediction explains transformer behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.