acceptodds
Under review as a conference paper at ICLR 2027

A Bayesian Perspective on the Role of Epistemic Uncertainty for Delayed Generalization in In-Context Learning

Abstract

Learning to generalize from finite data is a central problem in machine learning. In-context learning provides a concrete solution to this problem while exhibiting a striking phenomenon known as delayed generalization: transformers can perfectly fit observed task-input combinations long before they can infer a held-out task from in-context examples. Although prior work has substantially advanced our understanding of this phenomenon, the evolution of uncertainty during this transition remains comparatively less explored. We study this question in a modular-arithmetic benchmark with latent linear tasks, using Improved Variational Online Newton (IVON) and a Last-Layer Laplace approximation to estimate posterior-predictive epistemic uncertainty (EU). Across different levels of task diversity and training-input coverage, both posterior approximations show lower epistemic uncertainty in configurations that generalize to held-out tasks. Furthermore, we show that EU follows a non-monotone pattern of rise-peak-decline during training, with its peak preceding the late increase in accuracy, in held-out tasks. Finally, we attempt to formalize this behavior with a simplified Bayesian linear model that reproduces the qualitative dynamics, showing that delayed generalization and non-monotone uncertainty can arise from slow spectral modes of the training-data covariance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.