acceptodds
Under review as a conference paper at ICLR 2027

In-context classification: Calibration and noise in linear attention with softmax readout

Abstract

We study in-context binary classification in a simple attention model trained to minimise a logistic loss. Each prompt has a latent logistic teacher, and the learner must infer its direction from Bernoulli-labeled context examples. We derive an asymptotic prediction for the test logit in a high-dimensional regime with context length and total training prompts . Our formula cleanly separates calibration, test-time context noise, and non-isotropic parameter-estimation noise as interpretable effects on in-context performance. Our analysis also identifies the ridgeless separability regime and addresses the effects of context length, pretraining-set size, regularisation, and sampling temperature on in-context classification performance. The theory's predictions closely agree with finite-dimensional simulations and fully trained linear-attention models, while experiments with standard softmax attention show qualitatively related behaviour beyond the regime assumed by our analysis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.