Rethinking Algorithmic Explanations of Transformer In-Context Learning: A Systematic Study of Noisy Linear Regression
Abstract
Transformer in-context learning (ICL) is increasingly being used to learn statistical procedures, making otherwise costly or impractical inference accessible. Their reuse across datasets makes it important to establish how faithfully they approximate the intended statistical procedure as the inputs change. Nevertheless, constructions showing that Transformers can implement familiar algorithms, together with agreement on selected distributions and isolated mechanistic tests, do not establish that trained models inherit those algorithms’ generalisation properties. We address this gap by systematically studying Transformers trained to perform noisy linear regression in-context. We identify the main features organising their generalisation and how they respond to changes in pretraining, finding substantial competence beyond the training curriculum, but also significant sensitivity to input changes that leave the optimal prediction unchanged. Comparisons with the algorithms proposed in recent literature, supported by theory and targeted interventions, show that the trained Transformers violate structural properties those algorithms satisfy. The resulting characterisation supplies constraints for explanations of statistical ICL and insight for designing models that preserve the properties their intended use requires.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.