What Does Attention Retain? Information Loss and Sufficient Moments for In-Context Regression
Abstract
Success on Gaussian in-context regression can conceal information loss in a context representation. We show that, for any fixed moment budget, smooth noise laws can remain arbitrarily close to Gaussian and match its moments through twice the prescribed degree exactly, yet make the best decoder of that budget arbitrarily asymptotically inefficient. Each noise law is fixed before the sample-size limit. The mechanism is an exact characterization of the statistical experiment retained by complete mixed moments of covariates and labels: for centered isotropic absolutely continuous designs under the stated conditions, its local information is the squared norm of the noise score projected onto polynomials of one lower degree. This strong experiment limit covers every measurable decoder. Full local information is retained exactly for polynomial log noise densities of admissible degree; there, the same moments are sufficient at every sample size. Common-source affine attention realizes the degree-two boundary, while a scalar source-side score operation reaches the full-data endpoint. We connect the theory to standard causal softmax Transformers using independent risk evaluations and a task-preserving equal-Gram diagnostic. Reversible model drift witnesses a finite-context Bayes penalty for every Gram decoder, independently of a fitted control's quality. Fresh evaluations give positive numerical witnesses under Laplace and quartic noise, with a Gaussian negative control. The results separate computational performance from the statistical adequacy of a context representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.