acceptodds
Under review as a conference paper at ICLR 2027

When Are Linear Transformers Nearly Ridge-Optimal? Task-Varying Noise and Covariance Adaptation for In-Context Regression

Abstract

Test-time demonstrations for in-context learning may have noise levels that are shared within a task yet vary across tasks. We study this heterogeneity in in-context linear regression through a tractable one-layer linear-attention transformer. Under isotropic tasks, the population-optimal estimator reduces to scalar shrinkage of the empirical input-output cross-moment. We derive exact finite-sample bias–variance formulas and the optimal population scaling, then quantify two limitations of this rule. First, population tuning does not adapt to a new task’s noise level. Although average risk depends only on the mean noise variance, the dispersion of excess conditional risk is exactly proportional to the variance of task noise, with a coefficient determined by normalized estimation variance. In the proportional limit, this coefficient peaks below the interpolation threshold and decreases throughout the overparameterized regime. Second, scalar shrinkage cannot adapt to the realized empirical covariance. We decompose its risk gap to optimally tuned ridge into a squared-bias reduction and a variance increase, and derive rates showing that the normalized additive gap vanishes in data-rich, highly overparameterized, and weak-signal regimes. Experiments with multi-layer softmax-attention transformers show partial recovery of the scalar-to-ridge gap where covariance adaptation offers substantial gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.