When Are Linear Transformers Nearly Ridge-Optimal? Task-Varying Noise and Covariance Adaptation for In-Context Regression
Abstract
Test-time demonstrations for in-context learning may have noise levels that are shared within a task yet vary across tasks. We study this heterogeneity in in-context linear regression through a tractable one-layer linear-attention transformer. Under isotropic tasks, the population-optimal estimator reduces to scalar shrinkage of the empirical input-output cross-moment. We derive exact finite-sample bias–variance formulas and the optimal population scaling, then quantify two limitations of this rule. First, population tuning does not adapt to a new task’s noise level. Although average risk depends only on the mean noise variance, the dispersion of excess conditional risk is exactly proportional to the variance of task noise, with a coefficient determined by normalized estimation variance. In the proportional limit, this coefficient peaks below the interpolation threshold and decreases throughout the overparameterized regime. Second, scalar shrinkage cannot adapt to the realized empirical covariance. We decompose its risk gap to optimally tuned ridge into a squared-bias reduction and a variance increase, and derive rates showing that the normalized additive gap vanishes in data-rich, highly overparameterized, and weak-signal regimes. Experiments with multi-layer softmax-attention transformers show partial recovery of the scalar-to-ridge gap where covariance adaptation offers substantial gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.