Interpreting Multi-layer Transformers for In-context Linear Regression with Varying Covariance
Abstract
We study how multi-layer softmax-attention transformers perform in-context linear regression in the challenging regime where the covariate distribution varies across sequences, in contrast to the fixed-distribution setting considered in prior work. We first show that conclusions from prior work break down in this regime, and that transformer depth becomes essential for achieving low in-context error. Through a population-limit interpretation of the trained transformer, we reveal that it solves linear regression by implementing a variant of gradient descent whose parameters depend strongly on model depth, corresponding to a consistent structure in the learned weight matrices. We also discover the non-convergent property of the learned algorithm. Building on this insight, we show that with a chain-of-thought-style intermediate step, the transformer can solve in-context instrumental variable (IV) regression, applying the same algorithm whenever the same regime studied here appears in the IV regression. Our findings provide new evidence that multi-layer transformers learn distinct in-context algorithms under more complex regression scenarios, bridging the gap between empirical performance and interpretable algorithmic behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.