acceptodds
Under review as a conference paper at ICLR 2027

When Does Depth Matter For In-Context Learning? Adaptive Inference in Deep Transformers

Abstract

Transformers perform computations through many successive attention and feedforward blocks, allowing them to learn complex correlations between a large collection of coupled variables. When does stacking successive attention-feedforward computations over many layers provide a computational advantage over a single transformer block? We address this question by examining in-context learning in generalized linear attention transformers. We first introduce a general theory of distributed inference in such transformers, subject to constraints on communication and depth. We show that such systems can exploit internal representations (`function vectors') to infer a latent context variable at increasingly finer scales over its layers. For an in-context linear regression task, the theory predicts that while one-layer transformers without feedforward blocks are optimal for Gaussian priors over the context variable, multi-layer transformers are superior for non-Gaussian, tree-like priors. Trained linear attention transformers reproduce quantitative predictions from the theory. Using causal key-patching experiments, we verify that function vectors in intermediate layers mediate adaptive routing of information. Our results suggest that depth and feedforward blocks enable transformers to implement adaptive inference, and this is advantageous when the distribution over latent variables has hierarchical structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.