acceptodds
Under review as a conference paper at ICLR 2027

The Second Layer of Linear Self-Attention Implements a Context-Adaptive Update

Abstract

Recent work interprets transformer depth as iterations of an optimisation algorithm. We demonstrate that context-adaptive computation first occurs at layer two in linear self-attention. We prove that one-layer models cannot implement context-adaptive updates, that two-layer models can, and that, for every finite depth, there are task distributions where context-adaptive updates achieve strictly lower risk than any composition of the same number of fixed optimisation steps. These results identify the first layer where linear self-attention can condition its update rule on the observed context rather than only the training distribution. Depth therefore enables more than additional iterations of a fixed optimisation algorithm. We introduce a framework for measuring update behaviour in trained transformers by comparison with reference optimisation algorithms. The framework is validated on models with known mechanisms and applied across a range of transformer architectures and in-context learning tasks. One-layer models match the best single fixed optimisation step. Two-layer models show a context-adaptive correction, which disappears when the second layer is ablated. Our framework identifies the known two-layer induction transition from the update kernel in associative recall and extends to in-context classification and softmax attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.