acceptodds
Under review as a conference paper at ICLR 2027

A Systematic Study of Functional Attention Heads in Language Models

Abstract

A Transformer spreads its computation over many attention heads, and interpretability work now ties individual heads to specific functions. Induction heads copy repeated patterns, retrieval heads fetch information from long contexts, and truthfulness heads separate true from false answers. These labels are used as if they transferred across models, settled early in training, and could be turned up at will. Each label, though, was found in one model, with one detector and one intervention. This paper tests the assumptions together. We use six head-level signals, five base models, two Pythia pretraining trajectories, and matched-control interventions on Qwen, Llama, and Mistral. One answer recurs: a repeatable signal fixes a head type's coarse structure, not its fine identity. (1) Distribution: local-pattern heads precede induction heads in every model. No complete layer order is shared. (2) Attribution: induction and retrieval detectors pick largely the same heads, 20 and 19 of 30 on Qwen and Llama. The retrieval cost of an “induction” mask falls on those shared heads, not on either label's exclusive heads. (3) Formation: detector strength saturates thousands of steps before top-head membership settles. Fine-tuning changes membership without moving layer distributions. (4) Control: patching the selected retrieval heads restores most of the retrieval loss, and the labels improve KV-cache allocation. Doubling a head's output does not improve its task. A label therefore transfers three things: the relative order of head types, the task contribution of the selected set, and a ranking for memory allocation. It does not transfer absolute depth, an exclusive function, head membership, or gains from amplification. Reuse a label's coarse structure, and re-derive its membership and coefficients for each model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.