acceptodds
Under review as a conference paper at ICLR 2027

Interpretability-Guided Design from a Head-Binding Failure in Grokking

Abstract

Transformers often generalise long after fitting their training data, a delayed transition known as grokking. Yet on some tasks the transition never comes, and it has been unclear what fails inside a model that does not generalise. We study two-layer transformers trained on math tasks, such as taking a derivative, whose answer combines several inputs. There the failure lies in variable binding: which input each first-layer head attends to. On the derivative task, most models that fail have one head bound to the constant term, which differentiation ignores, and no model that succeeds has one. Ablating that head costs far less than ablating the others, because the input it reads plays no part in the answer. Transplanting only a generalising model's embedding and readout layers, 12.5% of its weights, repairs all six failed models we tried. With the inputs shuffled, all 19 bound heads read their input through a learned tag instead of its position. The wrong binding is usually set within the first few hundred of fifty thousand steps, as each head's attention sharpens onto one input. SoftBind keeps the first layer's attention soft over the first 500 steps, so no head locks onto an input before the task can steer it. On a sum task this more than doubles the grok rate, and keeping the softening for the whole run raises it from 26.2% to 85.2%. Interpretability here does more than explain a failure: it can tell us where to intervene.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.