Interpretability-Guided Design from a Head-Binding Failure in Grokking
Abstract
Transformers often generalise long after fitting their training data, a delayed transition known as grokking. Yet on some tasks the transition never comes, and it has been unclear what fails inside a model that does not generalise. We study two-layer transformers trained on math tasks, such as taking a derivative, whose answer combines several inputs. There the failure lies in variable binding: which input each first-layer head attends to. On the derivative task, most models that fail have one head bound to the constant term, which differentiation ignores, and no model that succeeds has one. Ablating that head costs far less than ablating the others, because the input it reads plays no part in the answer. Transplanting only a generalising model's embedding and readout layers, 12.5% of its weights, repairs all six failed models we tried. With the inputs shuffled, all 19 bound heads read their input through a learned tag instead of its position. The wrong binding is usually set within the first few hundred of fifty thousand steps, as each head's attention sharpens onto one input. SoftBind keeps the first layer's attention soft over the first 500 steps, so no head locks onto an input before the task can steer it. On a sum task this more than doubles the grok rate, and keeping the softening for the whole run raises it from 26.2% to 85.2%. Interpretability here does more than explain a failure: it can tell us where to intervene.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.