What Attention Control Can Reach: Exact Output Geometry
Abstract
Editing where a transformer’s attention heads look is now a standard control, and the same patterns are read as evidence of what it computes. Both rest on an unchecked assumption: that pushing a head’s attention produces the intended change in its output. For one head at a fixed input we settle an edit’s reach and its cost. Reach is fixed before any budget is spent: a bias shared by a group of keys moves each row of the head’s output only inside the hull of one average value vector per group, a region the head’s values and unedited attention fix. Every interior point is reached by a closed-form bias, a target outside stays out at any budget, and a linear readout’s largest displacement follows from one number per group on a single unedited pass. Cost is charged only to edits that cross the softmax: a row’s distance from uniform and its first-order response to logit edits obey a sharp product bound, response vanishes as a row becomes deterministic, and a finite edit pays along its whole path, not at its start. In two live models the excluded directions do not move beyond the float32 floor of the head’s own output; in 7 pretrained transformers the upper tail of trained rows sits nearer the bound than at random initialization; and on one model’s pre-registered held-out split the routing law read at the spent budget, not differentiated at zero, orders real interventions its derivative misorders more often than not. An attention edit can therefore be screened before it is run; neither its reach nor its cost is legible in the pattern itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.