When Routing Must Be Right: Irreversible Lock-In in Attention Routing
Abstract
Hybrid sequence models pair a cheap recurrent path with a small budget of attention, and something has to decide which tokens get it. On multi-query associative recall that decision is worth three orders of magnitude: spending a fixed 6.25% budget on the right tokens gives 0.997 accuracy, a static schedule tuned to the same budget gives 0.0002. Learned routers are notoriously hard to train, and the usual explanation, that top- selection is discontinuous, so unselected candidates get no gradient, is incomplete. REINFORCE supplies unbiased gradient to unselected tokens and still lands at chance, and one router that had been correct and stable for 25,000 steps produced no improvement at all. We argue the variable that matters is timing, and measure it. Holding the number of correctly-routed steps fixed and varying only their position, misrouting the first 8,000 steps leaves accuracy at chance while delaying the identical window by fifty steps recovers almost everything; the failing arm is the one with more time to recover afterwards. Ten seeds at each of nine onsets put the half-way point at 23 steps, and a sweep over the deficit's length puts its threshold at 48, two measurements agreeing on the scale from opposite directions, and settling the outcome inside the first 0.06% of training. The obstruction is that the signal a router needs does not exist when it is needed: the marginal utility of attention is under misrouting and 2.72 under correct routing. We probe four ways for a usable substitute and none survives, because each returns the model's existing reliance on attention, which is itself a product of the routing it was trained under. Two explanations that would make this an artifact are ruled out by measurement: learned routing stays at chance with four times the budget, and a four-block model in which attention feeds later positions locks in more readily than one block, not less. We also give a mechanism, and a pre-registered prediction it makes about learning-rate scaling that we test and falsify.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.