Repairing Attention Collapse Through the Loss with a Spectrally Gated Penalty
Abstract
Attention collapse can stall transformer training without a numerical failure: attention concentrates on a few tokens and the loss stops improving. We ask whether a term in the loss can bring a run back from collapse. This is not obvious, because a penalty on attention entropy acts through the softmax gradient, and that gradient becomes small as attention saturates. We propose SpecGate, an attention-entropy penalty whose weight is set for each head by a gate that follows the spectral sharpness of the head's attention scores and passes no gradient of its own. In a 124M-parameter language model, SpecGate brings runs from an established collapse to 0.07 nats above the healthy run of the same seed, against 1.05 nats without treatment, and raises the loss of healthy runs by 0.004 nats. The same penalty at constant weight, including a separately tuned one, ends 0.12 to 0.14 nats above SpecGate. Starting SpecGate at the first spectral alarm rather than after collapse has set in lowers the final loss by 0.22 nats, and stopping it 250 updates after the start of repair raises the final loss by 0.10 nats relative to continuous treatment. At equal numbers of updates, a gate on the entropy deficit and a fixed schedule are equivalent to the spectral gates within 0.02 nats; across the tested interventions, when and how strongly the penalty acts moved the final loss more than the signal that sets its weight. After collapse, QK-Clip, which rescales the weights, ends 0.03 nats above SpecGate at its published threshold; with a threshold tuned on this task for the final loss, it ends 0.02 nats below SpecGate and runs faster.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.