What Does Key Quantization Actually Change in a Delta Rule? The Step Size.
Abstract
Delta-rule linear recurrent neural networks such as DeltaNet, Gated DeltaNet and DeltaProduct update a matrix state with one least-mean-squares step per key. This step shrinks the residual it corrects only if the effective step size lies in , where is the step size and is the squared norm of the key. The models keep with a key normalizer, but a quantizer placed after the normalizer breaks this: with 4-bit integer (INT4) keys, deviates from one by up to . In models with step sizes below , every update then has a relative step-size error of . In models with step sizes up to , of the steps have , where the update flips and enlarges the residual. If and are independent, the share of such steps depends only on the tail of near and the distribution of , and measuring the two separately predicts the observed shares within for INT4 keys. NormFix divides the step size by the squared norm of the key that the kernel receives, . This restores the trained residual factor for every and makes the propagation of state errors nonexpansive for any quantizer. With 4-bit weights and INT4 keys, NormFix recovers of the long-context state-tracking loss of a Gated DeltaNet with negative eigenvalues and of the loss of a standard one, at less than extra kernel time. On the negative-eigenvalue model it is better than clipping, skipping, key renormalization and 8-bit keys on every state task and on retrieval. Its gain holds across six models and three scale axes, and where the key is normalized after the quantizer, so that , it changes no score by more than points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.