Beats and Aliases Can Break Rotary Length Extrapolation
Abstract
Rotary position encodings degrade beyond their training length, and a common account blames planes whose period exceeds the training window. We show that a spectrum whose every period is below 13 tokens can still collapse: what is missing is not another frequency but a difference of frequencies. For fixed content a rotary logit is a trigonometric polynomial of the offset, and in that fixed-kernel problem classical super-resolution bounds say that gaps above about let the training offsets determine it stably, while a gap with can raise the worst-case error near the half-beat offset by a factor . In a 0.46M-parameter character-level model, twelve pairs detuned by cycles per token make training on 512 tokens collapse after position 1,024; across a sweep of eight beat periods the onset tracks the beat, and wherever a model has failed, stretching only the gaps at inference removes 52–86% of the damage while an equal shift that keeps them removes at most 9%. Both recur at 3.3M parameters, after ten times longer training and on a second corpus, although in the first two the collapse begins earlier than the beat predicts. In spectra of harmonics of coprime periods, 59% of the strongest out-of-window attention weights fall where the Chinese remainder theorem predicts (base rate 4%), and masking them helps at least three times more than random masking. For RoPE's own geometric spectrum gaps and periods rank the planes alike, our experiments do not separate them, and a beat-specific prediction fails. Spectral arithmetic is an additional diagnostic for rotary extrapolation, computable before training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.