acceptodds
Under review as a conference paper at ICLR 2027

Rotary Attention in the Frequency Domain: Value Rotation as a Distance Filter

Abstract

Rotary position embedding (RoPE) encodes position in the attention score but leaves the values unrotated, so what a head reads carries no record of where it came from. We study the variant that also rotates the values and inverse-rotates the head output (QKVO-RoPE), which adds no parameters. We show that this turns each attention head into a filter over distance. Content that a head reads from a single position passes and keeps a phase that records its distance, whereas content that it averages over many positions partly cancels. To test this account, we train matched Transformer pairs from 130M to 1.3B parameters that differ only in the rotation and measure the cancellation in every head. We find that value rotation helps where one position must be read against a diffuse context. Beyond the training length the larger rotated models retrieve needles that their unrotated counterparts miss entirely, and up to 760M the rotated models extrapolate with lower perplexity. We also find that the costs fall where the analysis predicts, in heads that average and in associative recall, whose source lies at a distance set by the query, while the cost at the training length is small. A larger rotary base reduces the cancellation. Our results suggest rotating the values in models that must retrieve from long contexts, preferably with a large rotary base.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.