acceptodds
Under review as a conference paper at ICLR 2027

The Fates’ Spindle: Rotation Representation Emerges for Relative Position in Transformer

Abstract

Large language models rely on relative positional information across diverse tasks, while the mechanisms underlying relative position reasoning have yet to be elucidated. Using a controlled synthetic task that isolates relative position retrieval, we systematically study Transformers across positional encoding schemes and uncover a unified geometric mechanism. Across NoPE, ALiBi, T5 relative bias, KERPLE, and learned absolute embeddings, models converge to nearly identical geometry: keys and queries collapse into an approximately two dimensional subspace and trace a uniform rotation completing one cycle over the training context. This structure is equivalent to generalized RoPE, with all active channels sharing a frequency of approximately per position, where is the training context length. We further show that this single frequency representation is theoretical inferior to suitable multifrequency alternatives, explaining the advantage of richer frequency components in RoPE and related representations. Guided by this analysis, we evaluate a new frequency allocation scheme, providing support for the framework and its practical utility. Our findings provide a understanding of how Transformers learn relative position representations, establish synthetic tasks and geometric analysis as a principled approach to investigating positional capabilities, and inform future positional encoding design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.