Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval
Abstract
Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention. However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. On a controlled pre-training ablation, we show that confining positional encoding to local layers is sufficient to produce this semantic shift, and find that it selectively degrades representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: while SWA NoPE improves over RoPE on multiple-target retrieval and QA, it degrades when distinguishing semantically similar keys. These behavioral differences track the mechanism shift from positional toward semantic mechanisms rather than a uniform improvement in long-context retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.