acceptodds
Under review as a conference paper at ICLR 2027

MoPE: Learning One Positional Structure for Short and Long Contexts

Abstract

Positional encoding is central to Transformer language models, yet explicit positional bias is not equally beneficial across attention components. Recent hybrid designs combine rotary position embeddings (RoPE) with no positional encoding (NoPE), but typically prescribe their allocation through fixed rotary fractions or layer schedules. We introduce Mixture of Position Embeddings (MoPE), a simple framework that learns this allocation from data. MoPE jointly optimizes the model backbone and binary gates that select RoPE or NoPE for each layer and rotary frequency. At inference, these gates define a fixed positional structure without input-dependent routing. The resulting gate configurations exhibit distinct layer- and frequency-dependent patterns, revealing heterogeneous positional requirements within Transformers. Experiments on models ranging from 70M to 1B parameters show that MoPE maintains competitive short-context performance while achieving strong long-context retrieval. After 4K pretraining and context extension to 64K, the 410M and 1B models achieve 94.9% and 98.8% NIAH accuracy at 128K, respectively. Combined with YaRN, the 410M model further achieves 99.1% NIAH accuracy at 256K.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.