Science of Long Context
Abstract
Language models increasingly operate over long inputs, such as long documents, code repositories, and their own reasoning traces, and must retrieve information from any position in them. However, language model performance degrades as input length increases, and the extent of the degradation depends on the architecture. Every architectural choice is an inductive bias with its own limitations, and these limitations are poorly understood. We study them by training many architectures from scratch in a controlled setting, where the models differ only in architecture and serve as controls for each other. We find that removing explicit positional encoding alone does not improve length generalization, and that RoPE's retrieval beyond the training length degrades due to keys at offsets never observed during training. Hybrid models that restrict most layers to a sliding window still read far keys at these offsets in their global attention layers. We then show that in a relative positional encoding, keeping positions within the trained range comes at the cost of resolving nearby positions. Building on this insight, we design Log-Ratio RoPE, which rotates queries and keys by their log-position to keep offsets close to the trained range. We use it in the global attention layers of a hybrid model, while sliding-window layers retain RoPE to resolve nearby positions within the trained range. In 0.5B models trained at 8K, this design improves accuracy on the RULER long-context benchmark at 64K, eight times the training length, by 18.7 percentage points over RoPE and 8.1 percentage points over the RoPE–NoPE hybrid.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.