acceptodds
Under review as a conference paper at ICLR 2027

RoPE Knows Its Limits: Frequency Bands Show Effective Context Length in LLMs

Abstract

Recent large language models claim to support extremely long contexts. However, as demonstrated by the RULER benchmark, the maximum sequence length a model claims to accept can diverge substantially from its effective context length, namely the length over which the model can actually process contextual information. This observation raises a natural question: is effective context length merely an empirical property that can only be measured by running long-context benchmarks, or can it be predicted from a model’s internal structure? In this work, we focus on the frequency bands induced by Rotary Position Embeddings (RoPE), which are widely used in modern LLMs. Through our analysis, we find that the trends in effective context length observed in RULER are proportional to the base frequencies of the dimensions in which these frequency bands appear. This suggests that the context length a RoPE-based LLM can truly exploit may be determined by the scale of the frequency bands it actually uses. In hybrid attention architectures that combine Full Attention and Sliding Window Attention, the frequency bands differ depending on the attention type, and these bands indicate the effective context length available to each layer. Our findings recast effective context length as a model-specific spectral quantity, enabling its dataset-independent estimation in RoPE-based LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.