RoPE-Filter: Identifying Long-Range Dependencies for Long-Context Data Selection
Abstract
Long-context modeling is a core capability of large language models. Continued pretraining on long-context data is a standard approach for extending model context windows. Effective long-context extension requires data with strong long-range dependencies, but simply increasing the length of training sequences does not reliably improve long-context capabilities. In this work, we propose RoPE-Filter, a training-free method that probes each document's reliance on distant context by selectively attenuating the contributions of low-frequency RoPE blocks to attention logits in long-distance interactions. The resulting relative increase in perplexity is used to select documents with strong long-range dependencies. Experiments across widely used long-context benchmarks show that models trained on data selected by RoPE-Filter achieve better performance than those trained on data selected by existing open-source methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.