Measuring the Attention Span of Long Context Training Data
Abstract
LLMs are typically pretrained with short context lengths which are later extended through additional long-context training. As long context training is expensive, this stage requires careful data selection to maximize effectiveness. Data selection methods should distinguish a cohesive document with long dependencies across distant sections from a simple concatenation of parts with no real long distance structure. We propose to score documents based on their , which measures the look-back distance covered by the attention mechanism during a forward pass. Naive measurements of attention span are heavily influenced by attention sinks, making them a useless proxy for document structure. We present simple tricks for ignoring attention sinks, resulting in meaningful metrics of long-context structure that are useful for data filtering. When extending OLMo-3-7B from 8K to 32K using 5B tokens sampled from its own long-text mixture, additional AttentionSpan curation improves HELMET-Recall from to . The advantage remains even in a domain-matched experiment, in which context is extended using the same code, arXiv and book proportions for each checkpoint, and AttentionSpan is only used for filtering within domains. These results support sink-corrected attention distance as an interpretable curation signal for long-range retrieval training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.