Calibrating Long-Range Attention with Semantic Needles for Long-Context Pretraining Data Selection
Abstract
Effective long-context training requires documents with meaningful long-range dependencies, not merely long documents. Existing model-based selectors broadly follow two approaches: prediction-based contrast and attention-based measurement. Prediction-based selectors can measure the contribution of distant context but are costly to apply across a corpus, while attention-based selectors are more efficient yet often aggregate over heads without identifying those responsive to useful distant evidence. We propose SNAS (Semantic-Needle-guided Attention-based data Selection), which uses prediction contrast on a small calibration corpus to identify naturally occurring context-dependent target tokens and localize their distant evidence intervals. SNAS then calibrates model-specific attention heads by their enrichment over these intervals, without synthetic probes or downstream-task labels. For corpus-wide selection, it scores each document in a single forward pass using both distant-attention allocation and retrieved-content signals from the calibrated heads. Across two scoring models and three data domains, SNAS matches or numerically exceeds the strongest baseline in overall score across all six settings. With the Llama-3.1-8B scorer, SNAS achieves LongFilter's corpus-scoring throughput and uses 44% fewer GPU-hours; these measurements cover corpus scoring and exclude the one-time calibration cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.