Selecting Natural Long Documents for Context Extension
Abstract
Natural long documents are widely used to extend language models to longer contexts, yet length alone does not reveal whether they provide useful long-range supervision. We present a controlled study of how such documents should be selected, and scored. Using Qwen2.5-3B, 128K-token continued training, and matched 6B-token budgets, we compare repacking shorter documents from the preceding training stage into long sequences, random sampling of natural long documents, and generic-quality, LLM-judge, and likelihood-based selection; HELMET tasks are used for model selection and RULER for held-out evaluation. Random sampling is only slightly better than repacking at shorter evaluation lengths, but leads by approximately 2.5 HELMET points and 3.3 RULER points at 128K. Within a fixed pool of natural long documents, the evaluated generic-quality and LLM-judge pipelines do not exceed random sampling, whereas the likelihood-based long-dependency pipelines achieve higher observed scores. Selecting documents by their density of LongPPL key tokens—tokens whose prediction improves markedly when the long context is available—achieves the highest observed scores among the evaluated pipelines. Its most consistent cross-benchmark gains are retrieval-oriented and widen toward 64K–128K, whereas QA effects are benchmark-dependent and aggregation and variable tracking show no consistent gain. Because scoring every document over the full 128K span is expensive, we introduce an efficient variant, , that restricts scoring to the first 64K target positions and yields downstream scores close to those of the full-128K pipeline while providing a measured end-to-end scoring speedup. Finally, in one evaluated 66/34 short/long mixture, replacing random long documents with selected ones improves macro-averaged HELMET by 4.6 points and RULER by 3.3 points without a material additional observed reduction on seven short-context tasks, and this advantage is further corroborated after supervised fine-tuning on LongBench v2 and MRCR. Overall, we hope the controlled and comprehensive comparisons will provide more practical references for long-context training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.