LongFocus: Focusing on Key Tokens for Efficient Long-Context Mid-Training
Abstract
Despite extensive progress in data selection for long-context mid-training, the objective typically remains a uniform token-averaged cross-entropy. Since only a small fraction of tokens substantially benefit from distant context, uniform averaging can dilute long-range supervision even within dependency-rich documents. Existing token-reweighting approaches address this mismatch, but their fine-grained token dependency scoring method introduces substantial training overhead. We therefore investigate whether a coarse context contrast can suffice when coupled with appropriate preprocessing. We propose **LongFocus**, an online objective that combines an inexpensive non-overlapping segmented reference with eligibility filtering and bounded rank-percentile upweighting for long-context mid-training. Specifically,LongFocus computes an inexpensive online dependency score, *SegRatio*, by measuring each token's relative loss increase from the full-context view to a coarse segmented view with no cross-segment attention. *PreSelect* restricts additional weighting by document length and full-view loss, while *Reweight* assigns bounded weights according to rank percentiles within the eligible set. We evaluate LongFocus on Llama-3-8B-Instruct, MoE-16B-A3B, and MoE-125B-A10B using a two-stage 64K-to-512K mid-training recipe, with evaluations before and after SFT. LongFocus improves Long-Avg over Vanilla CE in all matched model-stage comparisons. On MoE-125B-A10B, the post-SFT gains are 7.28 points (13.80% relative improvement) after 64K mid-training and 5.48 points (10.24% relative improvement) after the subsequent 512K stage. On Llama-3-8B-Instruct, it outperforms LongCE in Long-Avg at both post-SFT endpoints, with gains of up to 2.0 points (3.29% relative improvement), while reducing per-step training time and peak memory usage by up to 20.2% and 31.0%, respectively. Code is available at https://anonymous.4open.science/r/LongFocus-vsfnhyta.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.