Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
Abstract
Long-context adaptation is often treated as window scaling, but packed training with document masking creates a token-level supervision mismatch: most target tokens still receive short effective context. This paper introduces Effective-Context Exposure Balancing (EXACT), a supervision-allocation objective that upweights rare long effective-context targets according to their inverse frequency in the long tail. Across seven Qwen and LLaMA continued-pretraining configurations, EXACT improves all 28 trained-context and extrapolated-context comparisons on NoLiMa and RULER. On Qwen2.5-0.5B, EXACT improves NoLiMa by +10.09 in the trained setting and +5.34 under extrapolation, and improves RULER by +10.69 and +5.55. On LLaMA-3.2-3B, RULER improves by +17.91 and +16.11. Standard QA and reasoning performance is preserved, with a +0.24 macro change across six benchmarks. A distance-resolved probe shows that the gains arise when evidence is thousands of tokens away, while short-context cases remain stable. The results support a supervision-centric view: long-context adaptation depends on how strongly training supervises long-context predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.