acceptodds
Under review as a conference paper at ICLR 2027

LDIS: Efficient Long-Context Training Data Synthesis via Interleaved-Study

Abstract

Long-context modeling enables language models to process complete documents, code repositories, and extended interaction histories. However, natural text provides sparse supervision for long-range dependencies, while recent hybrid models further face additional constraints from the finite memory capacity of the recurrent blocks. We propose Long-context training Data synthesis via Interleaved Study (LDIS). A teacher generates question–answer (QA) pairs from growing document prefixes, which we then interleave with source text at the chunk boundaries. This changes the training prefix distribution across the sequence, adding targets that encourage finite-state blocks to retain and use earlier information. We normalize document and answer losses separately to prevent document tokens from diluting answer supervision. Under matched student training token budgets, LDIS improves long-context performance across model sizes and teacher sources when adapting full attention to sliding-window attention and extending context length, while largely maintaining general capability. Ablations support interleaving and the loss design. We also explore self-improvement without an external teacher, using an internally trained model to supervise a new student initialized from the same base checkpoint.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.