acceptodds
Under review as a conference paper at ICLR 2027

ComposeLong: Self-Supervised Composition of Long-Context Training Data from Short Texts

Abstract

Effective long-context training for large language models (LLMs) requires long training sequences with meaningful dependencies across distant parts of the context. Constructing such data presents two fundamental challenges: naturally occurring long documents are scarce, and document length alone does not guarantee useful long-range dependencies. Existing approaches address these challenges either by synthesizing long sequences from short texts or through quality-based filtering of existing long documents. However, both directions rely on hand-designed rules or criteria: synthesis methods do not learn from the structure of real long documents, while filtering methods retain only a fraction of the already limited supply of naturally occurring long documents. We introduce ComposeLong, a self-supervised framework that learns how to construct long-context training data using real long documents with strong long-range dependencies as supervision. Once trained, ComposeLong applies the learned patterns of document organization to large collections of short documents, synthesizing long-context training data without generating new text. ComposeLong outperforms prior synthesis methods on most reported metrics across RULER and LongBench v2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.