acceptodds
Under review as a conference paper at ICLR 2027

Prequential Pretraining: Learning to Continually Compress the Future

Abstract

A major approach within computational theory of intelligence is prequential coding, which views continual compression of a stream of data as identifying parsimonious models for the underlying data. Prequential coding has been an important motivation for pretraining, leading to language models that can accurately complete single documents. Despite these successes, there has been little work on studying prequential coding on the more general setting of temporal streams of text data. Success in this setting would enable models to make predictions of future texts, across documents and potentially serve as an important primitive for continual learning. We study the empirical feasibility of this problem and curate a high-resolution, verifiably timestamped corpus with a single temporal stream of text up to 10B tokens spanning 2010-2022, including revisions to the same documents over time. We evaluate standard language modeling approaches as coding algorithms and find significant room for improvement in exploiting temporal and long-range structure. Natural interventions, such as long context, reach only 0.69 bits per byte (BPB), compared to retrieving exact-match prefixes from earlier in the stream which reaches 0.49 BPB, revealing substantial headroom for long-context models. Finally, we identify that in principle, it is possible to attain far better compression rates of 0.43 BPB through masking tokens that are in principle knowable from the past.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.