acceptodds
Under review as a conference paper at ICLR 2027

What Does a Training Token Buy? Contextual Information Gain Beyond Perplexity Across Open Model Suites

Abstract

A lower perplexity on held-out text can come from two sources. The model may use the context better, or it may make a better guess about the text before any context is given; perplexity adds the two into one number, so a falling perplexity cannot say which improved. We separate them. For every held-out token we score the model twice, with its context and under its own unconditional next-token distribution, obtained by averaging its predictions over reference contexts. The difference, contextual information gain (CIG), is the pointwise mutual information between context and token under the model, so a higher value means the context did more of the work; the first term is the usual held-out cross-entropy, and the second tracks how the model’s unconditional expectations move during training. We measure both on 790 public checkpoints from seven model families (14M to 12B parameters, up to 1071× the Chinchilla-optimal data budget) on 206 held-out texts from ten domains, in bits per byte, and judge every trend against the model’s own checkpoint-to-checkpoint noise. In the three training runs whose learning rate and training-data composition stay constant, held-out cross-entropy keeps falling to the end of the window, at up to 173× the optimal budget, while the model’s prior moves away from held-out text by 18 to 29% of the loss improvement without becoming more peaked. The additive scaling law predicts that after the same number of training tokens every model size improves at the same rate per factor of two in tokens; instead, larger models improve faster, at all 58 matched comparisons on four model ladders across three corpora. Only the 70M to 160M Pythia models get worse late in training, whereas models trained far beyond their optimal budget keep improving. In warmup-stable-decay schedules the decay phase lowers held-out loss 17× faster per factor of two than the stable phase in three runs and 36× faster in a fourth; every documented decay phase coincides with a change of training data, but within the one decay phase with 13 checkpoints, at unchanged data, the rate rises 3.5× as the learning rate falls to zero. The measurement thus turns the model’s own noise into a stopping diagnostic, and tells, from checkpoints and held-out text alone, whether the next factor of two in training tokens still improves the use of context or only shifts the model’s expectations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.