What Is a Repeated Token Worth? Beyond Token Counts in Language Model Pretraining
Abstract
Data reuse in language-model pretraining is usually summarized by token counts: processed tokens, unique tokens, passes and tokens per parameter. We ask whether these counts determine what a repeated token is worth, and find that they do not, using single-seed sweeps over model sizes up to 2B parameters and pre-registered multi-seed paired studies on three data sources. First, the value of a repeat depends on the comparison: relative to stopping after one pass, a second pass is worth most of a fresh one (an estimated , and at least within measurement resolution), yet at a fixed compute budget, replacing fresh tokens with repeats raises loss for every model in a five-size sweep and mean loss for every pre-registered allocation on all three sources. Second, equal counts do not give equal loss: with all counts matched, concentrating repeats on fewer samples raises loss slightly but consistently, and replaying each shard consecutively rather than in spaced passes raises loss at all three sizes tested, by up to bpb. Third, scaling depends on which count is held fixed: at fixed unique data, across eight model sizes, larger models reach their loss minimum after fewer passes (about 15 down to 4 over a 16-fold size range), whereas at matched tokens per parameter the larger of two models loses less to repetition. A count-based loss surface fitted to smaller models predicts larger ones where reuse helps, but below its fitted range of unique data it predicts a benefit from repetition where every measured case gets worse. Finally, in a 60-model study, re-tokenizing repeats with BPE-dropout lowers loss at high repetition but not at matched compute. Reuse claims should therefore report the comparison, the allocation and order of repeats, and the tokenization, not only aggregate counts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.