T^3: Training-Time Testing for Data-Constrained Pretraining
Abstract
High-quality data for pretraining is limited and grows more slowly than compute, which makes it important to get the most out of a fixed corpus. Prior work improves generalization under fixed data with heavy weight decay, ensembling, distillation, and layer looping. We ask whether the training data can also serve as validation data for every update. We propose Training-time testing (): it splits each batch into two halves, takes a temporary descent step on the first half, evaluates the gradient of the second half at the adapted weights, discards the temporary step, and updates with the mean of the two gradients. Added to strongly tuned data-constrained recipes, lowers validation loss by 0.012 to 0.015 nats per token at 300M, 600M, and 1.4B parameters. On SlowRun, adding the step alone improves the top one-hour script by 0.010 to 0.012 within its time limit, and scripts built on it set the leaderboard records on both the 15-minute and one-hour tracks. On both benchmarks the gain persists after ensembling. Probes show that the gain requires a descent step tested on other tokens; the cosine between the gradients of different examples does not increase, and flatness alone is insufficient. In a nonlinear model of repeated training, we prove that at a common training-loss target, cross-example testing learns more shared signal and less sample-specific noise than ordinary training, yielding lower population risk.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.