Scaling Changes the Direction of Correctness Churn More Than Its Amount
Abstract
Scaling improves aggregate language model performance, but individual predictions do not improve monotonically. Larger models correct some errors made by smaller models and also turn some previously correct predictions into errors. We characterize these changes using two quantities. Correctness churn, C, measures the fraction of positions whose correctness changes. Directional imbalance, D, measures whether those changes favor the larger model. Across three released model families, their base pretrained checkpoints, three corpora, and six data controlled Pythia transitions from 160M to 12B parameters, 3.1–5.2% of correct next token predictions become incorrect after scaling, with no consistent decrease as model size grows. A similar pattern appears in log likelihood. Positive and negative shifts of 0.289 and 0.275 nats yield a net improvement of only 0.014 nats. PolyPythias provides nine independent runs at each of five model sizes, allowing scaling transitions to be compared with retraining at a fixed size. Reinitialization and data reshuffling reproduce most of the observed churn but produce little directional imbalance. At the 160M→410M transition, correctness churn is 11.3% under scaling and 9.5% under retraining, with a bootstrap 95% confidence interval of [+1.45, +2.13] percentage points for the difference. Directional imbalance is D = +0.273 under scaling and |D| = 0.024 under retraining. The directional gap remains positive at every model size after Holm–Bonferroni correction, while larger models increasingly make the same errors. These directional differences also support selective routing. We derive a total variation bound on the maximum gain obtainable from a candidate routing statistic. Peer model disagreement provides an effective signal for separating rescued and broken predictions. A consensus based cascade that escalates positions with peer disagreement improves over universal large model escalation across all evaluated transitions while using less compute, and approaches the empirical ceiling for the routing rules considered.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.