Perplexity Can Reverse the Verdict on Model Collapse: A Frequency-Band Decomposition
Abstract
Recursive training on model-generated data is reported to erode the rare tail of a model's distribution, yet it is almost always evaluated with held-out perplexity. Perplexity weights each token by how often it occurs, so rare tokens barely count: on WikiText-2 the rarest band of training tokens carries 1.3% of the validation mass and the commonest 23.7%. We decompose held-out loss by each token's training frequency, an exact identity that needs no additional evaluation data. In five of ten recursive-training experiments, the aggregate and the rare bands give opposite verdicts on the same intervention, raising the sampling temperature. A conclusion drawn from perplexity can therefore be reversed by the numbers it already contains. A pre-registered forced choice that tests rare-token knowledge directly sides with the rare bands wherever the generator is not already well fit, and separates the arms by 1.95 accuracy points on checkpoints where three standard benchmarks find nothing. The mismatch cannot be outgrown. Under a Zipf corpus we prove that the validation mass on the tokens recursive training endangers decays as , the same exponent that governs how fast the tail is lost, so a larger corpus degrades the measurement as fast as it improves the process. The disagreement survives the band definition, a measured seed floor, a four-times-larger corpus, five further interventions and two controls for alternative explanations. We also prove that error stays bounded under unbiased accumulate-subsample training, resolving a case left open by Dey & Donoho (2024). We recommend reporting the band decomposition alongside perplexity. Code and results: https://anonymous.4open.science/r/band-resolved-collapse-1319.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.