Learning What Is Worth Training: The Data-Value Bottleneck in LLM Self-Improvement
Abstract
As high-quality human data runs short, reasoning models are increasingly improved by training on their own verified solutions. These self-improvement loops usually stall after a few rounds, and the usual remedy is to generate more, or harder, verified data. We argue that the harder problem is deciding what to train on. The same verified data can help one model, do nothing for another, and damage abilities that should be kept, depending on the checkpoint it is applied to and on how it is used. We call this the data-value bottleneck and ask whether the value of a candidate update can be known before paying for a full training run. In a pre-registered study of mathematical reasoning models from 1.5B to 27B parameters, we find that training value depends on both the current checkpoint and the training method. At a 9B checkpoint, a set of 322 problems had little effect when their solutions were used for supervised fine-tuning, but moved reinforcement learning one round ahead when the problems were used as prompts. The same teacher-generated solutions improved 4B and 9B checkpoints but reduced performance on protected tasks at 1.5B. None of the diagnostics we computed before training reliably distinguished useful updates from harmful or ineffective ones. Short pilot runs were more informative about harm. Every candidate whose estimated value was negative after full training already had a negative estimated value after 25 optimizer steps, on the evaluation sets used to make the decision. The converse did not hold: some useful candidates also looked negative early, and pilots did not reliably rank candidates with similar final performance. We therefore propose a conservative veto-and-commit policy: reject candidates with evidence of harm, select a candidate only when its estimated value is credibly positive, and otherwise stop. In four prospective tests, the policy never selected an update with negative value on a sealed evaluation, one that is read only once, after the decision is committed. This caution had a measurable cost: at a 2B checkpoint, the policy stopped even though the best candidate later produced a 5-point gain. It also prevented a much larger failure: at a near-ceiling 27B checkpoint, it rejected a teacher update that reduced sealed target accuracy by as much as 39 points after full training. We release the policy, evaluation protocol, code, pre-registration records and per-example results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.