Data Value Depends on the Full Training State
Abstract
A post-training pipeline has to choose its next dataset before it can read any outcome, so it decides from a score computed ahead of the run. Every such score returns one verdict per dataset, read from the data alone or from the data together with the current parameters. We show that the usefulness of a dataset is a property of the data together with the state that reads it, and that this state carries more than the weights: it carries the optimizer memory and the step clock as well. Two states can therefore share a verdict on a dataset whose effects at those states have opposite signs, and the shared verdict is wrong at one of them. We call the effect at the receiving state causal trace value (CTV) and measure it directly, by restoring the saved state and continuing once on the candidate and once on the incumbent at equal tokens and steps, so that the paired difference in held-out accuracy credits the choice of data alone. The measurement confirms itself before it reports, replaying one continuation and returning no verdict when the replay disagrees. Any rule whose verdict two states share pays a summed regret of at least the smaller of the two effects on a dataset whose effects at those states have opposite signs, so the price of a shared verdict is a quantity. On a frozen benchmark of text-to-SQL decisions, measurement achieves positive mean validation utility over twelve training configurations at a matched admission count, admits a fraction of the candidates that every rival policy must match, and trains continuations whose price we report beside what they buy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.