Data2Delta: Gradient Geometry for Budgeted Pretraining Data Selection
Abstract
When deduplication leaves more data than a pretraining budget permits, a selector must choose among examples that are not literal copies. We introduce Data2Delta, which measures redundancy in a fixed probe's gradient space: it clusters normalized gradient sketches and globally removes examples with low within-cluster ridge leverage. The cluster-average score is its effective dimension divided by its size, linking deletion allocation to directional redundancy without explicit cluster quotas. Two controlled Pythia-1B comparisons support the representation and membership choices: under the same selector, gradients improve endpoint perplexity (PPL) over dimension-matched BGE-M3 features; with each cluster's deletion count fixed, leverage-ranked membership improves PPL over random membership. End to end, Data2Delta lowers PPL by 0.0154 relative to uniform selection. Curve interpolation translates this gap into about 19% fewer effective prediction tokens to reach uniform selection's endpoint PPL. Low-NLL filtering achieves similar PPL, yet loses over 13 times as many source families in a separate natural-news evaluation at 10% removal. Training-pool provenance analysis also distinguishes the selections, while mean zero-shot accuracy is lower for Data2Delta than for the controls on the perturbed pool. These results support gradient geometry as a candidate-relative selection signal and distinguish its training-fit gains from source coverage and downstream accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.