Rethinking Data Scaling Through Data Geometry: A Uniformity Perspective
Abstract
Scaling training data does not necessarily scale the amount of useful learning signal available to a model. We study this discrepancy through the geometry of training data and ask why carefully structured subsets can learn more efficiently than substantially larger datasets. Across controlled regression and LLM fine-tuning, subsets with more uniform geometry optimize faster than equal-size random subsets, and in several settings a half-size subset even outperforms the full dataset, revealing that scaling data can add redundancy faster than useful geometric coverage. We provide two complementary theoretical explanations for this phenomenon. From an optimization perspective, closely spaced examples can induce nearly dependent rows in the model Jacobian, leading to poorly conditioned data geometry and slower convergence under gradient descent. From an approximation perspective, data spacing and local simplex shape determine the conditioning of the interpolation geometry, with more uniform data simplices yielding tighter approximation bounds. Both views connect learning efficiency to data uniformity through the minimum pairwise separation and related geometric quantities, with consistent empirical benefits across datasets, model sizes, and optimization settings. Together, these findings provide a uniformity-based explanation for why training on smaller, geometrically structured subsets can be more efficient than training on larger datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.