Spectral Data Valuation: What Data Teaches in the Model’s Information Geometry
Abstract
Data curation has become the central lever of large language model (LLM) development, yet the worth of a training example is still priced by surface proxies such as token counts, text space diversity, or post hoc benchmark deltas. Such proxies describe the data itself, not what the data teaches the model. Upon close inspection of how models internalize their corpora, we are struck by a hidden regularity: wherever learning is measured from inside the model, whether through the influence of examples, the rotation of weights, or the coverage of features, value concentrates on a narrow spectral head while sheer volume spreads over a nearly worthless tail. We trace this regularity to the model's information geometry, in which these seemingly unrelated measurements read one shared learning spectrum, and we make it operational with \ourmethod, a Spectral Data Valuation framework that prices a training example by the spectral energy its gradient delivers to dominant yet unsaturated directions, thereby turning data curation into maximizing information gain per token. Extensive experiments across model families and scales demonstrate that \ourmethod is (1) unifying, as the subspaces recovered from influence, weight rotation, and feature coverage align with a mean principal overlap of against a random baseline of , (2) efficient, matching full-corpus continual pre-training on GSM8K with only of the tokens while accelerating induction-head emergence by , and (3) predictive, forecasting downstream gains of unseen data mixtures with a rank correlation of .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.