Estimating Dataset Learning Utility Before Training: A Manifold Volume Perspective
Abstract
Assessing the quality of a candidate dataset before model training is increasingly important for data acquisition and emerging data marketplaces. However, dataset size can be easily inflated by duplicate or highly redundant samples, while existing intrinsic quality metrics do not directly reflect downstream learning utility and model-aware valuation methods often require training on the candidate data itself. In this work, we study label-free and fine-tuning-free dataset-level quality assessment, aiming to estimate the future learning utility of an entire candidate dataset before training on it. We propose a simple geometric criterion based on the manifold volume of frozen pretrained representations. Our key hypothesis is that informative and non-redundant data span richer and more diverse feature directions, whereas redundant data concentrate in a more restricted subspace. Theoretically, we show that manifold volume is invariant to trivial sample replication, becomes dimensionally constrained under severe redundancy, and exhibits diminishing returns along already well-covered feature directions. We evaluate the proposed criterion through controlled redundancy experiments on nine datasets spanning three modalities, four architecture families, and 13 dataset-backbone pairs. Manifold volume consistently exhibits a strong positive linear association with downstream learning utility, achieving the highest average Pearson correlation of 0.816.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.