acceptodds
Under review as a conference paper at ICLR 2027

Estimating Dataset Learning Utility Before Training: A Manifold Volume Perspective

Abstract

Assessing the quality of a candidate dataset before model training is increasingly important for data acquisition and emerging data marketplaces. However, dataset size can be easily inflated by duplicate or highly redundant samples, while existing intrinsic quality metrics do not directly reflect downstream learning utility and model-aware valuation methods often require training on the candidate data itself. In this work, we study label-free and fine-tuning-free dataset-level quality assessment, aiming to estimate the future learning utility of an entire candidate dataset before training on it. We propose a simple geometric criterion based on the manifold volume of frozen pretrained representations. Our key hypothesis is that informative and non-redundant data span richer and more diverse feature directions, whereas redundant data concentrate in a more restricted subspace. Theoretically, we show that manifold volume is invariant to trivial sample replication, becomes dimensionally constrained under severe redundancy, and exhibits diminishing returns along already well-covered feature directions. We evaluate the proposed criterion through controlled redundancy experiments on nine datasets spanning three modalities, four architecture families, and 13 dataset-backbone pairs. Manifold volume consistently exhibits a strong positive linear association with downstream learning utility, achieving the highest average Pearson correlation of 0.816.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.