acceptodds
Under review as a conference paper at ICLR 2027

Rethinking representativeness and diversity in dynamic data selection

Abstract

Dynamic data selection accelerates training by sampling a changing subset of the dataset while preserving accuracy. We re-examine two core notions underlying sample selection: representativeness and diversity. We define representativeness as weighted coverage of dataset-level high-frequency feature factors, moving beyond local geometric centrality, and treat diversity as a process-level constraint that requires the selection trajectory to gradually include complementary rare factors over training, rather than dispersion within a single subset. Based on this view, we propose a dynamic selection framework with three components. First, we score representativeness in a plug-in feature space to prioritize samples covering high-frequency feature factors. We instantiate this with a sparse autoencoder trained on the target dataset, using sparse unit activations to summarize both individual samples and dataset-wide factor statistics. Second, we realize process-level diversity by combining rare-factor sampling with a usage-frequency penalty that promotes sample rotation, provably discourages monopoly, and reduces gradient bias. Third, we couple the two-dimensional scoring with a smooth scheduler that transitions selection from core-pattern consolidation to rare-factor exploration, without extra gradients, influence estimates, or second-order computations on the training model. Extensive experiments across vision, text, medical, language modeling, and generative multimodal tasks and different models show that our method yields a better accuracy-efficiency frontier than prior selection methods: it retains most of the full-data accuracy at – lower training cost, and under a matched-propagation protocol it even surpasses full-data training on ImageNet-1K. Code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.