acceptodds
Under review as a conference paper at ICLR 2027

What Does a Data-Selection Gain Measure?

Abstract

A gain from selecting training data depends on the representation used to score examples and train the learner. It also depends on the candidate supply, the selected group mix and the labels used for guidance. We study what these comparisons measure using frozen feature maps and linear heads. Our central control redraws random examples at the same class–group counts as each selection. Across tasks from eleven benchmark sources, five registered comparisons show no clear worst-group benefit; secondary overall-accuracy residuals are negative in all 55 correlated source–implementation means. Two music comparisons are positive. Selecting real examples from unseen work families rather than generated variants of training works raises test accuracy by 2.88 percentage points. Random acquisition from the two supplies already gives a descriptive 2.60-point difference. On identical selected rows and random controls, retaining original descriptors alongside constructed features increases the selection advantage by 0.55 points relative to constructed features alone. This changes the learner, not the ranking rule. A larger reference gives a positive source-level estimate but an interval over repeated draws that spans zero; guidance shows no clear advantage over direct label use. Feature refresh fails its positive primary, and scorer–learner matching outside music shows no clear benefit. These findings do not provide a winning selector. They show why a selection gain should be reported with its candidate supply, scorer and learner, group-count control, and final accuracy. Code and numerical evidence are open-sourced at: https://anonymous.4open.science/r/data-selection-gains-ready-B863/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.