Beyond Scalar Rankings: Nonlinear Attribution Aggregation for Data Selection
Abstract
Selecting high-quality training data is increasingly important when annotation, storage, and computation budgets limit how much of an available data pool can be used. Data-attribution methods support this process by assigning a value to each training example, but they are commonly paired with a top-m rule that compresses an example's effects into a scalar score, which can lead to redundant or poorly balanced subsets. In this paper, we introduce NADA (Nonlinear Aggregation of Data Attribution Components), a method-agnostic selection framework that preserves component-level attribution information and selects complementary subsets through nonlinear aggregation to obtain more informative training data. We demonstrate the effectiveness of our method on standard logistic-regression and ridge-regression models, as well as prompt-based fine-tuning models, showing that NADA outperforms the corresponding scalar top-score selector in most matched settings and demonstrating the advantage of nonlinear component aggregation in mitigating the information loss caused by scalar attribution ranking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.