acceptodds
Under review as a conference paper at ICLR 2027

Beyond Scalar Rankings: Nonlinear Attribution Aggregation for Data Selection

Abstract

Selecting high-quality training data is increasingly important when annotation, storage, and computation budgets limit how much of an available data pool can be used. Data-attribution methods support this process by assigning a value to each training example, but they are commonly paired with a top-m rule that compresses an example's effects into a scalar score, which can lead to redundant or poorly balanced subsets. In this paper, we introduce NADA (Nonlinear Aggregation of Data Attribution Components), a method-agnostic selection framework that preserves component-level attribution information and selects complementary subsets through nonlinear aggregation to obtain more informative training data. We demonstrate the effectiveness of our method on standard logistic-regression and ridge-regression models, as well as prompt-based fine-tuning models, showing that NADA outperforms the corresponding scalar top-score selector in most matched settings and demonstrating the advantage of nonlinear component aggregation in mitigating the information loss caused by scalar attribution ranking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.