Beyond the Last Layer: Rethinking Gradient Proxies for Coreset Selection
Abstract
Gradient-based coreset selection offers a theoretically rigorous framework for data-efficient learning; however, its computational cost necessitates practical approximations. To scale these methods, existing heuristics, such as in CRAIG, approximate full gradient similarity by restricting gradient computation to the final layer, implicitly assuming that the penultimate representation suffices for data selection; a phenomenon we refer to as the Last-Layer Bias. In this work, we critically examine this assumption and investigate how the informative gradient signal is distributed across network depth. We propose a family of alternative gradient proxies that exploit information beyond the final layer, including input-level gradients (CRAIG-First), data-driven intermediate layer selection via FlexTuning (CRAIG-FT), and a distributed proxy based on Global Node Importance (CRAIG-NI), which selects sparse, non-contiguous subsets of high-activity neurons. We conduct one of the largest empirical studies in the coreset literature, evaluating 663 tabular classification and regression datasets, and extend our analysis to image classification and instruction tuning in large language models. Our results show that last-layer proxies often exhibit diminishing returns and barely outperform random selection (CRAIG-RL). In contrast, input-level proxies consistently accelerate convergence on tabular datasets, while distributed proxies such as CRAIG-NI generalize robustly across vision and language tasks, with particularly strong gains on visual benchmarks, effectively mitigating the Last-Layer Bias.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.