acceptodds
Under review as a conference paper at ICLR 2027

Refreshing Less, Selecting Better: Cache-Efficient Diverse Influence Selection for LLM

Abstract

Gradient-based data selection methods such as LESS score each candidate example by the alignment between its gradient and the gradient of a target validation set, and the scoring stage dominates their cost because per-example gradient features must be recomputed whenever the model checkpoint changes. Across three selection seeds, two model families, two candidate pools, and two target tasks, gradient features cached at a post-warmup checkpoint and paired with fresh validation gradients preserve the ranking 40 optimizer steps later with Spearman correlation between 0.952 and 0.991, while the top-10% subset they induce misses 10 to 22% of the examples that full recomputation selects. We therefore propose Cached Diverse Influence Selection (CDIS), which recomputes gradient features for the top-ranked fraction of candidates under the stale scores, fits an affine calibration on the recomputed examples, and selects the final subset under source and length quotas. A refresh fraction at or above the selection fraction recovers the exact top- subset whenever the calibrated stale scores have bounded error, and the measured budget curves follow this rule: at the recovered top- subset coincides with full recomputation in every setting, allocating the same budget per stratum recovers the stratified subset at 0.92 to 1.00, and the wall-clock of the gradient stage drops by a factor of 3.5 to 3.6. Iterating the cache over four checkpoints keeps top- overlap at 0.98 or higher at 1.9 gradient features per example against 4 for full recomputation, and the agreement between stale and recomputed scores on the refreshed examples provides a free check for unsafe reuse. Downstream, unconstrained top- selection on influence scores collapses to a single data source and scores 13 points below random selection on GSM8K. CDIS scores 12 points above random selection with paired confidence intervals that exclude zero and trails full recomputation by 4.3 points, one training-run standard deviation, at 3.4 times lower selection cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.