ReGIFT: Refreshed Geometry for Instruction Fine-Tuning via Dynamic Data Selection
Abstract
Instruction-data selection aims to improve fine-tuning performance under a limited training budget. Static selectors fix a subset before training, whereas dynamic selectors revisit it as the model evolves. Repeated selection therefore requires a ranking criterion that adapts to the model and remains affordable to recompute. We introduce ReGIFT (Refreshed Geometry for Instruction Fine-Tuning), a dynamic selector that constructs and updates a geometric reference from the current model's internal representations of the training pool at every refresh. It estimates layer-wise principal directions from a small uniform pool probe and adjusts loss–entropy rankings using candidate representations obtained during the existing scoring forward passes. This design requires no separately curated reference set, external judge, or learned selection policy. Using Davis–Kahan analysis, we bound changes in the geometric reference between refreshes under bounded covariance change and positive spectral separation. In single-run comparisons, the method selects 10% of the pool per refresh and achieves the highest eight-task average among all compared methods in the main setting, including full-data fine-tuning. Across six backbone–pool settings, it ranks first on the task average in five and second in one among the methods evaluated in each. In the main setting, total runtime is 33.0% lower than Data Agent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.