acceptodds
Under review as a conference paper at ICLR 2027

DRIFT: Data Selection for LLM Instruction Tuning via On-Policy Attribution

Abstract

Data selection for instruction tuning is usually studied for efficiency, reaching strong performance with a fraction of the training data. We ask whether data selection can instead raise the performance ceiling, producing a better model than training on all the available data. To test this, we take a large language model (LLM) that has already completed supervised fine-tuning (SFT) on its full corpus and train it further on examples selected from that same corpus. Any gain therefore lifts the model above the ceiling set by full-data training. Existing selection methods yield little or no gain in this setting. We propose DRIFT (Data Refinement via On-Policy Influence Functions for Supervised Fine-Tuning), a data attribution method built on influence functions. DRIFT samples the model's own responses to validation queries and weights them by correctness to form a validation objective. It ranks the examples of the original corpus by their estimated influence on this objective, and the model then continues SFT on the top-ranked subset. This design is motivated by the proximity gap in influence function theory: we hypothesize that these on-policy responses are better validation targets than external references in terms of locality. DRIFT also corrects influence scores for their dependence on gradient norm within each validation task. DRIFT outperforms all evaluated data selection baselines on Olmo3-7B-Instruct-SFT and OpenR1-Distill-7B, raising average accuracy across eight benchmarks from 39.17 to 40.14 and from 57.23 to 58.66, respectively, with gains on benchmarks not used for attribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.