acceptodds
Under review as a conference paper at ICLR 2027

CoFlow: Stabilizing LLM Instruction Tuning via Gradient Flow Compatible Data Selection

Abstract

Supervised fine-tuning is fundamental for adapting large language models (LLMs), yet training on comprehensive instruction datasets incurs significant costs and diminishing returns. Current data selection methods typically rank examples independently by quality, diversity, difficulty, or influence, overlooking that selected examples are optimized collectively through minibatches. We find that individually valuable examples can still conflict when optimized together, as their gradients may point in incompatible directions. Consistent examples form compact gradient clusters with high consistency scores, whereas conflicting examples scatter across directions and often receive low or negative scores. This indicates that efficient instruction tuning requires selected examples to be both individually useful and jointly compatible in gradient space. We introduce \ourmethod, a training-aware data selection framework for constructing stable instruction tuning flows. \ourmethod operationalizes gradient flow compatibility through compact flow probing, cone-compatible subset construction, and smooth minibatch scheduling. Experiments on NuminaMath, AlpacaEval, HumanEval, and MultiPL-E show that \ourmethod improves average downstream performance by 10.6% over strong selection baselines, reduces gradient conflict by 34.8%, and reaches full-data-level performance using only 4.1% of the training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.