acceptodds
Under review as a conference paper at ICLR 2027

Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning

Abstract

Gradient-based data selection scores a sample by how it moves the model. Online LLM fine-tuning complicates this: utility depends on the current parameters, the effective update geometry is set by an adaptive optimizer, and per-example gradients are too large to materialize, so online rules score samples independently and weight the selected batch uniformly. We formulate online selection instead as subset-level gradient matching: choose non-negative weights whose composite update aligns with the target under the optimizer-induced geometry. Solved with the coupled greedy solver that gradient matching conventionally uses, this degrades on LLM gradients—the solver refits coefficients while the support is still being built, and the estimation error is comparable to the correlation gaps that decide each choice. Two conditions recover it: coefficient fitting must leave support construction, and the optimizer geometry must enter the alignment term alone, with redundancy measured in the candidates' own metric. Filter-then-Weight meets both. To run it inside the training step we hold per-example gradients as factorized outer products, project the two factors rather than the gradient, and contract the token axis before the feature axes; selection then costs under 9% of a step and carries no quadratic dependence on context length. Across backbones from 0.6B to 8B and two optimizers, it ranks above every online selection baseline we evaluate on both main benchmarks. An ablation shows that the gain comes from filtering and reweighting together.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.