acceptodds
Under review as a conference paper at ICLR 2027

DRESS: Descent-Oriented Data Selection for Token-Efficient Language Model Training

Abstract

Improving token efficiency is increasingly important for training modern large language models. Training tokens differ in their contribution to learning, and their usefulness can change as training progresses, motivating online data selection that adapts to the current model. We propose Descent-oriented Reweighting for Efficient Sample Selection (DRESS), which scores training examples by their predicted contribution to reducing a held-out target loss. From first- and second-order Taylor expansions, we derive two variants, DRESS-FO and DRESS-SO, where DRESS-SO uses a generalized Gauss–Newton (GGN) approximation to capture interactions among candidate updates. The two variants offer a trade-off between scoring cost and the fidelity of the local target-loss approximation. We evaluate DRESS on GPT-2 pretraining across model sizes from 124M to 1.5B parameters and on Qwen3.5-9B-Base LoRA fine-tuning. On GPT-2 XL, both variants outperform the compared baselines in average zero-shot accuracy, and with only 30B update tokens, they even outperform random selection trained on 60B update tokens. When tested on Qwen3.5-9B-Base, both versions of DRESS improve over the standard fine-tuning baseline on evaluated benchmarks, while the second-order DRESS outperforms the first-order version.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.