acceptodds
Under review as a conference paper at ICLR 2027

OptiSelect: How does the Optimizer Shape Data Curriculum?

Abstract

Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate's value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first systematic study of how the optimizer shapes data selection. Our theory establishes a selection gain principle in which the advantage of online selection is governed by the discriminability of the optimizer-induced utility scores. We prove that sign-based and polar-tangential preconditioner of \Lion and \Muon would suffer from a discriminability collapse which caps attainable gains from OptiSelect, whereas diagonal-adaptive optimizers such as \AdamW and \Sophia admit strictly better upper bounds. The proposed principle also quantitatively derives the optimal candidate oversampling ratio. Pretraining experiments on 124M and 720M models conform with our theoretical analysis and show that \AdamW's diagonal-adaptive scoring geometry remains the strongest scoring geometry even under \Muon optimizer. We further demonstrate that OptiSelect maintains the benefits under modern rephrasing, which applies in modern data processing pipeline. Our findings provide theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.