acceptodds
Under review as a conference paper at ICLR 2027

Gradient Matching Meets Muon: Efficient Online Minibatch Selection for Language Models

Abstract

Training large language models with large batch sizes is often prohibitive due to memory constraints. Recent works in both language and vision models have shown that training on significantly smaller, carefully selected subsets can achieve performance closer to that of full-batch training. A substantial line of work uses gradient matching for such sample selection, where approximated last-layer gradients provide a useful proxy for aligning training samples with validation gradients, with importance scores tailored to standard SGD or Adam updates. However, recent spectral optimizers such as Muon, which have demonstrated gains over Adam, require a fundamentally different criterion: Muon orthogonalises the gradient momentum matrix before the parameter update, inducing an anistropic update. We show that the Euclidean inner-product score used by SGD/Adam-based gradient matching is systematically misaligned with Muon’s update direction, while an orthogonalised, momentum-aware score provides a tighter approximation to sample utility. We further study how Muon’s orthogonalisation affects coreset quality, particularly sample diversity. Building on these observations, we introduce Muon Spectral Matching for Adaptive Selection, an efficient greedy method that linearises the polar-factor orthogonalisation operator and applies a spectral correction to sample importance scores. Experiments on both large scale finetuning and pretraining regimes show the efficacy of gradient matching via Muon optimizer w.r.t Adam counterparts as evident across multiple subset selection baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.