acceptodds
Under review as a conference paper at ICLR 2027

The Implicit Bias of Mirror Descent on Multiclass Separable Data: Stochasticity and Momentum

Abstract

The success of deep neural networks has spurred growing interest in understanding optimization dynamics and the solutions selected by different algorithms. In this paper, we study the implicit bias of normalized mirror descent (NMD) for multiclass linear classification with cross entropy loss, elucidating how stochasticity and momentum shape the selected direction. Our key insight is that mini-batch sampling can alter NMD's implicit bias, while momentum counteracts this effect. More specifically, under homogeneous mirror potentials, we show that full-batch NMD converges to the corresponding maximum-margin direction, irrespective of momentum. In contrast, under mini-batch sampling without momentum, NMD can lose sight of the full-batch maximum-margin direction. Crucially, momentum restores the full-batch implicit bias, enabling NMD to recover the same convergence direction. Experiments on linear models and neural networks support our theoretical findings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.