acceptodds
Under review as a conference paper at ICLR 2027

How Bregman Divergences Shape Shampoo

Abstract

Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback–Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through spectral analysis and empirical investigation on the gradient second moment, we examine how divergence choice influences approximation behavior and preconditioner estimation. We find that some divergences can compensate for the bias introduced by approximating the gradient second moment with a finite number of gradients, and that this helps explain the behavior of their Shampoo variant in GPT-2 pretraining and synthetic experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.