The Loss Does Not See the Basis, but Adam Does
Abstract
Different optimization algorithms can find very different solutions to the same problem. Gradient descent (GD) exhibits a low-rank bias in overparameterized problems when initialized near zero, while Adam does not. GD is rotationally equivariant: rotating both factors of a factored model by an orthogonal matrix leaves the loss unchanged, and GD's solution rotates accordingly. In contrast, Adam is not. We prove that any algorithm using only the current gradient and respecting rotational symmetry amounts to a left-preconditioner depending only on the gradient's Gram matrix. This gives a norm-based test, which passes for the spectral norm in Muon and fails for Adam's norm. Empirically, nine optimizers fall into groups consistent with the test on underdetermined matrix sensing. Sharing Adam's second moment across coordinates gradually recovers a low-rank bias, and its solutions approach GD's as the stepsize decreases. On four hyperspectral scenes, Adam's solutions are hundreds to thousands of times further from GD's than two GD runs with different step sizes are from each other. In transformers and LoRA adapters, Adam separates two rotationally equivalent models from the first iteration. Overall, recovering particular solutions depends on the optimizer's invisible coordinate basis, as well as how fast it grows small singular values relative to large ones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.