Toward a More Favorable Curvature Structure: A Basis Perspective on Matrix Optimizers
Abstract
Matrix optimizers have shown promising performance in large language model training. Although prior explanations have linked their effectiveness to curvature, the underlying mechanisms remain incompletely understood. To better understand these mechanisms, we view optimizer updates as basis selection followed by a diagonally preconditioned update in the induced basis. We find that bases induced by matrix optimizers can make curvature structure more favorable for diagonal preconditioning, partly explaining their effectiveness. Empirically, we examine Gauss–Newton curvature during training and find that (i) matrix-optimizer-induced bases generally yield greater diagonal concentration of curvature in hidden parameter blocks than the ambient basis used by optimizers such as Adam; (ii) in the basis induced by each optimizer, the scaling used by its diagonal preconditioner closely corresponds to the diagonal curvature; and (iii) optimizers with higher diagonal concentration of curvature in their induced bases generally achieve better training performance. We further support these findings with an initial theoretical analysis of quadratic problems and neural networks, identifying conditions under which bases induced by matrix optimizers yield more favorable curvature structure and examining how this structure can contribute to improved optimization performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.