Preconditioned Sharpness-Aware Minimization Can Exploit Large Perturbation Radii
Abstract
When sharpness-aware minimization (SAM) is combined with Adam-type optimizers, its perturbation and update operate in different geometries. Preconditioned SAM (PSAM) (Singh et al., 2025) addresses this mismatch. We analyze PSAM in its own right and examine its practical use. We show that each PSAM step is standard SAM in scaled geometry when the perturbation and update share a preconditioner. This yields a stationarity bound and a fixed-metric PAC-Bayes bound. Our descent analysis under quasiconvexity introduces an admissible radius and suggests how PSAM's perturbation direction can allow larger radii. In practice, early preconditioner instability motivates a warm start for PSAM. Across vision and language tasks, warm-started PSAM performs best among the evaluated methods. In language modeling, the best-performing radius for warm-started PSAM is up to 25 times that for warm-started SAM, while other methods deteriorate at these radii. We observe larger radius gaps when the SAM and PSAM perturbation directions are less aligned, consistent with our geometric analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.