acceptodds
Under review as a conference paper at ICLR 2027

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

Abstract

We study differentially private (DP) training with Muon, a matrix-valued optimizer that updates hidden-layer weights using momentum followed by Newton-Schulz orthogonalization. While DP-SGD is well understood, the interaction between per-example clipping, Gaussian noise, momentum, and nonlinear orthogonalization in Muon has not been systematically analyzed. We formulate DP-Muon, a private Muon procedure that clips per-example gradients globally, adds Gaussian noise to the clipped minibatch average, and then applies momentum and Newton-Schulz orthogonalization as post-processing. We prove that DP-Muon inherits the privacy guarantee certified by standard subsampled-Gaussian accounting, with no additional privacy cost from Muon-specific post-processing. On the optimization side, we establish finite-horizon and vanishing stationarity guarantees under global clipping, with bounds that separate optimization error, clipping residual, privacy noise, and Newton-Schulz approximation error. We further show that Gaussian privacy noise, although centered before the nonlinear Newton-Schulz map, induces a systematic bias in the expected update after this nonlinear transformation. This motivates DP-MuonBC, an adaptive bias-corrected variant designed to cancel the leading bias induced by the current Gaussian perturbation while preserving the same privacy guarantee. Experiments on E2E at four privacy budgets show that the tested Muon configurations improve test NLL over Adam baselines, with further gains for the selected DP-MuonBC configuration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.