acceptodds
Under review as a conference paper at ICLR 2027

Implicit Bias of Factorwise Spectral Optimization in Self-Attention

Abstract

In self-attention, logits depend on the product , yet per-head Muon, used in recent LLMs, updates and separately. The training dynamics therefore depend on the factorization, not only on the product and its gradient. We ask which normalized product these updates select, and study this question through the momentum-free factorwise polar flow, a continuous-time idealization of per-head Muon. For a fixed gradient, sufficiently aligned factors keep full-rank gradients and, above the largest omitted singular value, converge to the leading singular directions. When factor alignment accumulates without bound, the minimum margin of the normalized product converges to the optimal width- margin characterized by the Ky Fan norm. For exact softmax attention, when each query has a prescribed target key, we prove local convergence of the product near the target, while the factors approach an orbit of common rotations. If the gradient weights on these target–competitor comparisons are also locally identifiable, they converge as well, at a polynomial rate in optimizer time. Numerical experiments illustrate that the product can converge even when its limit does not determine the gradient weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.