acceptodds
Under review as a conference paper at ICLR 2027

Can Feature Learning Decouple From Loss Minimization?

Abstract

Does feature learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates keep a fixed step length regardless of the gradient norm. Near the edge of stability, full-batch Muon on teacher–student problems enters approximately period- loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher–student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. Across ReLU, GELU and SiLU teacher configurations, direction-only alignment metrics and projected head refitting show that the directions learned during plateaus are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.