acceptodds
Under review as a conference paper at ICLR 2027

SGD-Induced Parameter Drift in the Saddle-to-Saddle Dynamics of Deep Linear Networks

Abstract

Stochastic gradient descent (SGD) and gradient descent (GD) can remain close in function space while selecting distinct parametrization. We study this separation in deep linear teacher–student regression across initialization scales, in both online and offline training. Empirically, SGD and GD learn similar end-to-end maps throughout training, and SGD does not achieve lower population loss. However, for small initialization, where the network undergoes saddle-to-saddle dynamics, we find that SGD systematically redistributes scale across layers. In the absence of label noise, once a teacher mode is learned, its output-layer singular value grows while the corresponding singular values in earlier layers shrink, leaving the mode's end-to-end contribution approximately constant. This corresponds to motion along the rescaling symmetries of the network. Layer imbalances are not conserved by SGD; near rank-deficient saddles, their mean drift is a gradient flow that reduces the gradient noise. For Gaussian inputs with teacher-aligned covariances and decoupled learned modes, we show that this flow is driven by the residual error of the unlearned modes, and prove that SGD follows it while near a saddle. The flow predicts that the output-layer singular value grows as , independent of depth, while the earlier-layer singular values decay as , where is the depth and is a learning-rate-scaled training time. With label noise, the layer singular values flow towards a unique residual-dependent fixed point until the next mode is learned. This fixed point shifts as additional modes are learned and, at full fit, recovers the noise equilibrium of Ziyin et al. (2024). These results characterize an implicit bias of SGD toward particular parametrization that can be substantial even when the learned function remains close to that of GD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.