Logarithmic Activation Clocks in Diagonal Linear Networks: When SGD Does and Does Not Follow Lasso
Abstract
Under a path-monotonicity condition, the from-start average of gradient flow in a diagonal linear network, rather than its raw output, recovers the Lasso regularization path. We ask what survives for stochastic gradient flow (SGF) and finite-step stochastic gradient descent (SGD). Integrating the square-factor log dynamics yields the Lasso KKT system plus coordinatewise Itô variance clocks and a self-normalized martingale. In an exactly solvable least-squares model, fixed SGF noise and a fixed SGD step produce different weighted-Lasso paths; ordinary Lasso therefore need not survive in either regime. In weak noise, only clock mass accumulated before future support motion enters the comparison. Under the stated nonexplosion, localization, and path-variation conditions, the averaged path returns to Lasso without requiring . For the standard signed product parametrization, an exact rotation produces two nonnegative channels driven by the same minibatch, with distinct finite-step clocks and . When the positive and negative Lasso channels have zero downward variation, the stopped, averaged actual-SGD path converges in prediction norm and signed-Lasso objective as and , again without an condition. For correlated row sampling, our positive results retain explicit tail-control assumptions; a universal off-face excursion bound remains open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.