acceptodds
Under review as a conference paper at ICLR 2027

Logarithmic Activation Clocks in Diagonal Linear Networks: When SGD Does and Does Not Follow Lasso

Abstract

Under a path-monotonicity condition, the from-start average of gradient flow in a diagonal linear network, rather than its raw output, recovers the Lasso regularization path. We ask what survives for stochastic gradient flow (SGF) and finite-step stochastic gradient descent (SGD). Integrating the square-factor log dynamics yields the Lasso KKT system plus coordinatewise Itô variance clocks and a self-normalized martingale. In an exactly solvable least-squares model, fixed SGF noise and a fixed SGD step produce different weighted-Lasso paths; ordinary Lasso therefore need not survive in either regime. In weak noise, only clock mass accumulated before future support motion enters the comparison. Under the stated nonexplosion, localization, and path-variation conditions, the averaged path returns to Lasso without requiring . For the standard signed product parametrization, an exact rotation produces two nonnegative channels driven by the same minibatch, with distinct finite-step clocks and . When the positive and negative Lasso channels have zero downward variation, the stopped, averaged actual-SGD path converges in prediction norm and signed-Lasso objective as and , again without an condition. For correlated row sampling, our positive results retain explicit tail-control assumptions; a universal off-face excursion bound remains open.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.