acceptodds
Under review as a conference paper at ICLR 2027

WIDTH AND DEPTH LIMITS COMMUTE DURING TRAINING IN LINEAR RESIDUAL NETWORKS

Abstract

Recent work showed that width and depth limits in residual networks with branches scaled with commute at *initialization* in the sense that no matter how those limits are taken (e.g. width first, then depth, or depth first, then width, or with a fixed width/depth ratio, etc), the network output converges to the same limit at initialization. In this paper, we extend this result to commutativity *during training* for linear residual networks: we show that under the same branch scaling, for any fixed number of gradient steps with gradient clipping, the network output converges in as , implying commutativity of the width and depth limits during training. We further show that such result does not hold without gradient clipping, and a weaker result can be obtained in that case under the weaker convergence in probability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.