When Does Gradient Descent Globally Train a Deep Linear ResNet?
Abstract
When does gradient descent globally train a deep linear residual network from a nonzero initialization? We show that finite-step trainability is governed by three distinct aspects of factorization geometry that are conflated by end-to-end conditioning: a lower radial coordinate measuring cumulative spectral contraction, an upper radial coordinate measuring cumulative spectral expansion, and a centered logarithmic shape variable measuring multiplicative inter-layer imbalance. Our first result establishes a geometry-sensitive Polyak–Lojasiewicz inequality, revealing how and control gradient transmission through depth. Our main result proves that an explicit nested spectral-balance basin, defined by simultaneous bounds on , is forward invariant under depth-normalized gradient descent. The proof exploits exact first-order cancellation in the discrete evolution of adjacent-layer Gram imbalance, converts the resulting second-order drift into logarithmic shape control, and couples it with relative-step smoothness and summable radial motion. Consequently, every trajectory initialized in the basin remains there for all time and converges linearly to a global optimum. We further show that these three geometric effects cannot in general be collapsed into a single conditioning parameter. Exactly balanced scalar families attain the predicted radial scaling laws, while a radially neutral depth-two construction, with , exhibits a pure-shape stability threshold of order where is the logarithmic inter-layer imbalance defined above. Numerical experiments recover the predicted scaling laws and show that internal factorization geometry strongly orders finite-step stability even among factorizations representing essentially the same predictor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.