acceptodds
Under review as a conference paper at ICLR 2027

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

Abstract

We study how each component of the Transformer feedforward block preserves rank across depth at initialization. We reinterpret skip connections and normalization as mechanisms for preserving gradient rank as well as magnitude across depth, since the operations that make a network expressive also reduce its rank. Skip connections route the gradient around the residual branch, where rank is lost, rather than along the long gradient paths that encourage the layers to compose. Across a combinatorial sweep of residual architectures that vary the normalization placement, initialization and branch scales, depth, branch shape and activation, we show that the effective rank of the input–output Jacobian is one function of the mean number of branches a gradient path passes through, which is determined by the branch-to-skip ratio. Deep residual networks therefore trade off rank collapse against ensemble-like behavior; preserving rank makes them ensemble-like by necessity. Normalization placement controls the branch-to-skip ratio across depth, which explains why rank collapses for Post-Norm but decays only slowly for Pre-Norm and unifies much of the normalization placement and depth scaling literature. The two-matrix structure uses additional parameters to preserve rank: the second matrix decorrelates a coherent mean spike that would grow across blocks with a single matrix and an uncentered activation, and expanding the width between the two matrices lets enough directions survive the activation to span the original space. In transformer language models, the rank at initialization follows the same function of path length, and most training failures coincide with rank collapse. Taken together, our results recast architecture design for deep networks as navigating an intrinsic tradeoff among rank collapse, ensemble-like behavior, and parameter count.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.