acceptodds
Under review as a conference paper at ICLR 2027

It's better to be second: Second-Order Residuals prevent vanishing gradients

Abstract

The residual connection has long been viewed as crucial for the stable training of deep architectures, from multi-layered perceptrons and ResNets to the transformer that modern large language models (LLMs) are based upon. The common understanding is that it is necessary to preserve gradient flow and speed up training convergence. However, is that really true? In this work, we prove that residual connections do not prevent vanishing (or exploding) gradients. Furthermore, they can lead to inefficient parameter usage by blowing up the output norm over depth. Instead, we show that using second-order residuals, where x_t+1 is a function of two previous layers x_t and x_t-1, can provably prevent vanishing gradients regardless of model depth. Second-order residuals can be substituted into modern deep architecture for little overhead cost, leading to comparable or better downstream performance on several tasks (synthetic multihop retrieval and multitask parity, as well as large-scale language modeling).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.