acceptodds
Under review as a conference paper at ICLR 2027

Certifying Residual Architectures from Their Primitives: A Sharp Stability Threshold

Abstract

Whether a deep residual architecture trains stably is usually determined by training it, which is expensive and answers the question only for the architecture, depth, and floating-point format that were tried. In practice stability is secured by heuristics on where in the residual block to place a normalization and which one to use; each is supported by experiments, and no common principle explains why and when they work. We show that stability can instead be certified before training, from the architectural primitives of the residual block, as an explicit function of depth and floating-point format. The certificate rests on two elements: (a) on a power-law growth bound for the residual block, , whose exponent is computed from the block's primitives by an arithmetic of exponents (forward); (b) on bounds on the gradients through a Lipschitz condition on the block over the reachable states (backward). The certificate determines the depth at which the state can reach the largest finite floating-point value : for , overflow requires at least order layers, whereas for it can occur within order layers; the threshold is sharp. Every normalization and every bounded activation sets , and the arithmetic identifies minimal modifications that bring a block from to without normalization. Empirically, in forecasting only blocks blow up, as the forward certificate predicts; in GPT-2 on OpenWebText, blocks without normalization also diverge, and there the forward certificate holds while the backward one fails, in the attention block with the largest query-key coefficient. Relaxing normalization from to improves out-of-distribution generalization in operator learning, and gives accuracy gains in time-series forecasting. As recent deep models, including foundation models, grow more complex and more costly to train, our results give a way to settle the trade-off between stability and representational flexibility of a residual architecture directly from the primitives of its blocks, before any training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.