acceptodds
Under review as a conference paper at ICLR 2027

d-STABLE: Learning Residual Allocations Under a Controlled Depth Budget

Abstract

The sensitivity of a residual network—how much it amplifies a small input change across all layers—is a product over layers, so it is governed by one quantity: the total *residual budget*, the summed strength of the per-layer residual branches. We show that this budget and its allocation across layers are two separable levers. The budget's growth rate in depth sets the *stability class*: a linear budget makes worst-case sensitivity exponential, a square-root budget (Fixup, Stable ResNet) sub-exponential, and a *logarithmic* budget polynomial, so is the critical rate. The allocation leaves the class unchanged and sets accuracy within it. Prior stabilizers couple the two: fixed schedules freeze both, and learned gates (ReZero, LayerScale) let the budget drift. **d-STABLE** is a reparameterization that holds the total budget at exactly at every optimization step while the optimizer learns the allocation. When every branch is norm-bounded this gives an a-priori polynomial bound on sensitivity; otherwise the pinned budget is an inductive bias whose amplification we measure. Our experiments separate the two levers. On `ogbn-arxiv` at 64 layers, unscaled residual GCN/SAGE collapse to 16–22% accuracy through gradient expansion, not oversmoothing. Keeping the residual mass small prevents the collapse: Fixup, ReZero, and a fixed logarithmic schedule all reach the same 69.5–69.8% plateau. The allocation supplies the margin above that plateau, reaching % (+1.7 over a fixed profile at an identical budget), ahead of GCNII, LayerScale, and Gradient Gating. On this benchmark the margin comes from the allocation's shape, not its training; enforcing the a-priori bound costs 1.8 points. On `ImageNet-1k`, where residual networks already train stably, the budget costs 0.5 top-1 points on ResNet-50 and 1.8 on DeiT-Tiny, where vanilla, LayerScale, and a DeepNorm-style rule are more accurate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.