d-STABLE: Learning Residual Allocations Under a Controlled Depth Budget
Abstract
The sensitivity of a residual network—how much it amplifies a small input change across all layers—is a product over layers, so it is governed by one quantity: the total *residual budget*, the summed strength of the per-layer residual branches. We show that this budget and its allocation across layers are two separable levers. The budget's growth rate in depth sets the *stability class*: a linear budget makes worst-case sensitivity exponential, a square-root budget (Fixup, Stable ResNet) sub-exponential, and a *logarithmic* budget polynomial, so is the critical rate. The allocation leaves the class unchanged and sets accuracy within it. Prior stabilizers couple the two: fixed schedules freeze both, and learned gates (ReZero, LayerScale) let the budget drift. **d-STABLE** is a reparameterization that holds the total budget at exactly at every optimization step while the optimizer learns the allocation. When every branch is norm-bounded this gives an a-priori polynomial bound on sensitivity; otherwise the pinned budget is an inductive bias whose amplification we measure. Our experiments separate the two levers. On `ogbn-arxiv` at 64 layers, unscaled residual GCN/SAGE collapse to 16–22% accuracy through gradient expansion, not oversmoothing. Keeping the residual mass small prevents the collapse: Fixup, ReZero, and a fixed logarithmic schedule all reach the same 69.5–69.8% plateau. The allocation supplies the margin above that plateau, reaching % (+1.7 over a fixed profile at an identical budget), ahead of GCNII, LayerScale, and Gradient Gating. On this benchmark the margin comes from the allocation's shape, not its training; enforcing the a-priori bound costs 1.8 points. On `ImageNet-1k`, where residual networks already train stably, the budget costs 0.5 top-1 points on ResNet-50 and 1.8 on DeiT-Tiny, where vanilla, LayerScale, and a DeepNorm-style rule are more accurate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.