Normalization-Free Residual Networks by Bounding Weights and Residuals
Abstract
Normalization layers are widely used for stabilizing deep residual networks, yet they incur computational overhead and obscure training dynamics. We propose a stable, normalization-free residual block rooted in constrained nonlinear optimization that replaces traditional activation normalization with an inexpensive, element-wise clipping operation on the output of each residual branch. We extend this framework to weights, combining it with a parameterization that enables a single global constant to bound the magnitudes of both activations and weights. This parameterization additionally results in controllable and uniform effective learning rates across all layers. Compared to networks that do not use clipping, our approach yields activation and weight distributions that are remarkably smooth and free of outliers, enabling direct post-training quantization of weights to INT4 with only a minor loss degradation. Our experiments demonstrate that this approach successfully eliminates the need for both activation and weight normalization, offering a simple and principled foundation for training deep residual networks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.