Breaking Balance for Speed: Rescaling Normalization at Initialization
Abstract
Gradient flow in homogeneous neural networks is known to preserve a conservation law between adjacent weight matrices, which follows from a known rescaling symmetry that has strong implications for the theoretical analysis of deep learning dynamics. Normalization breaks this symmetry, yet, as we show, its affine parameters introduce another conservation law with the downstream weights that determines how quickly the two parameter groups learn relative to one another. Contemporary initialization schemes bring them out of balance, which typically induces significant benefits for generalization. Leveraging this insight, we deliberately rescale the initialization to move normalization and weight parameters further out of balance and control their relative learning speeds. Across convolutional networks and Transformers, we find that this can accelerate optimization and yield consistent performance gains over standard and balanced initializations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.