Normalization and the Scalar Structure of Dropout
Abstract
Dropout randomly changes individual features in a neural network. We ask when its effect after normalization can be approximated by one common scalar rescaling. For a trained network and a given input, normalization divides the features by a shared factor. When many small independent contributions determine this factor, its relative variation can become small. We replace it by a value computed directly from the features and dropout probability. At this fixed radius, the scalar minimizes mean squared error. It rescales the part of the output affected by normalization and leaves the fixed part unchanged. We then quantify how much of the change caused by dropout this scalar removes in mean square and how much random variation remains. Under replication that preserves predictions and uses independent copy masks, the output error is bounded in probability at the inverse square root of the number of copies. For RMS normalization, a nonzero limiting covariance rules out a faster rate. Experiments on MNIST networks, a pretrained FaceNet embedding, and ten language models test the theory. In the language models, scalar rescaling reduces error against Monte Carlo means of probability vectors. In FaceNet, it reduces error against Monte Carlo means of cosine scores. Individual dropout outputs can still remain highly variable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.