Correlations Are Ruining Your Gradient Descent
Abstract
Natural gradient descent, and the preconditioned optimizers which have followed from it, correct gradient descent by estimating curvature and applying its inverse to the gradient. For any single layer of a deep neural network, this correction can be shown to approximately separate into two factors: one which accounts for correlations in the gradient signal, and one which accounts for correlations in the data arriving at that layer. Herein we show that this second factor can be expressed as a change of basis rather than of curvature and, considered thus, it can be removed from the optimizer entirely and placed within the network itself. Specifically, we show that, once forward weights are corrected, right-sided natural gradient descent can be implemented in a form which has three exactly equivalent forms: a gradient preconditioner, a decorrelating linear transform in the forward pass, or a set of recurrent dynamics in which all updates are Hebbian and rely only upon locally available signals. The (recurrent) decorrelation method we discover aligns closely with proposed methods for inhibitory learning and homeostatic plasticity in the brain. Notably, our transform requires no matrix inverse, no matrix decomposition, and no information beyond that which is available at an individual synapse. It can therefore be applied in distributed and physically constrained systems, including neuromorphic hardware and nervous systems, where existing preconditioners cannot. We demonstrate its effectiveness empirically in convolutional and residual network models while significantly improving a range of approximate biologically plausible methods for credit assignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.