Understanding Representation Alignment in Neural Networks through Normalized Gradient Dynamics with Periodic Resets
Abstract
Understanding how neural networks develop useful representations during training remains an open problem. We study this question by introducing feature-target alignment, a scale-invariant quantity that measures how well a learned feature map captures the target function. We show that normalized feature learning initially induces an alignment-improving direction in the training. Motivated by this observation, we propose normalized gradient descent with periodic resets, which repeatedly returns training to a regime that promotes target-aligned representation learning. We prove that the resulting updates monotonically improve feature-target alignment for a broad class of parameterized feature maps over any given finite horizon, in both population and finite-sample settings. We further study a full-matrix representation model and show that the same learning principle strictly decreases the misspecified risk, providing a concrete example in which representation learning improves prediction. We also provide empirical support for our theoretical results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.