Alignment and Convergence Analyses of Gradient Flow in Two-Layer Convolutional Neural Networks
Abstract
Understanding why gradient-based training of neural networks learns to extract signal and ignore noise remains a central open problem in deep learning theory. We take a step toward this goal by analyzing the gradient-flow dynamics of two-layer ReLU convolutional networks with sum-pooling, trained on the exponential or logistic loss for binary classification from a small, balanced initialization. Each input consists of one signal patch and some noise patches; the signal patches of the two classes lie in two well-separated, nearly antipodal cones, so that the label-signed signal patches share a common axis , and the noise patches are orthogonal to the signal subspace. Our main result indicates that during the early phase of the training, every filter that activates a signal patch of its own class at initialization aligns with either or , depending on the sign of its second-layer weight, to a prescribed cosine and within an explicit time determined by the signal geometry; such a filter thereafter activates every signal patch of its class and none of the other class. The alignment is driven by a self-reinforcing feedback mechanism. Since same-class signal patches lie in a narrow cone, every activated signal patch pushes the filter toward all the others, thus toward the axis. The noise patches, being orthogonal to the signal subspace, enter the analysis only through a single scalar that bounds their aggregate effect. In the second phase of the training ,the projection of the aligned filter onto grows monotonically, and the training loss decays at rate . Experiments on synthetic and MNIST data confirm the two phases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.