Structured Sample-Wise Gradients in Neural Network Optimization
Abstract
Stochastic gradient descent (SGD) aggregates gradients from individual training examples, yet the organization of these sample-wise optimization directions remains poorly understood. We analyze per-example gradients using pairwise cosine similarity and clustering throughout training. We find that gradient clusters are internally coherent, with highly aligned within-cluster gradients, while gradients from different clusters exhibit much lower similarity. This structure is reproducible across independently trained models and is stronger in deeper layers. Gradient clusters are strongly aligned with class labels early in training, but this correspondence weakens over time. Besides, lower within-cluster similarity is associated with better generalization performance, and pretrained models begin downstream training with lower within-cluster similarity than randomly initialized models. Our results reveal a structured sample-wise organization underlying SGD updates and connect its evolution to class structure, pretraining, and generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.