Gradient Entropy Dynamics Characterize Model Adaptability
Abstract
Gradients guide how a neural network learns, but can they also reveal how well a pretrained model will adapt to a downstream task? We uncover a systematic rise in the Rényi entropy of gradients across neuron units as task learning begins. This rise occurs when dominant gradient components diminish faster in relative terms, distributing gradient mass more evenly across units. Beyond this observation, we establish a theoretical connection between gradient Rényi entropy and adaptation dynamics, showing that higher gradient entropy yields a smaller upper bound on the remaining loss gap. This connection provides a theoretical basis for assessing adaptability through gradient entropy and motivates a Rényi-1/2 gradient-entropy score for ranking pretrained models, as well as a valley-restricted variant that limits the influence of high-gradient-energy regions. Experiments with CNN- and ViT-based models across ten downstream datasets show that the proposed score achieves stronger average ranking correlations than the compared transferability estimators. Even without labels, the method remains competitive with, and can outperform, most label-based alternatives. These findings highlight gradient entropy as an informative measure of pretrained-model adaptability to downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.