acceptodds
Under review as a conference paper at ICLR 2027

Gradient Entropy Dynamics Characterize Model Adaptability

Abstract

Gradients guide how a neural network learns, but can they also reveal how well a pretrained model will adapt to a downstream task? We uncover a systematic rise in the Rényi entropy of gradients across neuron units as task learning begins. This rise occurs when dominant gradient components diminish faster in relative terms, distributing gradient mass more evenly across units. Beyond this observation, we establish a theoretical connection between gradient Rényi entropy and adaptation dynamics, showing that higher gradient entropy yields a smaller upper bound on the remaining loss gap. This connection provides a theoretical basis for assessing adaptability through gradient entropy and motivates a Rényi-1/2 gradient-entropy score for ranking pretrained models, as well as a valley-restricted variant that limits the influence of high-gradient-energy regions. Experiments with CNN- and ViT-based models across ten downstream datasets show that the proposed score achieves stronger average ranking correlations than the compared transferability estimators. Even without labels, the method remains competitive with, and can outperform, most label-based alternatives. These findings highlight gradient entropy as an informative measure of pretrained-model adaptability to downstream tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.