Gradient Coherence Is Not Enough: Separating Influence from Utility in Continual Test-Time Adaptation
Abstract
Continual test-time adaptation (CTTA) updates a model on unlabeled test data, making it difficult to determine which gradient directions are useful for adaptation. Recent approaches exploit consistency in gradient history to identify structured update directions, based on the intuition that gradients sharing consistent structure provide more reliable adaptation signals. But does a consistent gradient direction actually lead to better predictions? We study this question through a mechanistic analysis of gradient-based CTTA, separating three properties of an update direction: its coherence with past gradients, its influence on model behavior, and its utility for the task. We first establish that gradient coherence is not merely a geometric property: differences in gradient structure propagate through the tracked subspace to actual updates, leading the model toward substantially different parameter and prediction trajectories. Having established that coherent gradient structure can meaningfully influence adaptation, we then ask whether this influence translates into task improvement by comparing candidate directions from identical model states. Across five independently permuted trajectories and 15 frozen states, the more coherent direction in a pair is also the more influential one in 68.8% of comparisons. In contrast, the more coherent direction provides greater supervised improvement in only 47.5% of comparisons, showing no corresponding positive utility-ranking pattern. These results show that gradient coherence captures real influence over adaptation, but influence is not the same as utility: a consistent direction can strongly change the model without being a better direction for improving its predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.