Decomposing the Entropy Minimization Gradient for Calibrated Test-Time Adaptation
Abstract
Entropy minimization (EM) is widely used for test-time adaptation (TTA), but its accuracy improvements under distribution shift often come at the cost of miscalibrated predictions. We study this behavior through the geometry of the EM gradient. After removing the softmax-invariant mean of the logits, we decompose the gradient into a radial component, which increases the centered-logit norm, and a tangential component, which changes relative class scores while preserving the norm to first order. Although these components are orthogonal in logit space, they induce conflicting parameter gradients, particularly in later network layers. Empirically, we find that the tangential component is the main driver of accuracy improvement, whereas the radial component is closely tied to confidence dynamics and becomes relatively stronger in later layers. These observations motivate two complementary interventions: Conflict-Aware Regularized Entropy Minimization (CAREM), which protects the tangential update by projecting conflicting radial gradients onto the tangential gradient's orthogonal complement, and freezing later layers during adaptation to improve calibration. CAREM modifies only the backward pass and introduces no additional hyperparameters or changes to the model architecture or adaptation objective. Across multiple entropy-based TTA methods, architectures, batch sizes, and distribution-shift benchmarks, CAREM improves classification performance, while its combination with late-layer freezing substantially improves calibration and retains strong accuracy. Our results show that accurate and calibrated entropy-based TTA can be achieved with simple changes to its optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.