LAST: Stabilizing Test-Time Adaptation via Logit Standardization
Abstract
Test-time adaptation (TTA) improves model performance under distribution shift by leveraging unlabeled test data during inference. However, existing TTA approaches often suffer from model collapse, where predictions progressively concentrate on a limited subset of classes. Despite its prevalence, the underlying mechanism of this failure mode remains insufficiently understood. In this work, we investigate the dynamics of TTA collapse and identify logit norm instability as a key factor driving prediction degeneration. We show that adaptation-induced logit norm drift disturbs the balance of prediction confidence across classes, amplifying inherent prediction biases through a self-supervised process. Specifically, excessively increased logit norms lead to overconfident predictions that strengthen erroneous class preferences, while reduced logit norms weaken class separability and impair stable adaptation. Based on this insight, we propose **L**ogit **A**daptation via **S**tandardization at **T**est time (**LAST**), a lightweight and plug-and-play module that performs class-wise logit standardization before the layer. By normalizing the distribution of logits for each class across test samples, LAST suppresses instance-dependent logit norm fluctuations and regulates prediction confidence. Extensive experiments across disserve benchmarks show that LAST consistently improves existing TTA methods and has ability to alleviate collapse, achieving up to a 36.5% improvement in Top-1 accuracy over TENT baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.