Towards Unbiased Test-Time Adaptation for Vision–Language Models
Abstract
Statistics-based test-time adaptation (TTA) provides an efficient way to adapt vision-language models (VLMs) by summarizing test streams into fixed-size class statistics without backpropagation or instance storage. However, there are two limitations in this paradigm. First, persistent class-wise prediction preferences can repeatedly misallocate incoming features, leading to prediction-biased accumulation in the class states. Second, additive feature aggregation can retain class-shared components, leading to classifier homogenization even under correct class allocation. To address these issues, we propose _**U**nbiased **T**est-Time **A**daptation_ (**UTA**), which jointly regulates class allocation and classifier construction. Specifically, to mitigate prediction-biased accumulation, a *History-Guided Allocation Debiasing* mechanism is designed to regulate class allocation using historical class responses, thereby reducing allocation errors in the accumulated class states. Furthermore, to alleviate classifier homogenization, we introduce an *Activity-Conditioned Classifier Construction* strategy to reweight accumulated class states according to stream-wide coordinate activity, which limits shared-component dominance and preserves class-specific distinctions. Extensive experiments demonstrate strong adaptation performance across multiple settings while preserving 95% of zero-shot CLIP throughput on large-scale ImageNet-1K. Moreover, the proposed components can serve as lightweight plug-ins for existing statistics-based TTA methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.