Forecasting When Test-Time Adaptation Will Fail
Abstract
Test-time adaptation (TTA) updates a pre-trained model on unlabeled test data. Many TTA methods recompute batch-normalization (BN) statistics on every test batch. On skewed batches such as pathology slides which are processed one at a time, this can sharply reduce accuracy. We show, in controlled experiments that keep the whole stream class-balanced, that skewed batches alone cause this failure, that recomputing the BN statistics without any gradient step reproduces it, and that the degradation comes from normalizing a batch with statistics dominated by its own class. We turn these findings into a gate that forecasts, before each TTA update, whether adapting on the current batch will help, based on the batch composition predicted by the model without adaptation and its response to the batch's statistics. The forecast is fitted once on held-out labeled source data and needs no target labels. Across CIFAR-10-C, CIFAR-100-C and two pathology datasets, the gate prevents the collapse of every BN-based method we evaluate on class- and slide-ordered streams, retains most of the benefit on shuffled streams, and improves methods designed for skewed streams on pathology. (Code will be made public upon acceptance)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.