RA-TPT: Stopping Logit Inflation at Test time
Abstract
Vision-Language Models (VLMs) enable strong zero-shot transfer, yet their predictive probabilities can become miscalibrated under distribution shift. Test-Time Prompt Tuning (TPT) adapts continuous prompts using unlabeled target inputs and can improve target-domain accuracy; however, its entropy-minimization objective may increase predictive concentration for incorrect or ambiguous samples. We propose Reliability-Aware Test-Time Prompt Tuning (RA-TPT), a logit-space approach for improving the calibration-accuracy trade-off during test-time adaptation. RA-TPT replaces entropy minimization with an uncertainty-aware bounded-margin objective and a pseudo-label anchor that provide a self-training signal while limiting margin-expansion pressure on uncertain inputs. It further applies uncertainty-routed centered-logit dispersion regularization to discourage excessive relative logit variation when predictions are uncertain, with a complementary mean-logit regularizer. A soft conflict-aware gradient projection attenuates the regularization-gradient component that opposes the separability gradient. Across 31 benchmarks spanning standard classification, ImageNet distribution shifts, biomedical imaging, and remote sensing, RA-TPT improves the accuracy-calibration trade-off in CLIP-style adaptation, substantially reducing Expected Calibration Error (ECE) versus entropy-based TPT and matching calibration-aware baselines at competitive accuracy. Because the objective acts purely on logits, it further transfers to non-VLM test-time adaptation, matching strong corruption-robustness baselines on CIFAR-10/100-C and ImageNet-C.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.