Are We Measuring the Right Angles? Spherical Test-Time Adaptation of Vision-Language Models
Abstract
Test-time adaptation (TTA) for vision–language models (VLMs) aims to improve robustness under distribution shifts by exploiting unlabeled test data during inference. Existing CLIP-TTA methods either optimize prompts through backpropagation or estimate Euclidean statistics from normalized CLIP features, despite CLIP performing classification through cosine similarity on the unit hypersphere. We propose STRA, a Spherical Test-Time Robust Adaptation framework for closed-form, backpropagation-free online CLIP adaptation. Our central view is that closed-form CLIP-TTA should recover a target-domain angular coordinate system under a pullback-vMF model, rather than fit an ambient Euclidean classifier. STRA instantiates this view through label-free geometry calibration from streaming second-order statistics and text-anchored spherical prototype updates driven by sparse soft pseudo-labels. This design avoids raw-feature caching and the selection bias induced by confidence-filtered geometry, decouples marginal geometry estimation from semantic pseudo-label updates, and couples them only through a shared bilateral angular map, while maintaining a single fixed configuration across datasets and backbones. Extensive experiments across natural distribution shifts, fine-grained recognition, and corruption robustness demonstrate consistent improvements over recent CLIP-TTA baselines with favorable computational efficiency. The code is available at https://anonymous.4open.science/r/STRA. baselines at favorable computational cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.