Decoupling Geometry and Semantics for Test-Time Adaptation of Vision-Language Models
Abstract
Test-time adaptation (TTA) improves model robustness to distribution shifts by adapting to unlabeled test data during inference. For vision-language models (VLMs) such as CLIP, our analysis reveals that distribution shifts alter two structures: the global second-order geometry of visual representations and the class-wise alignment between visual and textual semantics. However, existing VLM-TTA methods either focus primarily on prediction-level adaptation or estimate target statistics through prediction-dependent class assignments, without explicitly separating these two adaptation targets. Thus, erroneous class assignments can jointly distort geometric and semantic estimates under domain shifts. Motivated by this observation, we propose Geometry-Semantics Decoupled Test-Time Adaptation (GeoSem-TTA), a closed-form Gaussian-LDA framework that decouples global geometric adaptation from class-wise semantic adaptation. Specifically, GeoSem-TTA recursively tracks target geometry from reliability-weighted class-independent second-order observations, while adapting class-wise semantic locations through reliability-gated evidence accumulation and coupled centroid synthesis. The resulting statistics directly reconstruct a target-aware Gaussian Linear Discriminant Analysis classifier without parameter updates, backpropagation, or a growing feature bank. Experiments across cross-domain, cross-dataset, and OOD benchmarks demonstrate consistent improvements over strong VLM-TTA baselines, while running 2.3 faster than DLAE under the same setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.