Beyond Holistic Representations: Dual Distribution Modeling for Test-Time Adaptation of Vision–Language Models
Abstract
Although Vision-Language Models (VLMs) exhibit strong zero-shot generalization, their performance degrades substantially under domain shifts and out-of-distribution (OOD) scenarios. Test-Time Adaptation (TTA) addresses this issue using only unlabeled test samples, among which training-free TTA is particularly appealing for its efficiency. Existing training-free methods refine predictions using cached samples or representation distributions, but their reliance on holistic image representations limits the extraction of more semantically meaningful features. To this end, we decompose visual features into in-distribution features and class-related OOD features based on their alignment with textual semantics, and empirically show that both exhibit clear class-wise clustering in the embedding space. Based on this finding, we propose Dual Distribution Modeling for Adaptation (DDMA), a novel training-free TTA framework that separates these two feature types to better exploit semantic information while reducing interference from semantically irrelevant features. DDMA independently models their class-wise distributions during test time and jointly refines the original predictions by their distribution statistics. Extensive experiments across cross-domain and OOD benchmarks show that DDMA outperforms state-of-the-art TTA methods, achieving average accuracy gains of 3.37% and 4.02% over the pre-adaptation performance. Moreover, DDMA maintains competitive efficiency with an average throughput of 29.37 items/s.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.