acceptodds
Under review as a conference paper at ICLR 2027

Candidate-Conditioned Alignment for Stable Test-Time Adaptation in Cross-Domain VQA

Abstract

Vision-language models (VLMs) have achieved remarkable progress in visual question answering (VQA), yet their performance often degrades under domain shifts. Test-time domain adaptation (TTDA) offers an attractive solution by adapting models directly on unlabeled target data during inference. Existing TTDA methods, however, typically optimize prediction confidence over a shared answer space, an assumption that is inconsistent with the question-dependent prediction structure of VQA. As a result, optimization may progressively become biased toward dominant answers and produce question-insensitive predictions under domain shift. To address this issue, we propose a Candidate-Conditioned Alignment framework for VQA test-time domain adaptation. Instead of adapting over the global answer space, our method performs adaptation within instance-specific candidate sets through a local energy-based image-text alignment objective, enabling stable optimization while preserving question-conditioned reasoning. A lightweight text adapter further supports efficient adaptation without updating the pretrained backbone. Extensive experiments on six vision-language models and six cross-domain VQA benchmarks demonstrate consistent improvements across diverse domains, achieving an average accuracy gain of 18.8% over direct candidate-based inference under the same candidate setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.