Beyond Feature Shift: Measuring and Mitigating Output Distribution Collapse in Cross-Domain VQA
Abstract
Although vision-language models (VLMs) achieve strong performance on general-domain benchmarks, their effectiveness often degrades substantially when applied to domain-specific visual question answering (VQA) tasks such as medical imaging, remote sensing, and art understanding. Existing studies mainly attribute this degradation to feature-level distribution mismatch and focus on representation alignment. In this work, we show that such explanations are insufficient to characterize the failure mechanisms of pretrained VLMs under domain shift. Through a systematic study across six domain-specific VQA benchmarks and multiple VLM architectures, we identify a previously overlooked phenomenon termed output distribution collapse, characterized by prediction distribution collapse toward dominant answer prototypes, confidence miscalibration, and decision-margin instability. To quantify this phenomenon, we introduce an entropy-based Collapse Index (CI) that measures the severity of predictive distribution degeneration. We show that CI consistently exhibits a strong negative correlation with cross-domain VQA accuracy across datasets and architectures. Motivated by these findings, we propose Distribution Collapse Stabilization (DCS), a lightweight output-space learning framework that jointly promotes distribution dispersion, reliability consistency, and decision separability. Extensive experiments on six benchmarks and ten representative VLM backbones demonstrate consistent improvements, yielding an average gain of 4.1% over pretrained baselines while outperforming other adaptation methods. Our findings identify output distribution collapse as a previously underexplored failure mechanism in cross-domain VQA and establish output-space stabilization as a promising alternative to conventional feature-space adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.