SARA: Semi-Supervised Multi-Domain Learning of Vision-Language Models without Test-Domain Labels
Abstract
Vision-language models (VLMs) such as CLIP provide a promising foundation for multi-domain learning (MDL) by introducing rich multimodal semantics and transferable pretrained knowledge. However, existing MDL methods that preserve domain-specific differences often rely on test-domain labels to select specialized parameters or computation paths. Without such labels, domain-specific information can provide useful cues for some samples but induce misleading prediction variation for others. We refer to changes in such information for a fixed image as domain perturbations: some predictions remain stable, whereas others become unstable and may even undergo prediction flips. Moreover, most MDL approaches rely heavily on labeled data, while existing semi-supervised learning methods do not explicitly account for such domain-specific prediction variation, which can aggravate confirmation bias during self-training. To address these challenges, we propose Stability-Aware Routing and Adaptation (SARA), a CLIP-based semi-supervised MDL framework without test-domain labels. SARA assesses sample stability under domain perturbations through predictive consistency and decision margin, and consists of two coordinated components: (i) Stability-Guided Self-Training applies a conservative stability assessment to filter unreliable pseudo-labels; and (ii) Stability-Based Prediction Routing exploits domain-specific information for stable samples while reducing its influence on unstable samples through sample-specific prediction mechanisms. Extensive experiments on four multi-domain benchmarks demonstrate consistent improvements over MDL and SSL baselines, including a 1.39% average gain over the state of the art.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.