Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations
Abstract
Traditional domain generalization (DG) methods commonly encourage domain-invariant representations to improve out-of-distribution (OOD) performance. However, we find that this does not transfer directly to finetuning CLIP, a widely used foundation model. Applying invariance regularization during finetuning gives little or no benefit on OOD data compared to standard finetuning. We hypothesize this is because invariance and the preservation of pretrained features can conflict with each other. Domain invariance discourages reliance on domain-specific features, but in foundation models these features were learned during pretraining and are part of what makes the model useful. Retaining domain-specific knowledge would require domain awareness, which standard invariance methods do not provide. To address this, we propose CLIP-DCA (**D**isentangling **C**lassification from enhanced domain **A**ware representations), which encourages invariance only at the classification head through disentanglement, while domain awareness is encouraged in the encoder using a separate domain head supervised by synthetic domain images and style descriptions generated by a multimodal LLM (MLLM). Rather than removing all domain information from the classifier, CLIP-DCA balances invariance with awareness, as we find that pushing the two heads further apart lowers accuracy. Since recent work shows that standard DG benchmarks can be contaminated by CLIP's web-scale pretraining, we evaluate on both established out-of-pretraining (OOP) splits and a 33-dataset benchmark indexed by a continuous multi-modal OOD score. CLIP-DCA shows improvements on both. Our results suggest that, for CLIP, traditional DG principles remain useful when invariance is balanced with domain awareness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.