Learning Subset-Shared Invariances for Domain Generalization with Mixture-of-Experts
Abstract
Domain generalization (DG) aims to learn a model from one or more source domains that generalizes to an unseen target domain without accessing target data during training. A common approach enforces invariance of representations across all source domains, assuming predictive structure is globally shared. However, exact conditional invariance can exclude predictive factors shared only within subsets of source domains. To address this limitation, we propose subset-shared invariance, where predictive structure is assumed stable only within domain subsets. We implement this principle with a mixture-of-experts architecture, where routing induces expert- and class-specific alignment strengths across source-domain pairs and softly composes the resulting expert representations for prediction. This creates routing-conditioned alignment that prioritizes domain relationships according to learned responsibilities rather than enforcing them uniformly. To facilitate effective decomposition, we develop training objectives that encourage selective alignment, confident and balanced routing, and diverse expert specialization. Experiments on DomainBed benchmarks demonstrate improved out-of-domain generalization and greater robustness under increasing domain heterogeneity. Our results suggest that DG should move beyond enforcing a single global invariance and instead model invariance through partially shared structure across domain subsets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.