When Should a Domain-Generalized Model Specialize?
Abstract
Limited-label adaptation raises two questions: whether a fixed policy improves a domain-generalized model, and whether episode-specific policy selection provides additional value. The Cross-Fitted Domain-Generalization Specialization Assessment (CF-DGSA) separates fixed adaptation gain, held-out selection premium, and total gain over zero-shot inference. Its comparator is the accuracy-maximizing fixed policy on a labeled audit panel, while a routing margin affects only selection. Symmetric cross-fitting and a matched-half diagnostic distinguish held-out value from same-fold reuse. Selection-premium intervals project paired component intervals, with operating characteristics evaluated under Gaussian and paired-image models. Across nine configurations on PACS, Office-Home, and VLCS, four selection-premium intervals are positive, but only three corresponding total-gain intervals are positive. VLCS ResNet-50 has a selection premium and a total gain: improved selection among adaptation policies can still underperform zero-shot inference. Reproduced ERM++ models further distinguish beneficial fixed adaptation without resolved selection value, observed policy degeneracy, and harmful adaptation. The assessment separates evidence for adaptation from evidence for policy selection under explicit audit-label access.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.