Which Sources Should Participate? Target-Specific Multi-Source Configuration
Abstract
Multi-source learning often treats source participation as fixed, either using all available sources or ranking sources independently for a target. We instead study source participation as a target-specific configuration problem over non-empty subsets of complete source domains. Empirical analysis reveals three recurring patterns. Adding sources can reduce target performance even when more training data become available, near-tied validation configurations frequently reverse on test data, and larger configurations do not reliably outperform smaller ones when validation support is comparable. These observations motivate ARCS, an Admissible Risk-aware Configuration Selection rule that constructs a near-optimal validation set, resolves ambiguity through parsimony, and uses worst-class validation performance only as a localized safeguard. We evaluate ARCS across multiple real-world heterogeneous multi-source benchmarks spanning wearable sensing, clinical multi-site data, and multimodal sentiment analysis, where it remains competitive while selecting compact source configurations. Further analyses show that the main gain comes from validation admissibility combined with parsimony, that performance remains stable under reduced target evidence, and that approximate candidate search can reduce configuration evaluation without changing the final selection principle. Overall, our results suggest that source participation should be treated as a target-specific configuration decision rather than a fixed assumption in multi-source learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.