WHEN DOES TASK AFFINITY JUSTIFY PAIRWISE MULTI-TASK CONFIGURATION SELECTION?
Abstract
Task-affinity scores in multi-task learning are typically validated by their correlation with multi-task gain and then used to decide which tasks share what. We argue that these are different claims: a score (diagnostic) can track a real mechanism, such as representation reuse or output coupling, without determining the right configuration decision. We decompose this inference into three links: score → intervention effect → preferred configuration → out-of-sample selection gain, and test them on 1914 ordered pairs from 172 tasks spanning CelebA, QM9, and 16 multi-output regression families. The first link holds selectively: a direct reuse probe tracks frozen representation reuse and residual dependence tracks output coupling, but no diagnostic reliably predicts the value of adaptation. The second is context-dependent: with the measured task relation nearly unchanged, replacing a small CNN with ResNet-18 reduces adapted-transfer selection from 30% to 7%, while preferences also shift with the candidate set, metric, and target-data budget. A fixed portfolio of two or three configurations reduces oracle headroom below replicate noise on CelebA and QM9. Finally, learned selectors that beat the best fixed configuration on held-out pairs generalize to held-out tasks only on CelebA under the log score, while family extrapolation remains unsupported. We distill these results into principles for more reliable task-affinity evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.