acceptodds
Under review as a conference paper at ICLR 2027

Can domain synergy mislead? Measurement confounds in multi-domain RLVR interaction estimation

Abstract

Multi-domain RLVR post-training mixes mathematical reasoning, coding, agentic tool use, and general knowledge in one training run, and particular domain combinations are often credited with synergy. We ask whether such synergy can be measured reliably. We audit the full coalition lattice of four RLVR domains (GOPO, Qwen3-1.7B), evaluating every coalition on a fixed source suite used for selection and an independent transfer suite standing in for deployment, and decompose the gains into Mobius interaction dividends. Uncorrected, the lattice shows significant higher-order interaction that grows with order. We show that this pattern is largely a measurement artifact. When a fixed total budget is split across a coalition's domains, each domain's training depends on coalition size, which produces spurious higher-order dividends, and Mobius sums amplify training noise with order. After correcting for the budget split and accounting for seed variance, little higher-order interaction remains on the transfer suite, while the source suite is not fully resolved. The practical risk lies elsewhere: the coalition that looks best on the selection suite is often not the best on the deployment suite, and a single lattice cannot reliably identify the best mixture. We also examine a curvature-based account of cross-suite transfer and propose GIFT, an adaptive design for choosing which coalitions to train.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.