What Makes a Batch Reliable? Balancing Multi-Domain RLVR with TaylorMix
Abstract
Multi-domain reinforcement learning with verifiable rewards (RLVR) can produce unstable policy updates even with self-generated responses. Our experiments identify batch composition as a key factor in update stability. Selecting high-scoring cases independently can overlook the risks of using them together, while concentrating training in favored domains can limit learning opportunities for other capabilities. These limitations motivate TAYLORMIX, a Taylor-motivated selector centered on balance. Before rollout, it balances domain priorities with continued opportunities across capabilities through KL-regularized allocation, and individual usefulness against composition risk through companion-dependent case selection. Our theoretical analysis characterizes this balance by quantifying the continuous-allocation trade-off and establishing conditional geometric-penalty reduction at a fixed selection prefix. Under matched training budgets, TAYLORMIX achieves the highest category means among the compared methods in mathematical reasoning, code, and general capability on two evaluated backbones. Ablations support both levels of balance, linking capability gains to preserved domain opportunities and geometric selection; the latter also yields fewer large gradient spikes and zero-advantage rollout groups in the observed trajectories. Our anonymous code is available at https://anonymous.4open.science/r/TaylorMix-D280.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.