acceptodds
Under review as a conference paper at ICLR 2027

ProCore: Protecting Task-Specific Gradient Cores for Multi-Task RLVR

Abstract

Multi-task RLVR enables a single model to jointly learn multiple tasks. We find that each task’s gradient is highly concentrated, with a small subset of components accounting for most of the gradient norm. We term them gradient cores. Existing methods dynamically adjust task sampling ratios or weights using reward signals or advantage estimates to balance learning progress across tasks. However, they lack explicit protection for gradient cores during gradient aggregation and struggle to characterize each task’s remaining learning potential. Accordingly, we propose ProCore, a plug-and-play method for multi-task RLVR. ProCore identifies gradient cores by magnitude within each Transformer block and partitions them into exclusive and shared cores based on dimension overlap across tasks. It computes a Remaining Potential Score (RPS) from the gap between current pass@1 and an estimated performance ceiling derived from pass@k. Building on the original mixed gradient, ProCore directly enhances exclusive cores and dynamically weights shared cores using RPS, giving greater protection to tasks with more remaining potential. We evaluate ProCore on eight heterogeneous reasoning tasks and nine benchmarks spanning mathematics, coding, and scientific QA. ProCore achieves higher average scores than all baselines across both evaluation settings and further improves DAPO and MT-GRPO when combined with them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.