Beyond Task Balancing: Coordinating Experience Groups in Multi-Environment Reinforcement Learning
Abstract
Multi-environment reinforcement learning (RL) trains language agents to retrieve information, use tools, and solve problems within a shared policy. Besides deciding how much experience each environment contributes, joint training must decide how this experience combines into each update. Group Relative Policy Optimization (GRPO) compares trajectories within each problem’s group but averages groups without considering whether their policy gradients reinforce or oppose one another. We introduce Compatibility-guided Group Reweighting (CoGR), which coordinates experience groups while keeping each environment’s rollout quota fixed. Within-group relative rewards decide which trajectories to reinforce, and between-group gradient compatibility decides how strongly each group contributes. CoGR scores each complete group in the current on-policy batch by its environment-balanced gradient agreement with other groups and adds an extra penalty for strongly opposing gradients. Bounded positive weights with unit mean per environment rescale only the policy-gradient loss, preserving within-group comparisons and reusing the collected trajectories. We train a 4B-parameter agent on Search (Natural Questions and HotpotQA), TextCraft, and a restricted CodeContests subset. Across two training seeds, CoGR improves environment-averaged success on the development sets by 2.39 percentage points (pp) over GRPO after 64 policy updates, including a 7.64 pp gain in Code, the weakest environment. In a single-seed comparison at the same endpoint, CoGR exceeds agent-adapted GradAlign-cosine by 5.33 pp and achieves slightly higher environment-averaged success than agent-adapted MT-GRPO with 82.51% fewer generated rollouts. These results suggest that experience-group coordination is a useful way to improve joint learning within prescribed environment allocations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.