Learning Across Roles: Role-Balanced Shared-Policy Optimization for Multi-Agent Code Generation
Abstract
Shared-policy multi-agent learning enables collaborative code generation without maintaining a separate policy for each role. However, common actor-loss reductions implicitly assign greater weight to roles that generate longer responses or are invoked more frequently, coupling optimization to workflow exposure. We propose Role-Balanced Shared-Policy Optimization (RBPO), which jointly trains a Planner, a Coder, and a Reflector through a single role-conditioned policy in an execution-guided refinement loop. RBPO averages token losses within each response and responses within each role, then combines the resulting role objectives with explicit weights. This hierarchical objective makes each represented role's nominal weight independent of its response lengths and invocation count, without adding role-specific policy parameters or rollout calls. On LiveCodeBench (LCB-v6), RBPO with Qwen3-8B achieves 60.7% accuracy, improving over Qwen3-8B GRPO by 9.8 percentage points and the same workflow without reinforcement learning by 10.4 points. It also exceeds a strong learning-based multi-agent method using Qwen3-14B by 1.3 points under the evaluated inference configurations. Controlled ablations support the contribution of both sequence and role balancing, while role-removal analyses demonstrate the importance of reflection despite its lower invocation frequency. These results highlight explicit role weighting as an effective design choice for shared-policy multi-agent code generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.