acceptodds
Under review as a conference paper at ICLR 2027

Learning Across Roles: Role-Balanced Shared-Policy Optimization for Multi-Agent Code Generation

Abstract

Shared-policy multi-agent learning enables collaborative code generation without maintaining a separate policy for each role. However, common actor-loss reductions implicitly assign greater weight to roles that generate longer responses or are invoked more frequently, coupling optimization to workflow exposure. We propose Role-Balanced Shared-Policy Optimization (RBPO), which jointly trains a Planner, a Coder, and a Reflector through a single role-conditioned policy in an execution-guided refinement loop. RBPO averages token losses within each response and responses within each role, then combines the resulting role objectives with explicit weights. This hierarchical objective makes each represented role's nominal weight independent of its response lengths and invocation count, without adding role-specific policy parameters or rollout calls. On LiveCodeBench (LCB-v6), RBPO with Qwen3-8B achieves 60.7% accuracy, improving over Qwen3-8B GRPO by 9.8 percentage points and the same workflow without reinforcement learning by 10.4 points. It also exceeds a strong learning-based multi-agent method using Qwen3-14B by 1.3 points under the evaluated inference configurations. Controlled ablations support the contribution of both sequence and role balancing, while role-removal analyses demonstrate the importance of reflection despite its lower invocation frequency. These results highlight explicit role weighting as an effective design choice for shared-policy multi-agent code generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.