When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
Abstract
Outcome-reward reinforcement learning (RL) can train every role of a multi-agent LLM workflow end to end. Its payoff is hard to predict, because the reward does not credit individual roles and a shared policy trades pooled data against specialization. We ask when this training improves a workflow over its base model and single-agent RL, and what each role learns. We train three workflows (Eval-Opt, Voting, Orch-Workers) with GRPO on math and code at three Qwen3 scales, with one policy shared by all roles (Shared-Policy, SP) or one per role type (Isolated-Policy, IP). RL lifts every workflow above its untrained version. The solving roles, which write candidate solutions, end no better than single-agent RL, so a workflow exceeds single-agent RL only through the structure above them, whose gain is set by the simple rule its decision role learns. Eval-Opt gains most because its revision loop gives truncated first attempts another turn, extending the token budget. IP learns faster and peaks higher than SP. Under SP, averaging each output's loss over its tokens lets a decision role with much shorter outputs take well over its uniform share of the shared gradient, and the solving roles learn slowly. Under IP, the adapter of a role run three times per episode can become unstable after the peak, and only that role's policy-ratio and perplexity rise. A stopping rule on training success fires only after the peak but can return a checkpoint near it, and SP runs decline little.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.