MAMR-GRPO: Reinforcement Learning with Execution-Structured Credit for Multi-Agent LLMs
Abstract
LLM-based multi-agent systems (MAS) extend the capabilities of language models by decomposing complex tasks into specialized roles, enabling parallel reasoning, information gathering, and iterative refinement. In such systems, the final answer is produced collaboratively by multiple agents or roles whose outputs interact and influence subsequent executions. However, prevailing reinforcement learning methods such as Group Relative Policy Optimization (GRPO) typically assign credit at the rollout level, making it difficult to distinguish the contributions of different executions within a multi-agent trajectory. We propose Multi-Agent Multi-Round GRPO (MAMR-GRPO), which retains rollout-level comparisons while introducing local credit assignment aligned with the execution structure of MAS. Specifically, round boundaries, role dependencies, and parallel outputs provide references for comparing task assignment, subtask execution, and information aggregation. These local credits are combined with the terminal rollout advantage and applied to the corresponding generated segments of a shared LLM policy. Across seven search-augmented QA benchmarks, MAMR-GRPO reaches 46.7 average exact match, 3.1 points above trajectory-level GRPO on the same graph. Budget ablations show further gains with additional subagents and rounds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.