RACA: Role-Aware Credit Assignment for Multi-Agent LLM Reasoning
Abstract
Recent research has sought to enhance LLM reasoning through collaboration among specialized agents. Training such systems, however, requires accurate credit assignment across roles and decisions. Existing outcome-level reinforcement learning assigns the same terminal reward to every turn, allowing an incorrect action to receive positive credit when later agents repair it. We introduce Role-Aware Credit Assignment (RACA), which aligns credit with each role's responsibility and decision scale. RACA retains trajectory-level credit for controller decisions, evaluates worker actions using locally observable outcomes, and computes relative advantages only among turns matched by role, blackboard-derived context, and reward channel. For primary proposer responses, it separately routes solution-quality and assistance-seeking advantages to their corresponding token spans. All signals come from the original rollout, without an additional value model or counterfactual execution. Using Qwen3-8B, RACA achieves 64.5% accuracy on MATH Level 5 and 16.0% on AIME 2022–2026, outperforming all compared baselines on both benchmarks. Under the same adaptive multi-agent framework, it improves over Outcome GRPO by 6.4 and 6.0 percentage points, respectively. On MATH, RACA averages 7.4 model calls per problem, compared with 8.7–10.7 for the other adaptive methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.