MASHAP-LLM: Contribution-Aware Policy Optimization for Multi-Agent LLM Post-Training
Abstract
Multi-agent large language model (LLM) systems face a credit assignment problem during training. Shared rewards give every agent the same feedback regardless of its contribution, so policy updates have no explicit signal of which agent's behavior to reinforce. Measuring each agent's contribution by rerunning the task with subsets of agents, called coalition rollouts, is costly. We propose MASHAP-LLM, a multi-agent LLM post-training algorithm that rewards each agent according to its contribution and needs coalition rollouts only early in training. MASHAP-LLM builds on proximal policy optimization (PPO) and assigns rewards using Shapley values. At its core is LLM-SHAP, a learned component that reads the task and the agents' trajectories and estimates the coalition values used to compute these contributions. LLM-SHAP is trained on coalition rollouts scored by a verifier during an early fitting phase and then frozen. After freezing, LLM-SHAP scores each coalition from the full-team trajectory with the outputs of absent agents masked, so no smaller team is generated or verified. MASHAP-LLM turns the resulting Shapley credits into per-agent rewards through a rescaling that preserves the total team reward. We evaluate 3B models from the Qwen2.5 family on three math benchmarks and three coding benchmarks. Against the strongest baseline in each of nine settings, MASHAP-LLM improves the average score by 2.51 percentage points with two agents and 4.36 points with three agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.