acceptodds
Under review as a conference paper at ICLR 2027

Who Deserves the Reward? Shapley Credit Optimization for Multi-Agent LLMs

Abstract

Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising paradigm for decomposing and solving complex problems. However, training multi-agent LLM systems with tool use remains challenging due to the credit assignment problem, as it is often unclear which agent interactions contribute to the success or failure of decision trajectories. Existing LLM agent training methods typically rely on globally broadcast rewards or augment them with heuristic or LLM-judge-based local feedback, failing to capture marginal contributions from individual agents and leading to inefficient reinforcement learning. To address these limitations, we introduce the Shapley-based Hierarchical Attribution for Reinforcement Policy (SHARP), a credit-aware optimization framework tailored to multi-agent LLM reinforcement learning with tool use. SHARP stabilizes training by normalizing agent-specific advantages across trajectory groups and decomposes the reward into three components: a global broadcast accuracy reward for task alignment, a Monte Carlo Shapley based marginal credit reward for agent-level attribution, and a tool format reward for valid tool invocation. Extensive experiments across real-world LLM agent benchmarks demonstrate that SHARP significantly outperforms recent baselines, achieving average match improvements of 35.01% and 24.97% over single-agent and multi-agent LLM agent approaches, respectively. The implementation is available at https://anonymous.4open.science/r/SHARP-DDB4.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.