Counterfactual Credit Assignment over Execution Graphs of LLM-based Multi-agent Systems
Abstract
Training LLM-based multi-agent systems (MAS) requires assigning credit to agents for how their actions influence subsequent interactions and outcomes. We present RollCredits, a method for assigning decision-level credit through counterfactual advantage estimation. At a shared history, the expected reward difference between an action and policy-sampled alternatives equals the acting agent’s advantage. We formalize these comparisons using execution graphs to represent agent interactions, including sequential, parallel, and multi-round structures, with potentially distinct rewards for different agents. Each candidate action completes the downstream execution it induces and is scored by the acting agent’s reward, and we establish sampling and replay conditions under which the resulting group-relative estimator is unbiased after a fixed group-size correction. To keep collection affordable, candidates branch from decision states saved during seed executions and their continuations run concurrently. Across mathematical reasoning, code generation, and competitive games, RollCredits leads in four code metrics, three math metrics, and three games, including a held-out game. Diagnostics against an independently sampled reference show nearly unbiased credit, while local feedback retains bias despite additional sampling. In controlled math and game experiments, saved-state reuse reduces prompt recomputation by 24–25% compared with ancestor replay.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.