Individual Credit, Joint Effects: Understanding Credit Assignment in Agentic RL
Abstract
Credit-assignment methods give each action of an agent its own learning signal, and most recent work improves how these signals are constructed. How the signals change the policy once they enter training, which should inform their design, remains poorly understood. In GiGPO training on a controlled agentic task, 37.5% of actions with nonzero credit change probability against it. To trace why, we decompose each Adam update into contributions from individual current and past actions, which we call sources. After the first update, removing the four most strongly opposing sources restores the credit direction in 75.8% of reversals with count- and norm-matched controls, against 16.7% for other opposing sources. From the third update on, most reversals also require removing historical sources, partly because past gradients gain weight in Adam's first moment. Compared with candidate actions, the removed sources more often perform the same kind of operation as the target when their credit is opposite, and share one of its immediate effects when their credit has the same sign. A credit signal may thus reach related actions in these two ways, within an update and, through momentum, in later ones, which raises the question of how much history the optimizer should retain. We also test a history-retention adjustment that uses only existing optimizer state: with GiGPO and GRPO on ALFWorld and WebShop, uniformly attenuating retained momentum has no consistent effect on final success, consistent with attenuation scaling both helpful and harmful history equally.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.