Shared Resources Distort Credit Assignment in Group-Relative Agent Reinforcement Learning
Abstract
Group-relative reinforcement learning trains language-model agents by comparing rewards across rollouts of the same prompt. We identify resource-state interference, in which shared quotas, queues, and mutable tool state make a rollout's reward depend on other rollouts' actions. A rollout can thus influence its own evaluation baseline by altering sibling rewards. Our theoretical analysis distinguishes independent-deployment performance from shared-service welfare and provides finite-group constructions in which expected group-relative updates oppose the gradient of either objective. To address this problem, we develop receipt-conditioned credit control, using resource receipts to identify affected rewards and selectively replay trajectories from reference states. This enables either censoring or propensity-corrected credit estimation. Resource receipts and fixed-action paired replay isolate execution effects on outcomes, credit, and parameter gradients. Search audits on Qwen3 and Llama-3.1 show that shared execution distorts raw parameter gradients. In matched training with Qwen3-8B and Llama-3.1-8B, shared execution lowers retrieval NDCG@3 by 10.74 and 35.63 points, respectively, relative to isolation. Partial replay improves raw-gradient alignment with paired isolated execution. In separate matched training experiments, it improves retrieval performance over shared execution while using approximately 38% fewer training retrieval executions than isolation on both models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.