State-Grounded Credit Assignment for Dual-Control Tool Agents: Separating Progress from Attribution
Abstract
Multi-turn tool agents in dual-control environments complete tasks jointly with users who can also act on a shared world. Standard GRPO with terminal rewards evaluates the resulting joint trajectory but does not distinguish progress produced directly by the agent from progress made through user-side operations. We introduce state-grounded process rewards that measure progress toward target states constructed from reference actions, assigning direct credit to assistant-side changes and factual delegated credit to user-side changes. We further introduce delegated credit estimation (DCE), which adjusts user-side credit by replaying user turns after replacing the assistant’s current message with a neutral acknowledgement. Across three -bench-Verified domains, Qwen3.5-9B jointly trained with direct and factual delegated supervision achieves higher mean scores than Terminal GRPO and an adapted tool-call-matching baseline across all reported metrics. On Telecom, the online variant of DCE reaches 84.8% Pass (success in all five trials), exceeding Terminal GRPO by 9.1 percentage points and factual delegated supervision by 2.1 points, with comparable single-trial success to the latter. Validation with separate replays shows that online DCE selectively reduces credit for user-side progress that recurs under the neutral message. These results support user-side state supervision and contextual refinement of delegated credit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.