acceptodds
Under review as a conference paper at ICLR 2027

State-Grounded Credit Assignment for Dual-Control Tool Agents: Separating Progress from Attribution

Abstract

Multi-turn tool agents in dual-control environments complete tasks jointly with users who can also act on a shared world. Standard GRPO with terminal rewards evaluates the resulting joint trajectory but does not distinguish progress produced directly by the agent from progress made through user-side operations. We introduce state-grounded process rewards that measure progress toward target states constructed from reference actions, assigning direct credit to assistant-side changes and factual delegated credit to user-side changes. We further introduce delegated credit estimation (DCE), which adjusts user-side credit by replaying user turns after replacing the assistant’s current message with a neutral acknowledgement. Across three -bench-Verified domains, Qwen3.5-9B jointly trained with direct and factual delegated supervision achieves higher mean scores than Terminal GRPO and an adapted tool-call-matching baseline across all reported metrics. On Telecom, the online variant of DCE reaches 84.8% Pass (success in all five trials), exceeding Terminal GRPO by 9.1 percentage points and factual delegated supervision by 2.1 points, with comparable single-trial success to the latter. Validation with separate replays shows that online DCE selectively reduces credit for user-side progress that recurs under the neutral message. These results support user-side state supervision and contextual refinement of delegated credit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.