acceptodds
Under review as a conference paper at ICLR 2027

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Abstract

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of and , WebShop success rates of and , and Search-QA aggregate accuracies of and , respectively. On 3B WebShop, UniOPSD improves over SDAR by percentage points. We will open-source the code upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.