acceptodds
Under review as a conference paper at ICLR 2027

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Abstract

Reinforcement learning(RL) with verifiable rewards constructs trajectory-level advantage, yet often fails to credit the few pivotal decisions that drive outcomes in long-horizon multi-turn agentic RL. Some recent works introduce privileged self-distillation into credit assignment for RL, offering denser supervision, but it still remains unclear how such a local signal should express sequential credit. We therefore propose AgentOPSD, a critic-free recursive turn-level credit assignment for agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence, recursively updates a Bayesian belief state in log-odds space. This provides a principled reweighting scheme that transforms sparse outcome supervision into turn-level credit signals and identifies pivotal turns by the marginal revision between consecutive states while remaining fully compatible with standard policy optimization, requiring neither additional rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA with three Qwen model scales (Qwen2.5-3B/7B and Qwen3-1.7B). AgentOPSD improves over GRPO and strong self-distillation baselines, reaching 89.1% success on ALFWorld with Qwen2.5-7B, and ablations attribute the gains to turn-level aggregation and history-dependent recursive belief updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.