acceptodds
Under review as a conference paper at ICLR 2027

ProjOPD: Conflict Projection for On-Policy Distillation in Agentic Reinforcement Learning

Abstract

Agentic reinforcement learning typically provides only a single scalar reward at the end of an episode, so the many intermediate decisions within a trajectory all share one signal. Recent work supplements this sparse supervision with on-policy self-distillation, but existing approaches mainly study how to construct the teacher signal or how to weight individual tokens, and pay little attention to whether the auxiliary update is compatible with the task objective. We identify an optimization-level conflict between the two: a context teacher expresses nothing more than the local preference of the same policy under an additional prompt, and its gradient can have a negative inner product with the policy gradient; adding the two directly then increases the RL objective to first order, i.e. the auxiliary signal cancels the descent direction dictated by the environment reward. We therefore propose ProjOPD, which builds both logit gradients explicitly on a shared sparse support, removes from the distillation gradient the component opposing the policy gradient within each environment turn, retains only its orthogonal part, and back-propagates through a detached surrogate. The projection is applied per environment turn—a turn corresponds to one complete environment action, whereas deciding token by token ignores how the tokens of a single action jointly point. The construction guarantees to first order that the auxiliary signal cannot cancel the descent direction of the policy gradient, and requires no additional rollouts. In addition, we propose a cheaper teacher construction: the teacher is simply the current policy conditioned on a plan context that the policy itself generates; it only re-scores the tokens the student has already sampled and never re-executes the environment. At test time the planner and the teacher are removed, so inference cost matches an ordinary actor. We evaluate ProjOPD on ALFWorld, WebShop and Search-QA with Qwen2.5-3B and 7B, where it attains the best result on all three benchmarks at 7B; ablations attribute the gains to the conflict projection and to the turn-level projection granularity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.