On-Policy Self-Distillation with Outcome-Aligned Privileged Information for Agentic Reinforcement Learning
Abstract
Agentic reinforcement learning (RL) trains large language model agents through long-horizon interactions, but sparse trajectory-level rewards provide limited guidance on individual decisions. Recent on-policy self-distillation (OPSD) methods supplement sparse rewards with dense supervision derived from natural-language skills or reflections used as training-time privileged information (PI). However, discrete textual PI prevents gradients from flowing directly through the generated tokens to optimize PI construction. Learnable latent PI alleviates this bottleneck, but learnability alone does not ensure that its influence on sampled actions is aligned with relative trajectory outcomes. To address this gap, we propose PILOT (Privileged Information Learning with Outcome-aligned Training), a framework that learns outcome-aligned latent PI to guide on-policy self-distillation for agentic RL. PILOT constructs latent PI from source trajectories and measures its influence on distinct same-task query trajectories through sampled-action log-likelihood differences with and without PI. With the actor held fixed, it trains the PI analyzer to align each query’s average bounded influence with relative outcome credit derived from the same-task on-policy rollout group. The measured influence weights an OPSD objective optimized jointly with direct RL, yielding an actor that requires neither privileged inputs nor the analyzer at inference. We evaluate PILOT with Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on ALFWorld and WebShop. PILOT outperforms the evaluated RL and OPSD baselines on all aggregate task metrics across both benchmarks and model scales.Our code is available at https://anonymous.4open.science/status/PILOT-478B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.