acceptodds
Under review as a conference paper at ICLR 2027

VALVE: Critic-Guided On-Policy Distillation for Agentic Reinforcement Learning

Abstract

On-policy distillation (OPD) provides dense supervision for training large language model agents, but enforcing teacher preferences can suppress valid alternatives in tasks with multiple solution paths. The challenge is to distinguish disagreements that warrant correction from those that reflect successful alternative behaviors. Our analysis shows that student entropy and teacher–student divergence provide limited separation between these cases, while the relative performance of fixed distillation strengths reverses over training. We propose VALVE, a critic-guided, return-aware method that adapts teacher supervision when combining OPD with reinforcement learning. VALVE uses action-level advantages from the proximal policy optimization (PPO) critic as return credit for task completion, with their signs indicating whether interaction actions are estimated to yield higher or lower return than expected from the current state. It combines this return credit with token-level teacher support, measured by teacher–student log-probability gaps, assigning larger OPD weights when the two signals agree in sign and smaller weights when they conflict. The weights adapt across tokens, interaction steps, and training stages, while retaining the standard PPO actor loss and requiring no additional rollouts or model evaluations beyond PPO and OPD. Across ALFWorld and WebShop with three Qwen students ranging from 1.7B to 7B parameters, VALVE achieves the best or tied-best task success among the evaluated methods. On ALFWorld, it outperforms the strongest competing baseline by 5.3–11.7 percentage points across students. With Qwen2.5-3B, VALVE also yields greater solution diversity than OPD and GRPO on ALFWorld. It reaches the GRPO teacher’s success rate on both benchmarks using 66–68% fewer GPU-hours than GRPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.