SCOUT: State-Conditional Counterfactual Privilege Distillation for Agentic Reinforcement Learning
Abstract
On-policy self-distillation gives a teacher privileged task information while training the student on its own interaction histories. Yet a correct solution does not need describe how to proceed after the student makes a mistake. We introduce State-conditional COUnterfactual privilege disTillation (SCOUT), which adapts privileged guidance to the student's current state. A frozen teacher scores the same student history and token prefix under three views: a successful task plan (Gold View), reusable skills with preconditions and recovery steps (Skill View), and action structure with instance-specific privileged content removed (Neutral View). An online buffer supplies skills from successful training trajectories. A parameter-free gate uses environment-specific compatibility and applicability signals to mix the three teacher distributions into one distillation target. All additional components are training-only and inference uses the student alone. Across three backbones on ALFWorld, WebShop, and search-based question answering, SCOUT improves aggregate performance over Gold-only distillation both alone and with Group Relative Policy Optimization (GRPO). With Qwen3.5-4B and GRPO, it raises ALFWorld success from 88.2% to 93.2%, WebShop success from 68.7% to 77.4%, and search-QA exact match from 45.74% to 49.93%. Ablations on Qwen3.5-4B support adaptive weighting and retaining all three views, although gains vary across task categories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.