AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic Reinforcement Learning
Abstract
Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards, which assign a uniform advantage to every step and cannot identify which decisions led to success or failure. Self-distillation methods can provide finer-grained supervision by augmenting RL with privileged information. However, existing approaches usually apply the same privileged information uniformly to every step, ignoring an asymmetry between steps: routine steps need little additional guidance, while critical error steps need explicit corrective direction. We propose AHEAD, a step-aware self-distillation method that matches different supervision sources to different step types. The teacher receives environment feedback on all steps as a grounded dense signal, and additionally receives LLM-generated corrective hints on error steps to supply the direction that environment feedback lacks. The method modifies only the GRPO advantage. On ALFWorld, WebShop, and Search-based QA at three model scales, AHEAD improves over GRPO ( points on ALFWorld and on WebShop success rate at 7B) and over self-distillation baselines on ALFWorld at 7B and 1.7B. It also reaches a given success rate in fewer training steps and solves tasks in fewer interaction steps than GRPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.