The Environment Knows Why: Bidirectional Self-Distillation from Augmented Tool Feedback for Agentic RL
Abstract
Self-distillation with privileged context has emerged as a promising paradigm for agentic reinforcement learning, providing denser supervision for long-horizon credit assignment. However, existing methods typically construct such context for individual tasks or trajectories, limiting reuse and requiring repeated generation during training. We observe that rejected tool calls expose a reusable source of local supervision: the relevant tool contract is encoded in the environment implementation but often omitted from the returned feedback. Building on this observation, we introduce (*Bidirectional Environment-Augmented Distillation*), which combines task-agnostic environment augmentation with bidirectional feedback self-distillation. The former mines tool contracts offline and converts rejected feedback into reusable local guidance across tasks and trajectories. The latter uses the same augmented feedback bidirectionally: backward scoring provides hindsight correction for the rejected call, while forward scoring provides recovery guidance for the subsequent action. The resulting token-level teacher–student log-probability gaps serve as self-distillation signals for RL advantage reweighting. Extensive experiments demonstrate the effectiveness of across both in-distribution and out-of-distribution multi-turn tool-use settings, while analyses support the complementary roles of backward and forward supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.