acceptodds
Under review as a conference paper at ICLR 2027

The Environment Knows Why: Bidirectional Self-Distillation from Augmented Tool Feedback for Agentic RL

Abstract

Self-distillation with privileged context has emerged as a promising paradigm for agentic reinforcement learning, providing denser supervision for long-horizon credit assignment. However, existing methods typically construct such context for individual tasks or trajectories, limiting reuse and requiring repeated generation during training. We observe that rejected tool calls expose a reusable source of local supervision: the relevant tool contract is encoded in the environment implementation but often omitted from the returned feedback. Building on this observation, we introduce (*Bidirectional Environment-Augmented Distillation*), which combines task-agnostic environment augmentation with bidirectional feedback self-distillation. The former mines tool contracts offline and converts rejected feedback into reusable local guidance across tasks and trajectories. The latter uses the same augmented feedback bidirectionally: backward scoring provides hindsight correction for the rejected call, while forward scoring provides recovery guidance for the subsequent action. The resulting token-level teacher–student log-probability gaps serve as self-distillation signals for RL advantage reweighting. Extensive experiments demonstrate the effectiveness of across both in-distribution and out-of-distribution multi-turn tool-use settings, while analyses support the complementary roles of backward and forward supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.