acceptodds
Under review as a conference paper at ICLR 2027

Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?

Abstract

On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored. In this work, we instantiate Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information. We observe that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal less informative. We term this phenomenon decoding collapse, a failure mode that can be missed by aggregate evaluation metrics. To understand why this occurs, we show that although the self-teacher achieves stronger performance, learning remains unstable under inconsistent supervision signals. We further analyze two sources of inconsistency, model inconsistency and prompt inconsistency, and find that prompt-conditioned feedback can degrade the quality of the supervision signal, limiting the effectiveness of self-teacher learning. To mitigate this inconsistency, we introduce an exponential moving average (EMA) teacher to stabilize the self-teacher and provide more consistent supervision signals. Although the EMA teacher requires a warm-up phase during which performance may temporarily regress, it partially improves FA-SD by providing more stable supervision, while leaving a gap to stronger RL and teacher-distillation baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.