acceptodds
Under review as a conference paper at ICLR 2027

Inside Failed Rollouts: Progressive Experience Self-Distillation for Agentic Training

Abstract

On-policy self-distillation complements sparse outcome rewards with dense supervision by re-scoring an agent's own rollouts under privileged training-time context. Failed rollouts are particularly challenging because useful progress and harmful decisions often coexist within the same trajectory. We introduce Progressive Experience Self-Distillation (PESD), a failure-aware framework that uses privileged likelihood shifts to separate realized actions into supported, rejected, and uncertain regions. Supported actions receive forward distillation, consistently rejected actions receive low-weight unlikelihood updates, and uncertain actions are left to RL. PESD further organizes outcome feedback, applicability-gated skills, and verified successful trajectories into a progressive privilege hierarchy, with stronger guidance reduced across repeated task visits and removed at inference. On ALFWorld, PESD reaches 92.2/66.1 with Qwen2.5-3B-Instruct/Qwen3-1.7B, and 46.0/44.0 on Search-QA. These results show that failed rollouts can provide useful action-level supervision when privileged signals are applied selectively rather than uniformly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.