acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Self-Distillation via Accumulated Experience Representation

Abstract

On-Policy Self-Distillation (OPSD) has achieved strong performance in languagemodel post-training by distilling reference-conditioned teacher predictions along student-generated trajectories. However, existing methods still typically rely on explicit references or solutions to construct supervision, leaving open how to further improve self-distillation without introducing such additional supervision. In this work, we investigate whether a model can use its own past responses as experience and turn them into self-distillation supervision. We find that a single self-generated experience can substantially alter the current prediction, but this effect is sensitive to which experience is sampled; aggregating multiple experiences markedly reduces this prediction variation. Based on this observation, we propose Accumulated Experience Self-Distillation (AESD). AESD aggregates predictions induced by multiple past experiences and progressively accumulates their predictive effects throughout training through a shared set of continuously updated Soft tokens, forming a reusable distillation target. The resulting target is then distilled into the student along its own on-policy trajectories, while neither the experience buffer nor the Soft tokens are required at inference time. Across seven recent mathematical reasoning benchmarks, AESD outperforms competitive baselines on both Qwen3-4B and Qwen3-8B, improving average accuracy by 1.34 and 0.69 percentage points, while also increasing average APT by 2.43% and 2.06%, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.