acceptodds
Under review as a conference paper at ICLR 2027

IOD: An Iterative On-Policy Distillation Framework for Self-Improving Language Models

Abstract

On-policy knowledge distillation reduces the mismatch between student training and inference by learning from student-generated trajectories. However, extending distillation across multiple student updates raises the question of how previously generated trajectories should be managed as the student evolves. We propose Iterative On-Policy Distillation (IOD), a multi-cycle framework that combines teacher-based trajectory filtering with cross-cycle accumulation. At each cycle, IOD jointly re-evaluates previously retained and newly generated trajectories using a fixed teacher, then updates the student through token-level distribution matching while preserving the original human-labeled data. Experiments on summarization and Machine translation, using encoder–decoder and decoder-only teacher–student architectures, show higher task-specific scores than reproduced distillation baselines under the evaluated settings. Ablation studies further examine the effects of trajectory filtering, accumulation, and the student-update objective. These findings highlight trajectory management as an important consideration when extending on-policy distillation across successive student updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.