On-policy Instruction Tuning through Few-Shot Self-Distillation
Abstract
Pretrained large language models are typically adapted into assistants through an initial phase of supervised fine-tuning (SFT) on instruction-response demonstrations. Recent work suggests that the on-policy nature of reinforcement learning can help preserve capabilities acquired during previous post-training phases, motivating the question of whether instruction tuning can similarly benefit from on-policy training. We investigate this question by introducing Few-Shot Self-Distillation (FSSD), a method for on-policy instruction tuning of pretrained base models. FSSD conditions a teacher copy of the base model on a small set of relevant demonstrations and uses the teacher to provide token-level supervision on trajectories generated by the student. The teacher does not receive the reference response for the current instruction; instead, few-shot conditioning provides the supervision used to train the student. Across three model scales, FSSD helps mitigate regressions on general-capability benchmarks relative to standard SFT, while achieving stronger open-ended assistant performance on AlpacaEval 2.0 and Arena-Hard v2 and inducing substantially less distributional shift, as measured by the KL divergence from the pretrained policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.