SMC-KD: Knowledge distillation using Teacher-Guided Sequential Monte Carlo
Abstract
Small language models are becoming more capable, and in many domains they can generate a correct answer across multiple samples, even when they do not do so reliably in a single attempt. Distillation from a stronger teacher can turn these occasional successes into more reliable behavior, but its effectiveness depends on which responses the student is trained on. For example, fine-tuning on teacher-generated responses can lead to exposure bias: the student trains on prefixes it would not produce itself, which differ from those it sees at inference. Training on the student's own responses avoids this mismatch but needs feedback on those responses. Selecting good responses with Best-of- (BoN) or verifier-guided search typically relies on correctness labels or separately trained reward models. On-policy distillation (OPD) instead has the teacher score each student token, but a poorly matched teacher can hinder performance, and different teacher and student vocabularies complicate token-level supervision. In this paper, we introduce teacher-guided Sequential Monte Carlo Knowledge Distillation (SMC-KD), a method that uses teacher feedback to identify desirable trajectories among those generated by the student and distills them into the student's parameters. SMC-KD requires no correctness labels, process reward model, or semantic step boundaries, and naturally supports teacher and student models with different vocabularies. The student generates a population of partial trajectories, the teacher scores them by the teacher–student likelihood ratio, and SMC resampling allocates subsequent generation to higher-scoring prefixes. We then train the student on these trajectories through supervised fine-tuning (SFT). Before any training, we show that teacher-guided SMC selects more accurate responses than BoN under the same likelihood score, for every student–teacher pairing and population size we evaluate. After training, SMC-KD matches or improves on the base student in every setting and outperforms OPD in 10 of 12 comparisons across four student–teacher pairings and three math benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.